Open-source ML systems, on-device inference, and developer tools.
I build the software between a model and the machine it runs on: GPU backends, accelerator drivers, compilers, inference servers, and the tools that make them usable.
I'm an Omarchy team member working on Omarchy M, bringing MLX and Apple Neural Engine workloads to Apple silicon Linux. I also build local-first agent infrastructure and native Mac developer tools.
omarchy-mlx — MLX on the Apple GPU under Linux
The familiar import mlx.core as mx, backed by Vulkan compute through Mesa's Honeykrisp driver rather than Metal. The project maintains a pinned upstream MLX overlay and patch set, with GPU execution on supported Apple silicon Linux machines.
My work spans GPU kernels and fusion, quantized inference, numerical validation, model serving, and distribution. The stack includes local chat, memory-aware model admission, approval-gated downloads, and reproducible release wheels. Its private Vulkan driver stays separate from the desktop's Mesa installation.
The speech path combines a Parakeet encoder on the Apple Neural Engine with decoding on the GPU.
mesa (honeykrisp-omarchy) — Honeykrisp Vulkan for Apple GPU compute
Fork of Mesa carrying Honeykrisp Vulkan patches that omarchy-mlx performance depends on: correctness fixes, cooperative matrix, and CDM work. Upstreamable series lives on upstream/correctness.
omarchy-ane + mil-hwx-compiler — the accelerator stack
omarchy-ane extends the original eiln/ane work with additional M1/M2-family bring-up, Linux kernel modules, a userspace library, firmware loading, device-tree integration, and hardware validation. Tested configurations include bit-exact operator checks and Parakeet encoder comparisons against macOS references.
mil-hwx-compiler compiles a supported subset of Core ML's textual MIL into ANE program formats without invoking Apple's compiler. Its H13 output runs on M1 hardware under Linux; other backends have their own documented validation status. Static shapes, supported operations, and numerical envelopes are explicit.
The work extends through installation and updates: DKMS, package-owned device-tree overlays, boot-chain verification, and opt-in Omarchy packages for the runtime, driver, and compiler.
Coreglass — understanding inference performance
Coreglass connects hardware signals to real model workloads: prefill and decode rates, time to first token, token gaps, energy per token, CPU time, and available GPU/ANE counters. It captures remote machines over SSH, compares engines, tracks regressions, and produces both human-readable visualizations and machine-readable findings.
Measured, modeled, replayed, and synthetic data are labeled separately.
Related work: pinned, cross-platform inference benchmarks and Linux ANE experiments.
| Project | Contribution | Status |
|---|---|---|
| oMLX | Persistent reuse of compiled ANE programs, with cross-process locking, invalidation, and fallback handling. | Merged |
| mlx-serve | Linux/Vulkan port of the Zig-based server, including MLX/MLX-C integration, platform compatibility, and end-to-end serving validation. | Merged |
| Omarchy M | Package-owned device-tree overlays and DKMS-aware boot verification. | Merged |
| Omarchy packages | Integration of the MLX runtime, private Vulkan driver, ANE driver, and MIL compiler into an opt-in package set. | Merged |
| oMLX runtime observability | Expose the effective DFlash engine and fallback reason rather than only the requested configuration. | Open PR |
A measured result: my oMLX ANE compile-cache contribution reduced fresh-process model-load time by 53–66% in the documented Qwen3.8-27B tests across M1 Max, M2 Max, and M1 Ultra. On the M1 Ultra, a cold cache miss took 65.17 seconds versus 22.42 seconds for a warm hit, with identical response text. Benchmark details and failure-path tests.
An iOS development workflow from Linux: Swift toolchains, SDK setup, physical-device deployment, debugging, and signing and packaging checks. The documented device workflow builds and installs SwiftUI apps without running Xcode or macOS; it still uses an Apple-supplied SDK.
My related xtool work addresses dependency resolution, dynamic-library linking and embedding, and app-extension packaging.
A native Mac terminal for coding-agent work across local and remote machines. It combines a custom VT engine and Metal renderer with real PTYs, direct SSH, workspace organization, agent attention routing, and an MCP control surface. It builds and runs on real hardware.
I'm the creator and maintainer of Remnic: open-source, local-first memory and context shared across coding assistants and other AI agents.
Memories remain inspectable Markdown files. Hybrid retrieval, graph recall, provenance, correction, and MCP/HTTP integrations make context useful without locking it inside one vendor's conversation history. Integrations include Claude Code, Codex CLI, OpenClaw, Cursor, Replit, Pi, and OMP.
As of September 2026, Remnic was seeing approximately one million monthly package downloads across its integrations.
Remnic Canvas took first place in OpenAI's 2026 Build for Good hackathon with What Helps Me, a support passport where the person approves what is shared and can revoke access. How it won.
Agent operations, browser memory, and the broader Omarchy ecosystem
- Remnic Canvas: shared browser-agent memory through WebMCP (first place, OpenAI Build for Good 2026).
- modelctl: an in-development, provider-neutral control plane for workload contracts, model state, budgets, and bounded agent recovery. The public repository contains generic contracts and synthetic fixtures.
- Tower: self-hosted agent-fleet monitoring with heartbeats, run-state transitions, stale-state detection, and artifact-linked receipts.
- Fleet Shepherd: a read-only Omarchy panel for local and SSH-connected agent fleets, with partial-failure and stale-data handling.
Omaloop turns an Omarchy theme into a playable groovebox using a Rust synthesizer and PipeWire. Dealt is a daily music-making tool built with deterministic generation and Web Audio, without a model in the product.
Other Omarchy projects include Omastorm, Omarchy Chase, Apple Bridge, Remnic integration, Hardstop, Shiplog, CodexBar integration, YouTube Mini, Plex Mini, and SomaFM.
I also contribute fixes and integration work to other maintainers' projects, including HookEcho, Blip, and Infinitty.
I use coding agents throughout the development process. My focus is defining the problem, designing the interfaces and constraints, investigating failures, and making results inspectable: hardware tests, reproducible benchmarks, numerical checks, regression gates, and explicit support boundaries.
Across these projects I work with C/C++, Objective-C++, Python, Swift, Zig, Rust, and TypeScript, alongside Linux, Vulkan, Metal, MLX, and the Apple Neural Engine.
I bring more than 25 years of software delivery, architecture, integrations, and developer education. Earlier work included founding and leading Creatuity, serving as the founding chair of the Magento Association, and speaking at more than 20 conferences during the Magento 2 transition.
Today I'm interested in hands-on ML systems, inference optimization, on-device ML, compilers and runtimes, and open-source developer tooling roles. Based in Dallas, Texas; focused on remote opportunities.
Website and field notes · LinkedIn · X
Every bit of support helps keep joshuaswarren alive and free. If you are able, sponsor on GitHub or send a Lightning donation to joshuaswarren@strike.me to directly fund continued development and new integrations.
If financial support is not an option, you can still make a big difference: star the repo, share it, or recommend it to a colleague. Word of mouth is how most people find joshuaswarren.





