Skip to content
View joshuaswarren's full-sized avatar

Sponsors

@earlvanze

Block or report joshuaswarren

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
joshuaswarren/README.md

Joshua Warren

Sponsor Buy Me a Coffee

Open-source ML systems, on-device inference, and developer tools.

I build the software between a model and the machine it runs on: GPU backends, accelerator drivers, compilers, inference servers, and the tools that make them usable.

I'm an Omarchy team member working on Omarchy M, bringing MLX and Apple Neural Engine workloads to Apple silicon Linux. I also build local-first agent infrastructure and native Mac developer tools.

MLX, Apple silicon, and local inference

omarchy-mlx — MLX on the Apple GPU under Linux

The familiar import mlx.core as mx, backed by Vulkan compute through Mesa's Honeykrisp driver rather than Metal. The project maintains a pinned upstream MLX overlay and patch set, with GPU execution on supported Apple silicon Linux machines.

My work spans GPU kernels and fusion, quantized inference, numerical validation, model serving, and distribution. The stack includes local chat, memory-aware model admission, approval-gated downloads, and reproducible release wheels. Its private Vulkan driver stays separate from the desktop's Mesa installation.

The speech path combines a Parakeet encoder on the Apple Neural Engine with decoding on the GPU.

mesa (honeykrisp-omarchy) — Honeykrisp Vulkan for Apple GPU compute

Fork of Mesa carrying Honeykrisp Vulkan patches that omarchy-mlx performance depends on: correctness fixes, cooperative matrix, and CDM work. Upstreamable series lives on upstream/correctness.

omarchy-ane + mil-hwx-compiler — the accelerator stack

omarchy-ane extends the original eiln/ane work with additional M1/M2-family bring-up, Linux kernel modules, a userspace library, firmware loading, device-tree integration, and hardware validation. Tested configurations include bit-exact operator checks and Parakeet encoder comparisons against macOS references.

mil-hwx-compiler compiles a supported subset of Core ML's textual MIL into ANE program formats without invoking Apple's compiler. Its H13 output runs on M1 hardware under Linux; other backends have their own documented validation status. Static shapes, supported operations, and numerical envelopes are explicit.

The work extends through installation and updates: DKMS, package-owned device-tree overlays, boot-chain verification, and opt-in Omarchy packages for the runtime, driver, and compiler.

Coreglass — understanding inference performance

Coreglass connects hardware signals to real model workloads: prefill and decode rates, time to first token, token gaps, energy per token, CPU time, and available GPU/ANE counters. It captures remote machines over SSH, compares engines, tracks regressions, and produces both human-readable visualizations and machine-readable findings.

Measured, modeled, replayed, and synthetic data are labeled separately.

Related work: pinned, cross-platform inference benchmarks and Linux ANE experiments.

Selected upstream contributions

Project Contribution Status
oMLX Persistent reuse of compiled ANE programs, with cross-process locking, invalidation, and fallback handling. Merged
mlx-serve Linux/Vulkan port of the Zig-based server, including MLX/MLX-C integration, platform compatibility, and end-to-end serving validation. Merged
Omarchy M Package-owned device-tree overlays and DKMS-aware boot verification. Merged
Omarchy packages Integration of the MLX runtime, private Vulkan driver, ANE driver, and MIL compiler into an opt-in package set. Merged
oMLX runtime observability Expose the effective DFlash engine and fallback reason rather than only the requested configuration. Open PR

A measured result: my oMLX ANE compile-cache contribution reduced fresh-process model-load time by 53–66% in the documented Qwen3.8-27B tests across M1 Max, M2 Max, and M1 Ultra. On the M1 Ultra, a cold cache miss took 65.17 seconds versus 22.42 seconds for a warm hit, with identical response text. Benchmark details and failure-path tests.

Apple-platform developer tools

An iOS development workflow from Linux: Swift toolchains, SDK setup, physical-device deployment, debugging, and signing and packaging checks. The documented device workflow builds and installs SwiftUI apps without running Xcode or macOS; it still uses an Apple-supplied SDK.

My related xtool work addresses dependency resolution, dynamic-library linking and embedding, and app-extension packaging.

A native Mac terminal for coding-agent work across local and remote machines. It combines a custom VT engine and Metal renderer with real PTYs, direct SSH, workspace organization, agent attention routing, and an MCP control surface. It builds and runs on real hardware.

Memory and context for user-aware agents

I'm the creator and maintainer of Remnic: open-source, local-first memory and context shared across coding assistants and other AI agents.

Memories remain inspectable Markdown files. Hybrid retrieval, graph recall, provenance, correction, and MCP/HTTP integrations make context useful without locking it inside one vendor's conversation history. Integrations include Claude Code, Codex CLI, OpenClaw, Cursor, Replit, Pi, and OMP.

As of September 2026, Remnic was seeing approximately one million monthly package downloads across its integrations.

Remnic Canvas took first place in OpenAI's 2026 Build for Good hackathon with What Helps Me, a support passport where the person approves what is shared and can revoke access. How it won.

More projects and experiments

Agent operations, browser memory, and the broader Omarchy ecosystem

Agent infrastructure

  • Remnic Canvas: shared browser-agent memory through WebMCP (first place, OpenAI Build for Good 2026).
  • modelctl: an in-development, provider-neutral control plane for workload contracts, model state, budgets, and bounded agent recovery. The public repository contains generic contracts and synthetic fixtures.
  • Tower: self-hosted agent-fleet monitoring with heartbeats, run-state transitions, stale-state detection, and artifact-linked receipts.
  • Fleet Shepherd: a read-only Omarchy panel for local and SSH-connected agent fleets, with partial-failure and stale-data handling.

Desktop tools and creative software

Omaloop turns an Omarchy theme into a playable groovebox using a Rust synthesizer and PipeWire. Dealt is a daily music-making tool built with deterministic generation and Web Audio, without a model in the product.

Other Omarchy projects include Omastorm, Omarchy Chase, Apple Bridge, Remnic integration, Hardstop, Shiplog, CodexBar integration, YouTube Mini, Plex Mini, and SomaFM.

Contributions beyond my own projects

I also contribute fixes and integration work to other maintainers' projects, including HookEcho, Blip, and Infinitty.

How I work

I use coding agents throughout the development process. My focus is defining the problem, designing the interfaces and constraints, investigating failures, and making results inspectable: hardware tests, reproducible benchmarks, numerical checks, regression gates, and explicit support boundaries.

Across these projects I work with C/C++, Objective-C++, Python, Swift, Zig, Rust, and TypeScript, alongside Linux, Vulkan, Metal, MLX, and the Apple Neural Engine.

Background and interests

I bring more than 25 years of software delivery, architecture, integrations, and developer education. Earlier work included founding and leading Creatuity, serving as the founding chair of the Magento Association, and speaking at more than 20 conferences during the Magento 2 transition.

Today I'm interested in hands-on ML systems, inference optimization, on-device ML, compilers and runtimes, and open-source developer tooling roles. Based in Dallas, Texas; focused on remote opportunities.

Website and field notes · LinkedIn · X

Support

Every bit of support helps keep joshuaswarren alive and free. If you are able, sponsor on GitHub or send a Lightning donation to joshuaswarren@strike.me to directly fund continued development and new integrations.

Sponsor Buy Me a Coffee

If financial support is not an option, you can still make a big difference: star the repo, share it, or recommend it to a colleague. Word of mouth is how most people find joshuaswarren.

Pinned Loading

  1. remnic remnic Public

    Open-source memory and context for user-aware agents: scoped memory, provenance, retrieval quality, correction, boundaries, evals, and MCP/HTTP access.

    TypeScript 214 27

  2. ane-linux-experiments ane-linux-experiments Public

    Apple Neural Engine on Linux: verified gemm/softmax/attention/transformer-block primitives, driven from Python on an M1 under Omarchy Mac

    Python 18

  3. omarchy-mlx omarchy-mlx Public

    MLX-compatible Vulkan and ANE backends for Apple Silicon running Linux.

    Python 83 9

  4. mil-hwx-compiler mil-hwx-compiler Public

    Forked from maderix/mil-hwx-compiler

    Research MIL-to-ANEC/HWX compiler for the M1 and M4 Apple Neural Engines

    C++ 1

  5. omarchy-ane omarchy-ane Public

    Forked from eiln/ane

    Apple Neural Engine support for Omarchy Linux: DRM accelerator driver and userspace library, M1 T8103

    C 6 1

  6. mesa mesa Public

    Forked from intel-lgci-fdo-gitlab-mirror/mesa.mesa

    Mirror: https://gitlab.freedesktop.org/mesa/mesa

    C