Skip to content

feat(offload): partial GPU offload for dense qwen3.5 — host-RAM spill + optional CPU execution - #793

Open
aldrouil wants to merge 49 commits into
warpfront:masterfrom
aldrouil:feature/partial-cpu-offload
Open

aldrouil wants to merge 49 commits into
warpfront:masterfrom
aldrouil:feature/partial-cpu-offload

Conversation

@aldrouil

Copy link
Copy Markdown
Contributor

Author's note

This is a rather large change, so I am going to include a small preface here explaining the goal and purpose of this PR.

First off, like a lot of people, my PC is limited to 16 GB of VRAM. I have an AMD RX 9070 XT (gfx1201), not a datacenter class GPU with more than that like a 9070 PRO, or even a 3090, 4090, or 5090 from Nvidia (which are out of scope for hipfire but are relevant for the discussion from a VRAM standpoint). Because of that, I am limited in what I can run and fit on this PC. Dense models seam to be releasing in a few different size stages - 2b/4b models, 9b/12 models, and then at least for home users - larger 27b/30b models. The problem is that the larger 27b/30b dense models are at the capacity limit of a single 16GB GPU - even before you factor in KV cache allocation.

And unfortunately, the most average large GPUs tend to max out at 16GB of VRAM. Yes I know the 20GB 7900XT and 24GB 7900XTX exist. I don't have one. And other than the 9070 PRO, the max for the current generation of AMD GPUs is 16GB of VRAM.

If someone wants to run these models on a 16GB GPU with hipfire this leaves with two options.

  1. Tiny context window effectively useless for agentic work.
  2. Downsize to a smaller quant. I tested this myself and noticed some tool calls started to break on qwen3.8 on some of the 3 bit quants. The model would correct itself, but this is likely an artifact of a smaller quant.

Or option 3 which is use a different project like llama.cpp - but when models fit hipfire is significantly faster.

The goal of this PR is to add the ability for Hipfire to offload dense qwen3.5 based model layers - at the time of writing qwen3.5/3.6 MoE and qwen flash next are NOT supported - to system RAM. This enables two things:

  1. A longer context window than would fit otherwise.
  2. Dense models that would otherwise not fit on the relevant GPU.

Once this was done it was found that there was a signficant speed penalty for doing so. Much more than llama.cpp (using it as a reference performance target). An analysis of the reason why determined that was because llama.cpp uses it's CPU inference engine on it's CPU offloaded layers to avoid passing data back to the GPU. As it turns out, despite CPUs being much slower at inference it is faster to have a CPU do inference on locally available data rather than having the GPU read objects in system RAM over the PCIe link. As a result I had a subsequent agent add a basic SIMD AVX2 accelerated code path to hipfire to use on offloaded layers. I have tested a few of the CPU SIMD paths, but not all and welcome anyone doing so before this merge request is committed.

All that in mind there are a few design decisions I made.
a) By default hipfire defaults to it's 'old' behavior and does not attempt to offload layers to system RAM.
b) Even if you offload layers to system RAM, the CPU inference engine is disabled by default. With the CPU inference engine disabled the system passes all of it's data back to the GPU for inference - on my machine with a PCIe 4.0 x16 link the penalty was significant.

The code should be extensible to other dense LLM architectures.

There is one caveat that is worth mentioning though. CPU inference is not byte identical to GPU inference. That is not to say it cannot be, but it isn't. This was a deliberate design decision - but it should be noted that llama.cpp treats CPU and GPU inference the same way. CPU inference on llama.cpp is not byte identical to its CUDA/HIP/Vulkan, etc runtimes.

Summary

Adds opt-in partial GPU offload for the dense qwen35 family:
memory.gpu_layer_budget spills a contiguous prefix of layers to host RAM, so a
model — or a context length — that does not otherwise fit in VRAM can run; and
memory.offload_exec=cpu executes those spilled layers' weight-reading GEMVs on
the CPU instead of reading them over PCIe once per token. Both default to today's
behavior (every layer resident, weights read by the GPU), so an unconfigured run
is byte-identical to a resident one.

typed key env default meaning
memory.gpu_layer_budget HIPFIRE_GPU_LAYER_BUDGET unset → fully resident N keeps the last N layers on the GPU and spills the rest; auto/-1 defers to the engine (today: fully resident)
memory.offload_exec HIPFIRE_OFFLOAD_EXEC pcie who multiplies a spilled layer's weights: pcie = the existing GPU kernels read the host-mapped weights over the link, cpu = execute them on the CPU
HIPFIRE_CPU_EXEC_TRACE — off (developer) =1 prints one line per distinct CPU step shape

Motivation is capacity, not speed: offloaded layers are bandwidth-bound by design.
Unset, empty and unknown values fail closed (pcie, fully resident), so a typo can
neither move work to the CPU nor force an offload.

Which surface(s) does this touch?

  • kernel — crates/rdna-compute, crates/hipfire-dispatch, crates/hip-bridge
  • load — runtime load path (hfq, weight_backend), arch load*, hipfire-config, Cargo manifests
  • serve — hipfire-runtime weight-GEMV seam (llama.rs, weight_backend.rs, hfq.rs); no hipfire-engine / hipfire-generate change
  • arch crate(s): hipfire-arch-qwen35 (placement + load path); hipfire-arch-llama, hipfire-arch-qwen2 (declare no offload support)
  • crates/hipfire-quantize / quant formats
  • control plane — hipfire-tui (two Settings rows)
  • docs / CI / scripts only
  • policy files

Design notes

  • Host memory is hipHostMalloc(hipHostMallocMapped), not a host-located VMM
    arena. Measured on gfx1201, holding 4096 MB: a hipMemCreate host arena is
    charged against the device heap 1:1 (device headroom 15360 → 11264 MB), while
    hipHostMalloc costs 0 MB. The VMM host primitive and its locality tag were
    deleted so the net-zero mechanism cannot be re-adopted; the invariant is asserted
    by crates/rdna-compute/examples/host_offload_headroom.rs, because byte-identity
    and coherence tests cannot see it.
  • No kernel changes and no memory-class tag on the tensor type. Kernels keep
    dereferencing a GpuTensor; the device pointer is the host-mapped alias.
  • Placement is resolved once at load (i_gpu_start): layers [0 .. i_gpu_start)
    spill, the contiguous tail stays resident, and token_embd/output_norm/lm_head
    are always resident. The policy arithmetic
    (hipfire_config::memory::largest_fitting_tail, resolve_i_gpu_start) is pure,
    GPU-free and unit-tested. auto currently resolves to fully resident — the
    arithmetic is landed, the device measurement that would drive it is not.
  • A second reader entry point, not a host flag. HfqBackend::read_proj_host
    is Option<fn(&HfqFile, &mut Gpu, …)> because the host upload needs &mut Gpu
    and the existing read_proj fn-pointer contract cannot hand one out. A spilled
    layer whose arch supplies no host reader is a hard error, never a silent device
    allocation.
  • Fail-closed around the f32-dequant fallback: the catch-all dispatch arm and
    qt 1 Native mode have no host upload, so they refuse for offloaded layers rather
    than quietly consuming the VRAM the offload exists to free.
  • The AWQ sidecar stays device-resident on purpose: a 1-D f16 vector of length K
    would cost PCIe bandwidth per GEMV to save kilobytes.
  • hipGraph capture stays on for the pcie spill (host-mapped addresses are
    stable for the model's lifetime, replay verified byte-identical). It is disabled
    for the model's lifetime only when CPU execution is active, because a CPU step is
    a host sync point.
  • One seam for CPU execution, in hipfire-dispatch's execute_steps: a fused
    entry spanning a CPU step is not matched, and weight_gemv_swiglu_residual splits
    itself (GPU silu_mul_f32, then the CPU GEMV + residual). Attention, softmax,
    rmsnorm, RoPE, qk-norm, the KV write, the flash-attention families and the
    DeltaNet recurrence stay on the GPU, and KV residency is unchanged either way.
  • Retained replay is refused, not emulated: memory.offload_exec=cpu with a
    Redline backend fails the load naming both keys, because the tape records GPU
    launches and would replay stale activations.
  • New leaf crate hipfire-cpu (depends on rayon only, GPU-free tests). It
    transcribes the canonical decoder; crates/hipfire-runtime/tests/cpu_quant_cross_check.rs
    holds the two copies together bit-for-bit. All 28 dense-capable formats decode,
    each with an AVX2 row dot and a scalar fallback (all-scalar on ARM). Map:
    docs/quant-formats/cpu-simd-coverage.md.

Numbers

Behind/after are the same fixtures with the spill dialed in; resident control
included. Arms interleaved, fresh process per point, host otherwise idle.

fixture spill pcie cpu resident control
qwen3.5-9b.mq4 (--spec off, 128 tok) 8 / 32 layers (888 MB/token) 21.6 tok/s 24.1 tok/s (+11.6 %) 65.7 tok/s at 0 spilled
qwen3.5-9b.mq4 16 / 32 layers 11.0 tok/s 14.0 tok/s (+27 %) —
qwen3.8-27b.mq3-xt (64 layers) 8 / 64 layers (1.18–1.24 GB/token) 13.20 tok/s 15.20 tok/s (+15.2 %) 40.7 tok/s
  • Fixture identity (the recorded runs below — not the battery.json attachment, whose 9B is a
    different artifact): 9B md5 31a8d8dc7603226801b08d8319015602, prompt
    benchmarks/prompts/humaneval_3_below_zero.txt md5 37c5aad9f9efe93b5c47f27256bdf149,
    daemon md5 6b2f85588aac1abdbd233e415044eb12, hipfire md5 7075cd546b1ce61d39a36d59ce149cba.
    27B md5 80bb9198e6a565fc006b2ae1b7c89eca, probe prompt md5 5835c71e471849b4a72e1dc8e39695e7,
    daemon md5 e39c3adb85f608db1439b128471764e8, hipfire md5 57a0cf5795a8ce376d274eba78bc733a.
    Host: 1× RX 9070 XT (gfx1201, 16304 MB), ROCm/HIP 7.2, Ryzen 7 7800X3D, 28 GB RAM.
  • vram_free_mb is identical across arms at the same budget, and the 27B's own
    footprint (~13070 MB) matches stock — the resident path costs no extra VRAM. An
    offloaded=N log line is not evidence on its own; this control is.
  • The CPU path is not a throughput trick, it changes the bound: the spill costs
    ≈41 ms/token on cpu versus ≈51 ms/token on pcie, against the link's measured
    27.1 GB/s. Remaining headroom is the kernel, not the design — the same row dot
    streams a 195.8 MB host-mapped buffer at 44.5 GB/s over 16 rayon threads.
  • One earlier pass reported the cpu arm as 3–6× slower; it was measured while a
    workspace build was running on the same 16 threads. Recorded as superseded rather
    than deleted, because the way it was produced is the reusable lesson.
  • These rows come from the records below, not from the battery.json attachment — the 9B fixture
    attached there is a different artifact (md5 296092bf…) from the records' 9B (31a8d8dc…).
  • Full records, methods and raw JSON: docs/perf-checkpoints/2026-09-27-*
    (lifecycle historical, exploratory, not product claims).

Test plan

  • ./scripts/no-gpu-ci.sh passes — red on this host, for a pre-existing reason:
    its Python phase fails 5 tests in tests/test_mq4c_repack.py
    (mq4c_repack has no attribute main / HfqmError / parse_hfqm_index). Reproduces
    identically in an unmodified checkout of this repository, and this change touches
    no file under tools/ or tests/. The Rust phases of the same script are green
    (cargo check --workspace --examples, hipfire-cpu lib 26 passed, hipfire-config
    89 passed, env-docs drift clean).
  • cargo build --release clean — cargo build --release --locked -p hipfire-daemon -p hipfire-cli
    finished, warnings only
  • cargo test --lib --workspace passes — every crate's lib target green
    (hipfire-runtime 691, rdna-compute 251, hipfire-dispatch 273, hipfire-arch-qwen35 199,
    hipfire-quantize 49, …), 0 failed
  • load / serve / kernel changes: ran serve_harness.py --mode battery (and --mode chain)
    myself on hardware, with a real spill on the cpu arm: 5/5 turns each,
    runaway=0 empty=0 attractor=0 retrieval_miss=0, avg_decode=26.9 tok/s (chain: prefix-cache
    hits at 388/717/828/952 tokens), decoded and eyeballed — each turn's assistant_content is in
    the attachment below. Fixture qwen3.5-9b.mq4 sha256 ba83acf5…, md5 296092bf…;
    target/release/daemon md5 1c6e1c0a… (the harness reports the same digest it ran). The daemon
    serve log proves the spill rather than assuming it: partial offload: 24 resident / 8 offloaded, i_gpu_start=8 and cpu exec: 8/8 spilled layers fully covered; uncovered quants: none.
    Separately, redline_daemon_harness.py on the pcie arm with a spill is recorded in
    docs/perf-checkpoints/2026-09-27-gfx1201-cpu-exec-offload.md (stable=True, launches=427,
    AQL/PM4 shadow exact=True; on the cpu arm it fails the load with the intended refusal).
  • If perf-relevant: ./scripts/speed-gate.sh within ±2% of locked baselines — not run; it
    needs a 4B fixture, and the resident path this gate protects is covered by the byte-identity
    (zero-diff) guard in the records above, while the spill's cost is documented rather than hidden.
  • If this raises a ceiling in scripts/leanup-thresholds.txt: RATCHET-RAISE …
local serve_harness battery.json (load / serve / kernel changes)

Command

HIPFIRE_GPU_LAYER_BUDGET=24 HIPFIRE_OFFLOAD_EXEC=cpu \
HIPFIRE_DAEMON_BIN="$PWD/target/release/daemon" \
python3 scripts/serve_harness.py --model ~/.hipfire/models/qwen3.5-9b.mq4 \
  --mode battery --sampling greedy --speculation off --dflash off --thinking off \
  --out battery_9b_cpu_b24.json

--speculation off --dflash off --thinking off: no τ in the picture, and a reasoning
model cannot close a think block inside this budget (the daemon fails that turn closed
and releases no text). --mode chain was run with the same flags (cached=388/717/828/952,
5/5 turns, avg_decode=26.9).

Fixture identity

artifact digest
qwen3.5-9b.mq4 (5,313,750,016 B) sha256 ba83acf5bfd5d4e334b0afc26d779734e31623bb7f74e807c3581dfecb3128ad, md5 296092bf1e6a45d78c1acf815eb93366
target/release/daemon md5 1c6e1c0ae86ba651d446fe1c833a729b (the harness prints this as daemon_binary_md5; it matches the build under test)
target/release/hipfire md5 dc305c1c6de2119ed03a9376b42cfdcb
host 1× RX 9070 XT (gfx1201, 16304 MB), ROCm/HIP 7.2

This is not the qwen3.5-9b.mq4 of the perf-checkpoint records (those pin md5
31a8d8dc7603226801b08d8319015602 at 5,297,456,128 B), so these numbers are a
separate fixture and are not comparable with that record's.

The spill actually engaged (daemon serve log)

  partial offload: 24 resident / 8 offloaded, i_gpu_start=8
cpu exec: 8/8 spilled layers fully covered; uncovered quants: none
cpu exec: hipGraph capture disabled (CPU-executed steps present)

Per-turn rows

turn ctx gen finish decode tok/s ttft s attractor empty runaway
code 44 343 stop 26.9 1.026 false false false
reason 55 286 stop 26.9 0.348 false false false
factual 25 81 stop 27.1 0.163 false false false
prose 33 102 stop 26.9 0.249 false false false
instruct 31 82 stop 26.9 0.165 false false false

turns=5 runaway=0 empty=0 attractor=0 retrieval_miss=0 avg_prefill=165.1tok/s avg_decode=26.9tok/s

harness --out JSON (verbatim)

[
{
"request_id": "chatcmpl-1023366-1",
"ctx": 44,
"cached": 0,
"gen": 343,
"finish": "stop",
"think_words": 0,
"ans_words": 178,
"prefill_ms": 969.5,
"prefill_tok_s": 45.4,
"decode_tok_s": 26.9,
"decode_estimated": false,
"tau": null,
"cycles": null,
"dflash": null,
"mtp": null,
"mtp_ngram": null,
"ngram_mod_windows": null,
"ngram_mod_drafts": null,
"ngram_mod_accepted": null,
"ngram_mod_accept_rate": null,
"mtp_windows": null,
"ar_windows": null,
"mtp_retired": null,
"mtp_window_timings": null,
"ttft_s": 1.026,
"wall_s": 13.713,
"attractor": false,
"empty": false,
"runaway": false,
"ans_preview": "```python\ndef merge_sorted(a, b):\n    \"\"\"\n    Merge two already-sorted lists into one sort",
"assistant_content": "```python\ndef merge_sorted(a, b):\n    \"\"\"\n    Merge two already-sorted lists into one sorted list.\n\n    This function assumes that both input lists `a` and `b` are sorted in ascending order.\n    It returns a new list containing all elements from both inputs, also sorted in ascending order.\n    The implementation uses a two-pointer approach with O(n + m) time complexity.\n\n    Parameters:\n        a (list): First sorted list of comparable items.\n        b (list): Second sorted list of comparable items.\n\n    Returns:\n        list: A new merged sorted list containing all elements from `a` and `b`.\n\n    Example:\n        >>> merge_sorted([1, 3, 5], [2, 4, 6])\n        [1, 2, 3, 4, 5, 6]\n    \"\"\"\n    i, j = 0, 0\n    merged = []\n\n    # Traverse both lists and append the smaller element to the result\n    while i < len(a) and j < len(b):\n        if a[i] <= b[j]:\n            merged.append(a[i])\n            i += 1\n        else:\n            merged.append(b[j])\n            j += 1\n\n    # Append any remaining elements from list a\n    while i < len(a):\n        merged.append(a[i])\n        i += 1\n\n    # Append any remaining elements from list b\n    while j < len(b):\n        merged.append(b[j])\n        j += 1\n\n    return merged\n```",
"content": "```python\ndef merge_sorted(a, b):\n    \"\"\"\n    Merge two already-sorted lists into one sorted list.\n\n    This function assumes that both input lists `a` and `b` are sorted in ascending order.\n    It returns a new list containing all elements from both inputs, also sorted in ascending order.\n    The implementation uses a two-pointer approach with O(n + m) time complexity.\n\n    Parameters:\n        a (list): First sorted list of comparable items.\n        b (list): Second sorted list of comparable items.\n\n    Returns:\n        list: A new merged sorted list containing all elements from `a` and `b`.\n\n    Example:\n        >>> merge_sorted([1, 3, 5], [2, 4, 6])\n        [1, 2, 3, 4, 5, 6]\n    \"\"\"\n    i, j = 0, 0\n    merged = []\n\n    # Traverse both lists and append the smaller element to the result\n    while i < len(a) and j < len(b):\n        if a[i] <= b[j]:\n            merged.append(a[i])\n            i += 1\n        else:\n            merged.append(b[j])\n            j += 1\n\n    # Append any remaining elements from list a\n    while i < len(a):\n        merged.append(a[i])\n        i += 1\n\n    # Append any remaining elements from list b\n    while j < len(b):\n        merged.append(b[j])\n        j += 1\n\n    return merged\n```",
"reasoning_content": "",
"tool_calls": [],
"request_md5": "b2d411eb58ec2648ebd087ac828a51bb",
"atem_leak": false,
"terminal_count": 1,
"terminal_reasons": [
"stop"
],
"post_terminal_bytes": 0,
"saw_done": true,
"stream_error": null,
"prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f",
"expected_substrings": [],
"retrieval_missing": []
},
{
"request_id": "chatcmpl-1023366-3",
"ctx": 55,
"cached": 0,
"gen": 286,
"finish": "stop",
"think_words": 0,
"ans_words": 147,
"prefill_ms": 304.1,
"prefill_tok_s": 180.9,
"decode_tok_s": 26.9,
"decode_estimated": false,
"tau": null,
"cycles": null,
"dflash": null,
"mtp": null,
"mtp_ngram": null,
"ngram_mod_windows": null,
"ngram_mod_drafts": null,
"ngram_mod_accepted": null,
"ngram_mod_accept_rate": null,
"mtp_windows": null,
"ar_windows": null,
"mtp_retired": null,
"mtp_window_timings": null,
"ttft_s": 0.348,
"wall_s": 10.946,
"attractor": false,
"empty": false,
"runaway": false,
"ans_preview": "To find the total distance traveled, we need to calculate the distance for each segment of",
"assistant_content": "To find the total distance traveled, we need to calculate the distance for each segment of the trip separately and then add them together.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first part of the trip.**\n*   Speed ($v_1$) = $60$ mph\n*   Time ($t_1$) = $2.5$ hours\n$$ d_1 = 60 \\times 2.5 = 150 \\text{ miles} $$\n\n**Step 2: Calculate the distance for the second part of the trip.**\n*   Speed ($v_2$) = $40$ mph\n*   Time ($t_2$) = $1.5$ hours\n$$ d_2 = 40 \\times 1.5 = 60 \\text{ miles} $$\n\n**Step 3: Add the distances together to find the total distance.**\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 + 60 = 210 \\text{ miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210** miles.",
"content": "To find the total distance traveled, we need to calculate the distance for each segment of the trip separately and then add them together.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first part of the trip.**\n*   Speed ($v_1$) = $60$ mph\n*   Time ($t_1$) = $2.5$ hours\n$$ d_1 = 60 \\times 2.5 = 150 \\text{ miles} $$\n\n**Step 2: Calculate the distance for the second part of the trip.**\n*   Speed ($v_2$) = $40$ mph\n*   Time ($t_2$) = $1.5$ hours\n$$ d_2 = 40 \\times 1.5 = 60 \\text{ miles} $$\n\n**Step 3: Add the distances together to find the total distance.**\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 + 60 = 210 \\text{ miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210** miles.",
"reasoning_content": "",
"tool_calls": [],
"request_md5": "85b33e9aedabf7a9fa0d1d81cc6f8356",
"atem_leak": false,
"terminal_count": 1,
"terminal_reasons": [
"stop"
],
"post_terminal_bytes": 0,
"saw_done": true,
"stream_error": null,
"prompt_md5": "640e0fd4f55996cb175a422f0a12cef5",
"expected_substrings": [],
"retrieval_missing": []
},
{
"request_id": "chatcmpl-1023366-5",
"ctx": 25,
"cached": 0,
"gen": 81,
"finish": "stop",
"think_words": 0,
"ans_words": 70,
"prefill_ms": 126.1,
"prefill_tok_s": 198.2,
"decode_tok_s": 27.1,
"decode_estimated": false,
"tau": null,
"cycles": null,
"dflash": null,
"mtp": null,
"mtp_ngram": null,
"ngram_mod_windows": null,
"ngram_mod_drafts": null,
"ngram_mod_accepted": null,
"ngram_mod_accept_rate": null,
"mtp_windows": null,
"ar_windows": null,
"mtp_retired": null,
"mtp_window_timings": null,
"ttft_s": 0.163,
"wall_s": 3.122,
"attractor": false,
"empty": false,
"runaway": false,
"ans_preview": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its or",
"assistant_content": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its orbital plane around the Sun. As the planet orbits, this fixed tilt causes different hemispheres to receive varying amounts of direct sunlight and experience changes in day length throughout the year. Consequently, when a hemisphere leans toward the Sun, it experiences summer due to more intense solar radiation, while the opposite hemisphere experiences winter.",
"content": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its orbital plane around the Sun. As the planet orbits, this fixed tilt causes different hemispheres to receive varying amounts of direct sunlight and experience changes in day length throughout the year. Consequently, when a hemisphere leans toward the Sun, it experiences summer due to more intense solar radiation, while the opposite hemisphere experiences winter.",
"reasoning_content": "",
"tool_calls": [],
"request_md5": "d34383b66fd5f7af0cfced19d3777983",
"atem_leak": false,
"terminal_count": 1,
"terminal_reasons": [
"stop"
],
"post_terminal_bytes": 0,
"saw_done": true,
"stream_error": null,
"prompt_md5": "8f66b4c97988825bd8e7840aaf44357e",
"expected_substrings": [],
"retrieval_missing": []
},
{
"request_id": "chatcmpl-1023366-7",
"ctx": 33,
"cached": 0,
"gen": 102,
"finish": "stop",
"think_words": 0,
"ans_words": 81,
"prefill_ms": 211.1,
"prefill_tok_s": 156.3,
"decode_tok_s": 26.9,
"decode_estimated": false,
"tau": null,
"cycles": null,
"dflash": null,
"mtp": null,
"mtp_ngram": null,
"ngram_mod_windows": null,
"ngram_mod_drafts": null,
"ngram_mod_accepted": null,
"ngram_mod_accept_rate": null,
"mtp_windows": null,
"ar_windows": null,
"mtp_retired": null,
"mtp_window_timings": null,
"ttft_s": 0.249,
"wall_s": 4.01,
"attractor": false,
"empty": false,
"runaway": false,
"ans_preview": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against t",
"assistant_content": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against the jagged rocks below. One stormy night, a strange object tumbled onto the shore, glowing with an eerie, pulsating blue light that defied the darkness. He approached cautiously to investigate and discovered it was not debris, but a small, intricate clockwork device humming with impossible energy. Realizing this invention could change everything, Elias knew his life of solitude had just become the most significant chapter of his story.",
"content": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against the jagged rocks below. One stormy night, a strange object tumbled onto the shore, glowing with an eerie, pulsating blue light that defied the darkness. He approached cautiously to investigate and discovered it was not debris, but a small, intricate clockwork device humming with impossible energy. Realizing this invention could change everything, Elias knew his life of solitude had just become the most significant chapter of his story.",
"reasoning_content": "",
"tool_calls": [],
"request_md5": "643a78544be2dab1edcfe5b8e85c3718",
"atem_leak": false,
"terminal_count": 1,
"terminal_reasons": [
"stop"
],
"post_terminal_bytes": 0,
"saw_done": true,
"stream_error": null,
"prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817",
"expected_substrings": [],
"retrieval_missing": []
},
{
"request_id": "chatcmpl-1023366-9",
"ctx": 31,
"cached": 0,
"gen": 82,
"finish": "stop",
"think_words": 0,
"ans_words": 62,
"prefill_ms": 126.7,
"prefill_tok_s": 244.6,
"decode_tok_s": 26.9,
"decode_estimated": false,
"tau": null,
"cycles": null,
"dflash": null,
"mtp": null,
"mtp_ngram": null,
"ngram_mod_windows": null,
"ngram_mod_drafts": null,
"ngram_mod_accepted": null,
"ngram_mod_accept_rate": null,
"mtp_windows": null,
"ar_windows": null,
"mtp_retired": null,
"mtp_window_timings": null,
"ttft_s": 0.165,
"wall_s": 3.182,
"attractor": false,
"empty": false,
"runaway": false,
"ans_preview": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep",
"assistant_content": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep functions small, focused on a single responsibility, and easy to test.\n3. Write clear comments only for complex logic or non-obvious design decisions.\n4. Follow consistent coding conventions and formatting rules across the entire project.\n5. Refactor code regularly to remove technical debt and improve readability over time.",
"content": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep functions small, focused on a single responsibility, and easy to test.\n3. Write clear comments only for complex logic or non-obvious design decisions.\n4. Follow consistent coding conventions and formatting rules across the entire project.\n5. Refactor code regularly to remove technical debt and improve readability over time.",
"reasoning_content": "",
"tool_calls": [],
"request_md5": "9af4a6122e824b100c1d31cb525d7816",
"atem_leak": false,
"terminal_count": 1,
"terminal_reasons": [
"stop"
],
"post_terminal_bytes": 0,
"saw_done": true,
"stream_error": null,
"prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4",
"expected_substrings": [],
"retrieval_missing": []
}
]

Correctness on the changed surface, in more detail:

  • cargo test -p hipfire-cpu --lib — 26 passed (GPU-free; the crate is in
    scripts/no-gpu-ci.sh).
  • cargo test -p hipfire-config — 89 passed.
  • python3 scripts/check-env-docs.py — clean.
  • Production-launcher vs hipfire_cpu::gemv parity on identical bytes, both arms of
    the rotation contract: worst 9.3e-7 relative against a 1e-4 tolerance, 15 of 23
    formats bit-exact. qt 49 (every projection of the 27B MQ3 fixture, and a format
    with no canonical dequantize_to_f32 arm) is oracled on a real 27B tensor:
    1.195e-7 / 3.530e-7.
  • dequant_group reproduces the canonical decoder bit-for-bit over 1695 real
    tensors. Zero-diff guard: stock ≡ this change ≡ spill+pcie, byte-identical greedy
    text (2B 1343 chars, 9B 2734 chars).
  • The CPU path's contract is coherence plus a measured divergence, not bit-identity:
    9B, 8 spilled, greedy — both arms agree for 723 characters (~190 tokens) and then
    differ by one whitespace token; the path is deterministic across processes and
    across RAYON_NUM_THREADS=1, so it is not CPU-side scheduling. On the 27B MQ3
    fixture both completions are 610 B and byte-identical.
  • Notes: the GPU-gated parity test is --ignored and fixture-dependent, so it is not
    in the GPU-free CI set; two formats (qt 11/12) report SKIP — production launcher has no kernel on this arch on gfx12. rustfmt --check --config skip_children=true
    is clean on every file this branch touches, and clippy is an advisory job.
  • scripts/verify-bind-thread.sh (a pre-commit hook, not part of the required CI
    gates job) reports 12 of 94 pub fn in impl Gpu missing the bind_thread
    contract against 8 of 86 before this change. The four added host-mapped entry
    points were fixed in this branch — the two allocators bind the calling thread, the
    two state queries carry the gate's reasoned skip marker. The remaining 8 are
    byte-identical to the pre-existing set and are left alone.

Known limits

  • Placement is whole-layer and qwen35-only. The CPU-execution seam itself is
    arch-agnostic (hipfire-runtime/src/llama.rs routes the llama-family
    down-projection through it); llama and qwen2 declare no offload support, and
    MoE/A3B is structurally out of reach — load_moe_ffn never sees the offload flag.
  • Prefill and the slots/serve body do not use the CPU seam: batched GEMM there
    never enters execute_steps, so their spilled weights still cross PCIe.
  • KV cache is not spilled (attention stays resident).
  • Pinned host memory is unreclaimable, so keep the spilled prefix small: on the
    27B MQ3 fixture, 8 of 64 layers (≈1.2 GB pinned) loads; 16 (≈2.5 GB) stalled twice
    with 0 GB free.
  • The 27B MQ3 fixture has one decoded read, not a sweep: a greedy completion, 610 B and
    byte-identical between the cpu and pcie arms (cpu.txt / pcie.txt in
    docs/perf-checkpoints/data-2026-09-27-cpu-exec-mq3-avx2/). The earlier record's "no decoded
    text was read on this fixture" is a gap in that earlier run, closed by the later one — but a
    second prompt or genre is still absent, and there is no ≥3-process arm-crossing byte diff on
    this fixture — the cross-process evidence above is the 9B cpu arm's determinism, which is a
    different property.
  • A dead-code warning remains in the new crate (mq4g256_row_dot_scalar is now
    unreachable).

Hardware validation request (optional)

Omitted deliberately: the CPU-executed path needs HIPFIRE_GPU_LAYER_BUDGET and
HIPFIRE_OFFLOAD_EXEC set at load, which a registry-tag route cannot express, and a
route without the spill proves nothing about this change.

How this merges (direct review)

Merge authority is direct maintainer review plus the required no-GPU CI checks.
hw-gate automation is optional evidence delivery, not a prerequisite or substitute.

  1. Required CI (must stay green): build (workspace, no GPU), unit tests (lib, no GPU), gates (ratchets, layering, registers) from .github/workflows/ci.yml.
  2. Required review: one approving maintainer review. The reviewer judges claim-matched evidence for the surfaces you ticked.
  3. Evidence you owe: pick routes from docs/VALIDATION.md. Prefer attaching local harness output (serve_harness.py, redline_daemon_harness.py, test_kernels, etc.). A lone hipfire run transcript is not evidence.
  4. Optional automation: hw-gate is not a required status check.
  5. Retired paths: scripts/coherence-gate*.sh and tools/change_gate / agentic-review are historical only — never acceptance evidence.

Architecture-trait change?

No — crates/hipfire-runtime/src/arch.rs (the Architecture trait) is unchanged.

Avery Drouillard and others added 28 commits September 27, 2026 18:24
…_gpu_start

Squash of 1f75880, efa8e6e, 60a56f8, ec83496.

1f75880 added the host-located VMM allocation primitive (backing pages in
system RAM, mapped into GPU VA). Root-cause fix it hinged on:
HIP_MEM_LOCATION_TYPE_HOST was defined as 0, which ROCm reads as
Invalid/None; corrected to 2 (Device=1, Host=2 per driver_types.h), and the
C VMM probe confirms gfx1201 host-location granularity 4096 once correct.
New surface: HipMemAllocationProp::host_pinned(), HipMemLocation::host(),
mem_get_handle_properties() fail-closed placement check, MemoryLocality
enum + allocation_prop router + reserve_host(), Gpu::alloc_vmm_tensor_host().
Verified gfx1201: HOST_OFFLOAD PASS smoke, hip-bridge vmm unit tests 6/6.
NOTE: this VMM mechanism is superseded later in the stack by hipHostMalloc
(95c7b6a) — the primitive, not the enum fix, was reverted.

efa8e6e added memory.gpu_layer_budget (Full | auto/-1 | N resident) plus
largest_fitting_tail() pure admission fn (refuse-cleanly None when nothing
fits). Tests: 9 memory module tests; hipfire-config suite 83 passed / 0
failed. Fixture pin: qwen3.6-27b.mq4 size 14984158208 (sha256 86a5f80f...).

60a56f8 added resolve_i_gpu_start (Full->Some(0), Auto->tail, Layers(n)
->n_layers-min(n,n_layers)); pure, unit-testable, no behavior change.
ec83496 added Qwen35Config::i_gpu_start (layers [0..i_gpu_start) spill);
data-only, default 0 fully resident, reversible.
Squash of 88cc8ff, defd51a, 2ef2dfb, a545410, 0455781.

88cc8ff: upload_f32_host / upload_raw_host (+upload_raw_with_copy_host
with arena-release-on-copy-fail). Allocation half of offload; no callers
yet, behavior unchanged.

defd51a: HfqBackend.host_local + read_proj_host: Option<fn(..., &mut Gpu)>
twin seam (&Gpu device reader cannot host-allocate since arena registration
needs &mut Gpu). None = hard error on offloaded layer, never silent device
fallback. CPU dequant factored into dequantize_norm/dequantize_to_f32 so
both paths share the attractor-critical math. All other construction sites
set host_local:false, read_proj_host:None — neutral.

2ef2dfb: load_weight_tensor_host via shared load_weight_tensor_raw_with
(one copy of 36 upload sites + K%256/blob-length guards; qt 32/33 guards
kept). [m,k]-shaped arms (qt 1/2/16) host-localized too. f32-dequant
fallback quants REFUSED on offloaded layers (fail-closed). AWQ scale stays
device-side deliberately (1-D f16 vector of len K).

a545410 + 0455781: host_offload_smoke proves byte parity on real data
(17825792 bytes identical, gfx1201 qwen3.5-9b.mq4, q_proj qt=13
[8192,4096]) AND locality (vmm_host_located asserts both directions +
blob-size check), with a fixture picker (2-D raw-code tensor, dims>=1024)
instead of a hardcoded name.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…cope

Squash of 21ff5c9, 95b54a1, 89a8567. Docs only; no code change.

21ff5c9: four settled decisions — proj needs the read_proj_host twin seam;
f32-dequant fallback fail-closed for offload; byte-identity claims require
greedy decode (default-temp engine nondeterministic, measured on branch and
stock); target fixture qwen3.6-27b.mq4 (14984158208 B, sha256 86a5f80f...).
Latent MoE hazard noted (load_moe_ffn never sees HfqBackend).

95b54a1: hipfire bench pins greedy itself (temp 0.0/top_p 1.0), so the -t 0
rule is run-only; arms must interleave across fresh processes (consecutive
same-arm runs drift ~2x — measured branch 38.5 then 22.8 vs stock 23.0 then
40.4; interleaved medians 40.25 vs 40.35 tok/s, 0.25% apart). Resident VRAM
control ~13.7 GiB (86% of 16304 MB) via rocm-smi card0 CSV field.

89a8567: scopes -t 0 to run; bench output not for eyeballing (empty-think
template); VRAM control from bench JSON (13042 MB branch vs 13098 MB
stock, ~13070 MB) — offload proven only if vram_free_mb lands well above
~3100 MB.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ng, repro

Squash of b8b6520, 71e81c5, 63224ba, f51fd44, 7364371, ea61cc7.
Intermediate diagnosis on the VMM path; mechanism superseded by hipHostMalloc
(95c7b6a), probes retained as regression coverage.

b8b6520: hipMemMap needs granularity-aligned size; both alloc_vmm_tensor
paths passed caller bytes through. Real MQ3 blob (99840 B vs 4096 gran)
died "must be a non-zero multiple of granularity 4096". Fixed on BOTH
paths (device twin had the identical latent defect). Found by the first
real offloaded load, not smokes (aligned F32 only). SURVIVES in final tree.

71e81c5: host_kernel_access probe — real gemv_f32 over host VMM vs device
twin at the exact faulting 99840 B size. HOST_KERNEL_ACCESS PASS (max diff
0): platform is NOT the limitation.

63224ba: wired apply_offload_policy -> i_gpu_start -> per-layer host_local.
Default unset = fully resident, byte-identical to stock (verified vs
hipfire3 on qwen3.8-27b.mq3-xt). Auto(-1) failed closed (no capacity
measurement at config build). EP/MoE pinned resident. Offloaded path
FAULTED ("Page not present"); ruled out platform/graph/scale — remaining
suspect: quantized GEMV over host DType::Raw blobs.

f51fd44 + 7364371: two-arm repro (F32 PASS max diff 0; MQ4 FAULT on host
VA) with non-vacuity guards (device ref must be non-zero, exact match) and
layout asserts (row_bytes=(K/256)*136, K%256==0). Fixture bug fixed en
route (f32 scale/zero headers, not arbitrary bytes).

ea61cc7: root cause — forward kernels overread the tensor tail; exact-fit
VMM (every observed MQ3 blob an exact 4096 multiple) made the overread the
first unmapped byte. Fix: pad reservation + map whole, owner_buffer still
exact prefix. Verified: BUDGET=62 run -t 0 decodes, md5
8826a76be712ad4b9d74c96021f18bda byte-identical to resident. +OFFLOAD_DEBUG
VA logging. NOTE: ea61cc7's TODO.md blocker log was deleted by 998e55a
(correctly — resolved); f51fd44's "MQ4-specific" narrowing was retracted
in that log (device ref all-zero) and is not re-asserted here.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Squash of ef4aa23 and 998e55a. The load-bearing pair of the stack.

ef4aa23: offload was net-zero for VRAM — host-located VMM arena charged
the device heap 1:1. Measured gfx1201/ROCm 7.2 via 1 GiB hipMalloc ladder
with 4096 MB held: control 15360 MB, VMM arena 11264 MB (cost 4096 MB),
hipHostMalloc(mapped) 15360 MB (cost 0 MB). Production uploads rerouted to
new Gpu::alloc_host_mapped_tensor (tail pad now costs host RAM only).
Also fixed latent gemv_mq4g256 bug: 7 kernargs to a 5-param kernel (FWHT
tables belong to mq_rotate_x), m/k misread, all-zero result — no production
callers, confined to the diagnostic. Verified greedy -t 0
(code_edit_rewrite_copy.txt md5 80b910784c456bbb5532eb0493a497b9,
qwen3.8-27b.mq3-xt): byte identity md5 8826a76be712ad4b9d74c96021f18bda
(resident, BUDGET=62, BUDGET=32); footprint resident 13042 MB vs budget-32
7992 MB (-5050 MB for 4946 MB spilled; pre-fix arms read 13042 vs 13480).
headroom PASS (cost 0 MB), kernel_access ARM1_F32 + ARM2_MQ4 PASS
(bitwise-equal, non-vacuous, via production path). Decode at budget 32
PCIe-bound by design (4.4 vs 39 tok/s) — capacity feature, not speed.
Follow-up noted: pinned host RAM stalls under memory pressure (layer 38
stall at 0 MB free) need watching.

998e55a: deleted the dead host-VMM allocator (MemoryLocality/HostPinned,
reserve_host, allocation_prop router, host_pinned(), alloc_vmm_tensor_host;
vmm_host_located -> host_located ownership check) — branch-only code, trap
for re-adoption. Migrated smokes to production upload_f32_host (+ no-arena
+ owner-count asserts). Also fixed pre-existing vmm_tensor_smoke failure
(non-granular map rejection contradicted b8b6520's round-up since
ea61cc7). Verified gfx1201: build clean, all smokes PASS,
host_offload_smoke qwen3.6-27b.mq4 PARITY PASS (27852800 B identical).
…g sweep

Squash of b2057dd, 2e7b65e, feb4bec, 50d0822, 7611a45.

b2057dd: GPU tests in dispatch::tests serialized on a module lock
(try_gpu returns Gpu + guard). upload_raw_copy_failure asserted byte-exact
hipMemGetInfo free while siblings shared the device — failed ~1 in 3 runs
with no source change. No assertion weakened. Verified 6/6 parallel runs
+ four-package gate.

2e7b65e: SlotEngine gate charged the whole model file; now charges the
resident share summed from the HFQ index (layers non-uniform; non-layer
tensors always resident). Off == file size exactly (no-op for existing
configs). Pays off: max_seq=262144 un-loadable resident (OOM, 330 MB
free) loads at BUDGET=32 and emits resident-identical md5
8826a76be712ad4b9d74c96021f18bda. Coherence: battery+chain coherent,
attractor=0; unclosed-think turns fail identically on resident control
(pre-existing). Offloaded decode 4.5 vs 40.0 tok/s (8.9x PCIe penalty).

feb4bec: copy-fail assert 64 bytes exact -> 64 MiB signal / 8 MiB
tolerance (leak of this owner is 64 MiB; kills cross-process settling
flakes). Pool-counter equality unchanged. NOTE: superseded twice more in
stack (1c2fa50: 512 MiB/64 MiB on measured 28-56 MiB never-returned
driver overhead; 9bb5992: 500 ms poll for late-vs-lost reclaim) — see
commit 8 for final form.

50d0822: prose sweep host-VMM -> host-mapped system RAM (i_gpu_start doc,
host_located doc); 5 deliberate "rejected alternative" mentions kept.
Verified smokes PASS.

7611a45: release_registered_host_mapped removes only-if-released (was
unconditional drain -> invisible leak on hipHostFree failure);
ensure_vmm_cleaned refuses on live host_mapped owners too (+ naming test).
Changelog: weight sweep 1814 ms resident vs 57592 ms at =32; offloaded
Redline parity unverified (resident PASS exact; offloaded load >900 s
timeout, 1073.5 s sweep under harness vs 57.6 s plain).
…ding

Squash of 3018ddf, 12fa133, e4647e2, c718d8d.

3018ddf: gpu_layer_budget help stopped promising auto-fit (-1 fails closed;
largest_fitting_tail unwired to a device measurement). Changelog: offloaded
arm passes routed Redline at completable budget (=62, identical kernel
hashes); =32 load exceeds 900 s under harness.

12fa133: Offload row added to easy-mode lists (all four positional lists +
knobs::KNOBS explainer). Unset renders "all resident" (label never
round-trips into config). Explainer: counts RESIDENT layers (32/64 spills
32); sharp two-sided cost (per-step PCIe + slow load); -1 unimplemented.
Verified on rendered surface; tui tests 153 pass.

e4647e2: render-level guard (label+value pair + explainer title via
TestBackend). First draft asserting contains("Offload") was near-vacuous
(passes on "OffloadXX"); now whitespace-collapsed "Offload all resident",
sensitivity checked both ways. 154 pass.

c718d8d: row editable (EDITABLE_FIELDS spec; reasoning_effort same fix;
REASONING_EFFORTS exported from hipfire-config; -1..65536 bounds, null
clears) + wording un-inverted (Offload/all-resident -> GPU layers/all on
GPU; direction with concrete example; auto as unavailable not failure).
Three guards proven to fail without fix (class guard, typed-value
end-to-end, help coverage). Arrows-on-integers left as-is (tab-uniform).
Verified: tui 156 + config 88 pass.
… hardening

Squash of 03e963c, 7b5eae6, 1c2fa50, 9bb5992.

03e963c: auto(-1) no longer fails the load (unresolvable at config build:
needs measured capacity + per-layer bytes, no Gpu in hand) — keeps every
layer on GPU with a log note. NOTE: framed as "placeholder" here; 7b5eae6
reframes as engine-decides (below) — latter reading stands. No-op budgets
reported (Layers(200)/64-layer saturates silent -> now explicit). Render
guard made hermetic (both renderings pinned, override marker handled).
Placement arithmetic factored to pure offload_split + residency_report
(apply_offload_policy infallible; silence still means unset-default).
Tests pin direction/saturation/degenerate; verified 23 ok-blocks across 6
crates.

7b5eae6: auto = "the engine decides" like every other key (kv_cache,
dflash/mtp/flash/mmap_screen); currently decides resident until capacity
measurement feeds largest_fitting_tail. Behavior unchanged (Auto=>0);
vocabulary + docs + CHANGELOG only.

1c2fa50: feb4bec's guessed 64 MiB/8 MiB still failed ~2/5 (28 MiB
retained exactly). Measured via free_retention probe: owner bytes always
return exactly; driver intermittently charges 0/28 MiB (64 MiB req, up to
56 MiB at 2 GiB) hipFree never returns. Now 512 MiB signal / 64 MiB
tolerance (no observed overhead reaches it; leak drops 512 MiB). Proven:
skip-hip.free reddens with "536870912 bytes came back short"; 8+3 runs +
full gate pass with user serve live on same GPU.

9bb5992: poll get_vram_info up to 500 ms — probe showed reclaim LATE not
lost (11/12 exact, missing 28 MiB on later pass). Tolerance covers
never-returned overhead; poll covers lag. Proven: leak still reddens
("never came back ... after 500ms"); 6 runs + six-package gate pass.
…4 path

Squash of 1d40e51, 6e6a6aa, c83f64f, ea74e1d.

1d40e51 (S1): hipfire-cpu leaf — decode_group_codes (affine/codebook only,
what gemv dots against forward-rotated x) vs dequant_group (+inverse FWHT,
canonical basis); mixing them is silent R^-2 not a crash. 10 formats
(qt 13/44/15/8/17/20/1/2/16/3); qt 20 measured from registry qwen3.5:2b-mq3
(plan's "mq3=qt17" would miss every shipping -mq3 artifact). Evidence:
cpu lib 19 tests (canonical-generated tables, WH oracle); cross_check 1695
real tensors bit-identical; gpu_gemv_parity 8 formats worst 6.7e-7 vs 1e-4
(real qt13/20/15/8 rows + synthetic bit-exacts). HIPFIRE_OFFLOAD_DEBUG via
developer_var (fixes env-docs check).

6e6a6aa (S2 seam): memory.offload_exec=pcie(default, byte-identical
today)|cpu; fusion skipped on CPU windows; Step::Gemv/GemvResidual over
host-mapped weights run D2H->[AWQ divide]->rotate->hipfire_cpu::gemv->H2D
(residual accumulates in place); swiglu-down splits (GPU silu_mul_f32 +
CPU GEMV). host_bytes registry resolve; memcpy_dtoh_auto refuses under
capture; graphs disabled for model lifetime when spill possible. Two real
bugs fixed (invisible to parity): AWQ sidecar divide (138 fixtures) + double
rotation of Prerotated inputs. Contract: llama.cpp-level coherence +
measured divergence, not byte-identity (2B: 1343-char pcie md5
5722aca7..., cpu 652-char agreement then diverges; 9B likewise 723-char).
Boundaries: slots/serve GEMMs + prefill stay GPU-side; lm_head
always-resident; Q8HFQ/Givens/Paro refused. Tests: 21 cpu / 89 config /
272 dispatch / 251 rdna-compute.

c83f64f (S3): coverage table (7 shapes, 10 roles, 2B dim 2048; trace "0
host-mapped steps still on GPU", rotation tags real). Boundaries verified:
slots/serve + prefill uncovered; lm_head never host-mapped; cpu+spill+
Redline = load error (verified both ways); exec=cpu with no spill = one
info line, no-op.

ea74e1d (S4): AVX2+FMA Mq4G256 row dot (32 nibbles/16 B, scale·Σ+zero·Σx
per group; 1e-5 vs scalar, scalar pinned to S1 arithmetic) + per-step
D2H/GEMV/H2D trace + cpu_exec_time_ns. Measured gfx1201/7800X3D 9B mq4
(binaries+prompt md5s in orig): spill-8 cpu 24.1 vs pcie 21.6 tok/s
(+11.6%), spill-16 +27%; vram_free identical (capacity unchanged); ~23 vs
~17 GB/s effective. Ceiling 44.5 GB/s host-mapped (44.8 heap), limit is
per-step serialization not codegen. Routes: cpu battery+chain 5/5 coherent;
pcie redline PASS exact; cpu redline = intended refusal; no-gpu-ci green
(modulo pre-existing mq4c_repack failure).
Squash of 0ffb07d and 733f350. Docs only; no code change.

0ffb07d: dated historical record
(docs/perf-checkpoints/2026-09-27-gfx1201-cpu-exec-offload.md) — S4 numbers
previously lived only in a commit message and /tmp. Supersedes the two
3-6x-slower readings (contaminated: concurrent release build + second sweep
on same 16 threads). Clean interleaved: 24.1 vs 21.6 tok/s cpu at 8/32
spilled (+11.6%; +27% at 16/32), vram_free_mb identical. Beyond medians:
72 steps/token x ~0.53 ms ~= 38 ms vs 41.5 ms/token (no hidden
serialization; copies ~18%); kernel ceiling 44.5 GB/s host-mapped (44.8
heap), 84 GMAC/s x16 / ~11 x1 (Zen 3, no AVX-512) — cpu arm at ~80% of own
ceiling; determinism (two processes + x16 vs x1 byte-identical); divergence
= single whitespace token at ~190 tokens then drift (near-tie signature).

733f350: shared-prefix length = argmax-survival count under 6.7e-7
per-step diff (sensitivity, not accuracy) — stated in 6.2.1 so it is not
read as an error rate.
…nt record

Squash of 853a386, 74b0886, 924d7fd.

853a386: CpuQuant 10 -> 28 formats (every dense-capable qt-register
entry). qt 1/2/3/16, 6-12, 13/15/17/18/31, 19/20/30/51 Lloyd, 40/41
ternary/binary, 44/45/47/48/49/50 V2+MQ4C. Packings from canonical decoder
(bit-exact tables for 7/11/12/18/19/30/40/41) or mq-v2-family.md wire spec
(qt 47/48/49/50, 45 — transcribed from gemv_mq*g256v2 kernels, byte-agree).
Evidence: cpu lib 24 (layout table all 28, f64-oracle GEMV, 19 expectation
tables); cross_check 1695 real tensors bit-identical; gpu_gemv_parity worst
per format (Mq6G256V2 1.9e-7, Mq5G256V2 2.9e-7, Mq3G256V2 7.2e-8,
MQ4C/MQ2-fam/HFQ* bit-exact); qt 49 real-27B q_proj (12288x5120) 2.4e-7 /
3.5e-7 — the row making all-qt49 27B CPU-executable. Fixes: coverage line
now resolves via loader dtype_from_quant_type (was RAW_CODECS passthrough,
blind to arch-loaded qt 31/MFP4/PARO); qt 11/12 reported arch-ineligible
(no dense GEMV kernel on gfx12) not failed. Structural exclusions named
(qt 5 stride, 14 Mq8Internal, 21/24/32-37 HFP4/MFP4 layout, 28/29 PARO,
38/39/22 MoE/non-weight). E2E 2B spill-12: 12/12 covered, 0 on GPU.

74b0886: PERF BUG (shipped in prior two commits): V2 family + qt 44
widened per-128 fp16 [s0 z0 s1 z1] per element (branchy half->f32 =>
conversion-bound). Hoisted to v2_halves/half_of: 27B qt49 m=12288 k=5120
19.95 -> 5.70 ms/step (3.5x; run 286 s -> 163 s). Arithmetic order
unchanged (24/24 tables bit-exact, cross_check still 1695 green). qt 44 =
current --format mq4 output, so all mq4v2 bodies were affected (2B/9B qt 13
flat-f32-header hid it). Parity now covers BOTH rotation arms
(check_prerotated all formats: qt 45 bit-exact both arms despite
"HFQ4-G256" header comment). Record corrections: Bf16 from_quant_type(16)
row was missing (S1-S3 claimed coverage wrongly — fixtures carry none, so
their lines stayed true); 6.2.1 coverage bullet updated to 28-format set.
Remaining lever: only Mq4G256 vectorized; qt 49 scalar ~11 GMAC/s vs qt 13
30-60 GB/s — AVX2-for-V2 next, not correctness.

924d7fd: dated amendment for qwen3.8-27b.mq3-xt (497 projections all qt
49): budget-56 (8 spilled) loads 8/8 covered; budget-48 (16 spilled, ~2.5
GB pinned) stalled 2x at layer 62/64 with 0 GB free — pinned RAM is a
load-time capacity knob. qt 49 post-hoist real-tensor rows re-measured
unchanged (49 7.19e-8; 44/47/48 + 45/50 rows). STATED GAP: no eyeballed
completion on this fixture (160-token run hit reasoning-budget gate; 2
loads stalled; 600-token run same gate) — qt 49 rests on parity +
coverage/trace, not decoded text.
Squash of ab090d7, 7133b3c, 8b2c3f0.

ab090d7: self-contained benchmark handoff runbook (fixture digests for
2B/9B/27B + prompts + binaries; knobs + split assertion; interleaved
fresh-process loop; five contamination modes that each cost a sweep leg;
run matrix; output-reading guide; sanity anchors + expected crossover;
27B footprint recipe; report requirements). Indexed in docs/INDEX.md;
HIPFIRE_CPU_EXEC_TRACE added to env inventory (env-docs green).

7133b3c: identity table gains git rev-parse HEAD beside binary md5s;
contamination table (symptom->cause->check: concurrent build ~300% on cpu
arm, surviving daemon, reused-pid kill, steps:0 = trace unset, uncovered
quant, prefill delta impossible, vram delta = pass-condition failure,
attractor-as-tight-stddev); knobs frozen between arms (leave
HIPFIRE_VERIFY_GRAPH alone).

8b2c3f0: simd::mq3g256v2_row_dot_avx2 (AVX2+FMA+F16C: one header widen,
3-byte chunks via broadcast+vpsrlvd+mask+convert+FMA, 8 contiguous 3-bit
codes per 24-bit LE field, no f32 materialization; affine header = 2 FMAs
per half). row_dot_enabled() per-format gating (qt 13 AVX2+FMA, qt 49
+F16C; AVX-512 needs own gate — Zen 2 k9lin in host set). Trace now
per-shape means (was process-wide mean misattributed per shape: "m=48
5.31 ms" was first-3-steps mean; real m=48 = 0.048 ms scalar / 0.024 AVX2).
Measured 27B qt49 BUDGET=56 (8/64 spilled), round-robin {cpu,pcie} x
{pre,post} x 3 fresh processes: cpu 3.10 -> 15.20 tok/s (+390%), pcie
control 13.20 both binaries (delta is the arm; -76.5% loss -> +15.2% win).
Steady state 0.05-0.96 ms GEMV vs 3.4-9.6 scalar; 7.48 s GEMV + 1.42 s
copies over 12800/15872 steps; non-uniform mix (16:8:8:8:4:4:1:1) so
per-token quoted as ranges. Evidence: 27b amendment + raw data/ kept;
crate tests 25/25 dbg+rel (f64 oracle + AVX2-vs-scalar with non-pow2
headers); qt 49 re-parity 5.364e-7/9.537e-7 vs 1e-4, others unchanged;
greedy run cpu==pcie 610 B. Docs: env inventory, runbook, design note,
CHANGELOG. Raw bench data (bench JSON, traces, drivers, kernel probe)
kept byte-identical per tree-parity requirement.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The coverage map (docs/quant-formats/cpu-simd-coverage.md) lists the V2
family as scalar-only, ranked "warpfront#1 highest" for reuse x shipment: qt 44
(Mq4G256V2) is what the mq4v2 bodies and qwen3.8-27b.mq4-xt actually
carry, and qt 47/48/50 are the same header over different payload
widths. All four share qt 49's fp16 per-128 header, so they were one
kernel each away.

x86.rs becomes a table of layouts over one decode core instead of a
kernel per format:

- codes8<BITS> — eight BITS-wide codes as i32 lanes, in element order.
  An 8-code chunk is exactly BITS bytes, so nibbles, the 2/3/5/6-bit
  cross-byte packs and (next) BQ1's sign bits are all one broadcast plus
  one variable shift: <=4-bit in 32-bit lanes (one vpsrlvd for all eight),
  5/6-bit in 64-bit lanes with the four low dwords compacted per half.
- codes_dot_sum — (sum c*x, sum x) over a run of chunks, so a group pays
  two horizontal sums, not one per chunk.
- f16_quad — the V2 header widened with one vcvtph2ps; v2_group_dot
  scores each 128-element half and then affines it.
- row_dot! — the row driver (sum over groups), monomorphized per format
  because a #[target_feature] body's features do not propagate into a
  closure or through a function pointer.

load_le reads *exactly* the chunk's bytes (the 3-bit width reuses qt 49's
one-preceding-byte 32-bit load, which stays in bounds because a payload
always starts after its header). A wider load is one instruction cheaper
and runs up to three bytes past the group on the last chunk, which the
row slice licenses only for rows that are not the tensor's last.

qt 49 moves onto the shared path; its arithmetic is unchanged (same
extraction, same per-half dot/sum, same combining order), so its recorded
numbers still hold. qt 13 keeps its hand-unrolled 32-codes-per-16-byte
nibble group — see the module docs for why.

Dispatch: row_dot_enabled(q, req) is now "the features q's kernel needs
are present" (features_for: F16C for fp16-header kernels, AVX2+FMA
otherwise), and row_dot_avx2 is the single format -> kernel map, returning
None for formats not ported yet so gemv's per-row match disappears.

Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C):
- crate tests 26/26 green, including gemv_matches_the_f64_reference, which
  now runs the 5 V2 formats through the SIMD path against an independent
  f64 oracle (AVX2 is active by default on this host).
- avx2_and_scalar_agree_within_tolerance now asserts the dispatcher hands
  the row to a kernel at all: a None would leave gemv on the scalar path,
  where the two outputs agree by construction and the comparison would
  pass vacuously. Headers are rewritten non-power-of-two so a mixed-up
  half or a swapped scale/zero is O(1) relative.
- kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best
  of 3, not a serve-path measurement): qt44 4468 -> 281 us (15.9x), qt47
  7025 -> 487 (14.4x), qt48 8863 -> 439 (20.2x), qt50 4132 -> 310 (13.3x),
  qt49 2520 -> 327 (7.7x), qt13 unchanged 3965 -> 297 (13.4x).

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Tier 1 (qt 15 Mq6G256, qt 31 Mq5G256), Tier 3 (qt 45 Mq4CG256) and the
rest of the f32/fp16 single-header family, all of which the coverage map
lists as scalar-only.

Two observations collapse most of that list into rows of a table:

- The kernel only sees bytes, and rotation is applied to the *activation*
  by the caller (crate::gemv / quant::is_fwht_g256). So an unrotated
  format and its rotated twin share a kernel: qt 6 has qt 13's geometry,
  qt 8 -> 15, qt 11 -> 17, qt 9 -> 18, and the G128 blocking formats
  (qt 7/10/12) are the same shape at 16 chunks per group.
- TQ2 (qt 40: (code-1)*d) and BQ1 (qt 41: bit ? +d : -d) are affine once
  rewritten as scale*d + zero — (d, -d) and (2d, -d) respectively — so
  they are one uniform_group_dot each with a derived header, not two
  kernels.

x86.rs gains uniform_group_dot::<BITS, PAYLOAD, CHUNKS> (the group body
the V2 family already used for one half) and an affine_row_dot! macro
that instantiates the header reader plus the row driver together, which
is what keeps 15 formats to one macro call each.

qt 41 is the first user of codes8::<1>: one byte of sign bits is eight
one-bit codes, i.e. the same broadcast-and-shift as every other width.

Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C):
- crate tests 26/26 green; gemv_matches_the_f64_reference covers all 15
  new formats through the SIMD path against an independent f64 oracle.
- avx2_and_scalar_agree_within_tolerance extended to every new format,
  with awkward_headers rewriting the header per family (f32 pair, fp16
  pair, fp16 d) so a swapped scale/zero is O(1) relative.
- kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best
  of 3, not a serve-path measurement), speedup vs the scalar decode:
  qt6 13.6x, qt7 13.9x, qt8 6.6x, qt9 14.0x, qt10 8.8x, qt11 7.7x,
  qt12 6.9x, qt15 6.6x, qt17 7.2x, qt18 12.8x, qt31 19.8x, qt40 10.7x,
  qt41 10.2x, qt45 13.0x. No format regresses.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Tier 2 of the coverage map: qt 20 (Mq3G256Lloyd, what the registry's
-mq3 tags actually carry), qt 19 (Mq2G256Lloyd) and qt 30
(Mq4G256Lloyd), plus qt 51 — byte-identical to qt 19 and unrotated by
design, which is a caller-side difference, not a kernel one, so it is a
second row_dot! over the same body.

The Lloyd tier has no affine header: the group's CB fp16 entries *are*
the decode and the payload indexes them, so codebook_group_dot needs one
accumulator instead of two and no scale/zero.

The lookup is vpermd, not vgatherdps: a 4- or 8-entry book is one
permute, and a 16-entry book is two plus a blend on the index's top bit —
which works precisely because vpermd reads only an index's low three
bits, so both halves' lookups are valid and the blend selects. For a
table this small the gather is the slower instruction.

Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C):
- crate tests 26/26 green; gemv_matches_the_f64_reference covers all four
  formats through the SIMD path against an independent f64 oracle.
- The tolerance test's fixture now rewrites the *codebook* for this tier,
  not a header: the shipped fixture books are powers of two and qt 30's
  repeats an eight-entry book into sixteen lanes, which would hide an
  off-by-eight table error. write_cb installs entries with alternating
  signs and a distinct magnitude per lane across four exponent classes.
- Mutation-checked that the new coverage bites: dropping the 16-entry
  blend fails qt 30 at rel 7.08e-1, and moving qt 20's payload offset by
  one chunk width fails at rel 3.50e+1. Both revert clean.
- kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best
  of 3, not a serve-path measurement): qt19 21.6x, qt20 8.6x, qt30 16.5x,
  qt51 21.1x over the scalar decode.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The last four formats without a kernel (qt 1 F16, qt 2 F32, qt 16 Bf16,
qt 3 Q8F16) close the table, so row_dot_avx2 no longer has a fallback arm
and row_dot_enabled is now purely a CPU-feature question. Both matches
are exhaustive: adding a CpuQuant variant fails to compile until its
kernel and its gate are stated, which is what keeps the two tables from
drifting apart.

F16/Bf16/F32 are element formats — no codes, no header, just weights to
widen — so dense_row_dot! runs them: 256-element groups, eight elements
per lane group, and the only per-format variable is the element width.
Bf16 is the one kernel in the crate that needs no F16C (a 16-bit shift is
the whole conversion), and F32 needs no conversion at all.

qt 3 Q8F16 is the last hand-written group: a 32-element block of one fp16
scale then 32 signed bytes, decoded with vpmovsxbd.

docs/quant-formats/cpu-simd-coverage.md arrives with this change (it was
the plan this series executed, and lived only in a sibling checkout): the
per-format kernel table with its ISA gate, how the kernels are organised,
what is still not covered (no CpuQuant variant, non-x86_64, no NEON), and
the verification story. Indexed in docs/INDEX.md; the CHANGELOG entry
covers the whole series (and the qt-49 entry's now-renamed kernel symbol
is corrected there).

Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C; GPU gfx1201 / RX 9070 XT):
- crate tests 26/26 green. every_cpu_decodable_format_has_a_kernel now
  takes its format list from from_quant_type over the whole u8 space, so
  it fails if the kernel table and the qt map ever disagree.
- gpu_gemv_parity (pre-existing, #[ignore]d, real fixtures + GPU): the
  production GPU launcher against this crate's GEMV with the vector path
  active, driving the real cpu_exec seam with a forward-rotated
  activation. Real tensors of qwen3.5-2b.mq4/mq6/mq3 and qwen3.5-9b.mq4 —
  qt 13/15/20, with AWQ and prerotated arms (18 checks) — plus synthetic
  buffers for the rest: 23 formats in its per-format summary. Worst
  relative error 9.3e-7 (Mq4G256, the +awq real-tensor arm at m=6144
  k=2048); then Mq5G256V2 2.9e-7, Hfq6G256 2.3e-7, Mq6G256 2.3e-7,
  Mq3G256Lloyd 2.2e-7, Mq6G256V2 1.9e-7; then Mq4G256V2 8.9e-8 and
  Mq3G256V2 7.2e-8; the remaining 15 of the 23 report exactly 0.
  Two honest gaps: qt 11/12 report SKIP (no production dense GEMV on
  gfx1201, so the comparison could not run for them), and F16/F32/Bf16
  are absent from the test's synthesised table, so those three rest on the
  crate tests alone.
- The tolerance test's fixture rewrites the *elements* (and, for Q8F16,
  the block scale) for these formats: the shipped fixture values are small
  powers of two, which would make every product exact. With them awkward,
  the measured worst case across every format and shape is 2.3e-5
  relative, so the bound there is now 1e-4 — documented as f32
  accumulation noise rather than a decode bound, against the 7.1e-1 /
  3.5e+1 a real decode error produces.
- Mutation-checked: dropping the bf16 shift fails at rel 1.0e0, moving
  qt 3's payload offset fails at rel 2.0e+1. Both revert clean.
- kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best
  of 3, not a serve-path measurement): qt1 20.2x, qt2 13.2x, qt16 15.4x,
  qt3 7.9x. Nothing in the table regresses; the scalar decode stays the
  ARM/non-AVX2 path and the test reference. No serve/bench run: per
  AGENTS.md the parity route is the claim-scoped gate for a kernel change,
  and this makes no model-level perf claim.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Two defects in the Settings surface the layer-offload work added
(`feat(tui): offload row in Settings easy list`), both on the
`GPU layers` row (`memory.gpu_layer_budget`) and both invisible to its
guards — those assert that a row has a spec, an explainer and draws, not
what a value means.

1. Clearing the row stored the literal string "null". `writer::write_value`
   returns `Value::Null` for the documented clear spelling (`null`), and
   `persist_setting` reflected it as `Value::Null.to_string()` — the string
   "null". The row's label then read "null on GPU" (its unset test is
   `is_empty()`), `easy_override_state` claimed a change for a key that
   was not in config.toml, and the next edit seeded its buffer from
   "null", so typing a count produced "null32" — rejected by the schema,
   which made a cleared row unsettable without hand-erasing the buffer.
   A cleared value now stores the unset spelling and drops the override
   marker (which has always meant "this key is explicitly on disk").

2. Enter on a row the user never typed into was an error. The editor
   opens empty exactly when the stored value is null
   (`current_setting_value` returns ""), and `parse_cli("")` on a
   nullable integer is `could not parse ''`, so Enter answered
   `rejected invalid value for gpu_layer_budget: ""` and stayed in edit
   mode. An empty buffer on a field that cannot hold the empty string IS
   the unset spelling, so `write_value` now retries it as `null` — which
   clears, and clearing an already-unset key is a no-op. A `FreeStr`
   field accepts "" directly and never reaches that arm, so
   `prefill_drafter` still means the empty string.

Also, the same row rendered its reserved auto spelling as "-1 on GPU",
claiming minus one layers were resident; it now reads
"auto (engine decides)", matching the schema's own wording ("defers
placement to the engine").

Verification (host: Zen 4 7800X3D; no GPU needed):
- tui 160 tests + config 89 tests pass.
- Each fix is mutation-checked, i.e. the new tests fail with the fix
  reverted: restoring `Value::Null.to_string()` fails
  `clearing_the_offload_row_returns_it_to_unset` with
  `left: Some("null") right: Some("")`; removing the empty-means-clear
  retry fails
  `empty_input_clears_a_nullable_field_but_stays_empty_for_a_string` with
  `Invalid { key: "gpu_layer_budget", value: "" }`.
- The clear path is driven through the real key handler, not the writer
  alone: Enter → type 32 → Enter (renders "32 on GPU"), then erase the
  seeded buffer → Enter (renders "all on GPU", no override marker, key
  gone from config.toml), then `null` → Enter (same state), then type 8 →
  Enter (renders "8 on GPU") — the last step is the one that used to
  produce "null8".
- The changelog line here is the clear-semantics fix; the row's own entry
  (the `Offload exec` setting) lands with the row's change.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
… exec`

`memory.offload_exec` (compat env `HIPFIRE_OFFLOAD_EXEC`, default `pcie`)
decides whether a spilled layer's weight-reading GEMVs run on the GPU over
the link or on the CPU. It was reachable from the config file and the env
var and from nothing in the TUI: the key existed, the engine read it, and
no Settings row offered it — the same gap the layer-offload row closed for
placement. This adds the second half of that decision next to the first.

- `Offload exec` easy row, right after `GPU layers`, in all four positional
  lists (`easy_keys`, `easy_override_state`, `easy_help_keys`, `easy_rows`)
  with a `knobs::KNOBS` explainer and per-option help for both arms. Cells
  read `over PCIe` / `on CPU`, in the same grammar as their sibling
  (`all on GPU` / `32 on GPU`).
- An enum row, so Left/Right/Space stage a preview and Enter commits — the
  path the existing `EDITABLE_FIELDS` spec mechanism already handles; no
  numeric buffer, so none of the empty-buffer semantics apply.
- Its option list is `hipfire_config::OFFLOAD_EXECS`, and that same const
  is now what the schema field validates against, so the two cannot drift.
  A local copy would have been the class that left `reasoning_effort`
  uneditable and still leaves the TUI's `thinking_budget` list without
  `off` (a value the schema accepts, so cycling from it jumps to the wrong
  option). Exporting the list for a two-value enum is the precedent set by
  `REASONING_EFFORTS` in the same file.

Guards (each exercised, not just asserted):
- `every_inline_editable_easy_row_has_a_field_spec` in the config crate
  now covers the new row by construction (it walks `easy_keys`), and
  `offload_exec_options_are_the_schema_values` pins the option list, writes
  every arm through the shared validator, and rejects a value outside it.
- `settings_easy_draws_the_offload_exec_row_and_its_explainer` renders the
  real Settings surface (ratatui TestBackend, both arms plus the explainer)
  with the values pinned in-test, because `render_with` builds on
  `App::load()`, which reads the developer's real `~/.hipfire/config.toml`.
- `easy_offload_exec_row_cycles_and_commits_every_schema_value` drives the
  real key handler end to end: cycle stages a preview without writing,
  Enter commits, the row re-renders, the key lands in config.toml, and
  cycling wraps.
- `offload_exec_row_renders_both_arms` pins both cell values.

Verification (host: Zen 4 7800X3D; no GPU needed):
- tui 163 tests + config 89 tests pass; `cargo run -p hipfire-tui` starts
  on a pty and runs without panicking.
- Not verified here: a real terminal session driving the row (the render
  tests are the crate's own surface harness, and the pty run was
  start-only); and the engine-side effect of flipping to `cpu` (that is
  the offload path's own measurement, docs/perf-checkpoints/2026-09-27-*).

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
… exec spellings

Follow-up corrections to the `Offload exec` row and to the clear semantics
around it — the sentences a reviewer needs stated exactly rather than
approximately:

- `writer.rs`: the `gpu_layer_budget` spec comment claimed "typing `null`
  clears the override". The editor cannot do that: its buffer seeds from the
  current value, so typing appends ("32" + "null" = "32null", which the
  schema rejects) — the end-to-end test in the previous commit had to be
  rewritten for exactly this reason. The comment now says what the code
  does: an *emptied* buffer is the unset spelling (`write_value` maps empty
  input on a nullable field to clear), and `null` is the config/CLI spelling
  rather than something the buffer can reach once a value exists.
- `knobs.rs`: the exec explainer now names both spellings the user sees —
  "over PCIe" is `memory.offload_exec = pcie`, "on CPU" is `= cpu` — so the
  row cannot reintroduce the un-inverted-wording problem the GPU-layers row
  was fixed for; and it carries the timing ("read once at load — changing it
  takes effect on the next serve/restart"), because the key is process-scoped
  and snapshotted at startup. Without it the row would promise an immediate
  effect it cannot deliver.
- `writer.rs` test: `empty_input_clears_a_nullable_field_but_stays_empty_for_a_string`
  now also pins that the empty-means-clear retry cannot hand a *non-nullable*
  enum a clear spelling its schema does not have:
  `write_value(path, "offload_exec", "")` must still error.

Scope note, since it is easy to misread the previous commit: the
empty-means-clear retry lives in `hipfire-tui`'s own writer, which has a
single non-test caller (`app.rs` `persist_setting`). The CLI does not go
through it — `hipfire config set` uses `hipfire_config`'s
`ConfigLayer::set_cli` → `field.parse_cli` — so
`hipfire config set <nullable key> ""` still answers `could not parse ''`.
This commit changes no CLI behaviour, and neither does the previous one.

Verification: tui 164 + config 89 pass; the pinned assertion above is
asserted against the shared validator, not assumed.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The knob that drives the whole feature had no env-var reference row: docs/env-vars.md gained HIPFIRE_OFFLOAD_EXEC and HIPFIRE_CPU_EXEC_TRACE but not the budget, and the AGENTS.md §7 table gained only the exec row. Adds both (typed-key table + compat-env table + the §7 quick-reference row). The generated-drift check passes because the field!() macro already declared the alias; the human-facing tables were the gap.

Also asserts the plan's shipped status and replaces §11's unchecked acceptance list with the evidence of record: the byte-identity and zero-diff criteria are ticked against the fixtures that evidenced them (2B/9B), the 27B tier is marked not-evidenced on MQ4XT (the run of record is qwen3.8-27b.mq3-xt), and the A3B criterion is marked half-evidenced (structural exclusion, no sparse trace compared).
…in the reader seam

Three doc comments on the projection reader still said the host path 'registers a VMM arena'. That mechanism was deleted in favour of hipHostMalloc(hipHostMallocMapped), which registers a host pointer the registry frees with hipHostFree. Comment-only.
…pins

The mq3-xt amendment cites benchmarks/prompts/gpu_offload_probe.txt (md5 5835c71e471849b4a72e1dc8e39695e7, 59 tokens) as the measured fixture; the 27B numbers are not reproducible without it, and AGENTS.md rule 4 requires canonical bench prompts to live here rather than at a scratch path. Byte-identical to the pinned md5.
…rds and raw captures

The records and their raw captures named this machine's layout: the signed-in user in /home/<user>/.hipfire_kernels and a dev-checkout path, and a host mount for the 27B artifact. Normalized to ~/.hipfire_kernels, <repo-root> and <models-dir>/qwen3.8-27b.mq3-xt. No number, digest, size or line count changed; 25 files, 38 lines. The new placeholders are angle-bracket substitutions, so the shell drivers in the data dir stay obviously non-executable as-is.
Four records were referenced but not present, so the links from the committed records, the benchmark handoff and three Rust doc comments did not resolve: the llama.cpp per-spilled-layer baseline (the source of the 27.1 GB/s link rate), its hipfire counterpart, the independent CPU-exec reproduction, and the 2B/9B/27B spill sweep the mq3-xt amendment corrects. Added as historical records, paths normalized like the rest of the family. The dated-reference graph now closes.
The repo's changed-file rustfmt gate flags 27 of this branch's own files, and three of the hunks are new lines rather than historical debt: an ungrouped import and two unwrapped calls in the qwen35 load path, plus the offload module's indentation in hipfire-config. Mechanical, produced with scripts/fmt-changed.sh against the ref the branch was cut from; `cargo check --workspace --examples` still passes and rustfmt --check --config skip_children=true is clean on all 27. The CI job is advisory, but this is what the tracked pre-commit hook would demand of any later commit touching these files.
scripts/verify-bind-thread.sh reports 12 of 94 pub fn in impl Gpu missing the bind_thread contract on this branch, against 8 of 86 on the ref it was cut from: the four new host-mapped entry points were the delta. Resolved per function rather than blanket-annotated.

- upload_raw_host / upload_f32_host issue device work (hipHostMalloc through alloc_host_mapped_tensor, then a raw memcpy_htod), so both now bind the calling thread first.
- host_located / vmm_handle are pure state reads — the tensor's own ownership tag and the arena registry — and take the gate's documented skip marker with that reason.

Remaining: 8 of 94, byte-identical to the pre-existing set on the base revision (last_launched_kernel, deadline_exceeded, prepare_mq4v2_fp8_x, upload_f32, upload_f16_bits, upload_raw, htod_uploads, pool_stats). Untouched: pre-existing debt, and the script is a pre-commit hook rather than part of CI's required gates job. cargo test -p rdna-compute --lib: 251 passed.
…ed placeholders

Three follow-ups to the path sanitization:

1. The A/B drivers were left non-runnable by it: '<repo-root>' and '<models-dir>' are redirection operators in shell, not placeholders. They now derive REPO_ROOT from git and take MODEL from the environment (MODEL=... ${MODEL:?}). 'bash -n' is clean on both.
2. Two statements denied the files' existence: the drivers' own headers said 'not committed', and the base record's reproduction notes said the arm driver and microbenchmark 'were scratch and are not committed'. Both are committed in this record's data directory; the text now says so.
3. '<models-dir>' -> '/path/to/models' across the records and their JSON/err captures: the angle-bracket form is a valid redaction but is also repo house style for unset variables in fenced commands, and the path form reads unambiguously as a placeholder in both prose and payload strings. No digest, size, count or number changed.
…wrappers

`clippy::not_unsafe_ptr_arg_deref` is deny-by-default, so `cargo clippy -p hip-bridge`
failed to compile the crate (3 errors, exit 101) and took the whole workspace run down
with it — hipfire-dispatch, hipfire-arch-qwen35, hipfire-runtime and rdna-compute only
ever reported that cascade, which is why four crates appeared to have errors of their
own. The same command on the pre-change tree is clean (5 warnings, 0 errors), so all
three are new: `host_get_device_pointer`, `host_free`, `mem_get_handle_properties`.

The file already carries this allow, with the same rationale ("HIP treats `device_ptr`
as an opaque GPU address; Rust never dereferences it"), on `mem_get_address_range`, so
these mirror the established pattern rather than turning three public signatures into
`unsafe fn` and rippling that to every caller. Regenerates hip-bridge's map for the
seven added lines.

@fivetide fivetide left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tackling this. Capacity offload is something we want, and the measurement discipline here is excellent. The mechanism layer is in good shape: the host-mapped primitive with an ownership registry and cleanup check in rdna-compute, the DeviceBuffer::is_host_mapped tag, the HfqBackend reader seam, and the CPU seam in execute_steps are all in the right places.

Our main concern is architectural. Hipfire's direction is that arch crates should contain as little code as possible, ideally typed declarations, with loading, placement and dispatch handled centrally. In this PR, the decision about which layers spill (and everything that follows from it) lives in hipfire-arch-qwen35. As a result, every new arch would have to copy it, and even qwen35's own PARO and MoE paths already fall through the gaps. Concrete suggestions:

  1. Placement belongs in the loader, not in Qwen35Config. model_load::Layout is already the arch-generic "where does each layer land" object. Give it a per-layer Residency { Device, HostMapped }. Resolve it once from memory.gpu_layer_budget in the shared load path, and pass it through WeightSource::read_layer(gpu, layer, residency) into HfqBackend. llama and qwen2 already build HfqBackend, so they would get offload almost for free. For the manifest route (weight_manifest::WeightPlacement / fulfill_manifest, currently piloted on llama), the same tier becomes a field on the placement. A contiguous spilled prefix is really just a pipeline split whose first stage is the host tier (DeviceMesh::stage_for_layer already bands layers).
  2. One upload path, not a qwen35 twin. Rather than parameterizing qwen35's private quant-type match (whose doc comment already says it should be pulled into runtime), consolidate it into weight_backend::decode_raw_codec / RAW_CODECS and give that function a residency/target parameter. Change the read_proj signature to take &mut Gpu. The "&Gpu-only fn-pointer contract" is our own choice and can change, which removes the need for read_proj_host as a second entry point.
  3. Downstream decisions should read residency from loaded weights, not config. Dispatch already asks gpu.host_located(buf), which is the right shape. Record the execution target (GPU over PCIe vs CPU) on the weight at load time. Then graph eligibility, the Redline refusal, admission, and device-side derived caches (e.g. ensure_fp16_shadow) can all ask the Gpu/weights "do you own host-executed weights?" instead of config.i_gpu_start plus a process-global LazyLock.
  4. Gate capture centrally. Put the refusal in GraphState::begin_graph_capture* and the replay controller's begin_capture, instead of two qwen35 call sites (forward.rs:1694, forward_slots.rs:2518). MTP proposal, DFlash draft, dense-TP and verify graphs aren't covered today, and mtp_spec.rs / dflash.rs call hip.stream_begin_capture directly, which bypasses the capture_mode guard in memcpy_dtoh_auto. The harness runs used --speculation off, so spec decode + offload_exec=cpu is currently unvalidated. The Redline refusal should live in hipfire-loader, where the replay backend is chosen, not in qwen35 arch.rs.
  5. Admission from what the loader actually allocated. Have the loader report device bytes and pinned-host bytes. Feed ModelFootprint from that, and charge pinned host memory to the existing host tier in AdmissionController (unreclaimable hipHostMalloc is exactly what that tier guards against — the PR notes 16/64 layers stalling with 0 GB free). That removes the tensor-name parsing (tensor_layer_index) in serve_engine.
  6. Fail closed on sources that can't spill yet. i_gpu_start is set from config regardless of source, but ParoSource::read_layer never sees it and load_moe_ffn places experts in VRAM unconditionally. Those loads log "N offloaded" while resident_weight_bytes subtracts the full layer bytes (experts included) from the admission charge, so a 16 GB card can be over-admitted into OOM. Until residency flows through PARO, MoE experts and VL, gpu_layer_budget should fail the load for them. Same for UMA devices (e.g. Strix Halo), where spilling to host frees nothing.
  7. One CPU decoder. hipfire-cpu is at layer 3 and hipfire-runtime at layer 5, so runtime can depend on it. Make hipfire-cpu the canonical CPU dequant and have weight_backend::dequantize_to_f32 call it; the cross-check test and the transcription then go away. Ideally the CPU decoder becomes part of the codec registry entry, so a new format is one row and cpu_quant_for disappears.
  8. Use Steps for the fused FFN. Instead of hand-splitting weight_gemv_swiglu_residual in llama.rs, express SiLU + GemvResidual as Steps so the generic seam handles it. That's also where prefill/batched coverage would eventually plug in.
  9. Trim scope so this can be reviewed. Please split out: the gemv_mq4g256 kernarg fix (a real bug, its own PR), the rdna-compute test lock and tolerance changes, the TUI rows (~640 LOC), and the ~30 perf-checkpoint data files. largest_fitting_tail / resolve_i_gpu_start and GpuLayerBudget::Auto aren't used in production yet: drop them, or land them in runtime next to admission.rs when auto actually decides something (placement arithmetic doesn't belong in hipfire-config, which is config vocabulary only).

Suggested order: (2) → (1) → (3/4/5) → CPU execution. After step (1) the RAM spill works for every arch that uses HfqBackend, and your qwen35 numbers should reproduce unchanged.

GpuLayerBudget::Auto, largest_fitting_tail and resolve_i_gpu_start had no
production caller: the qwen35 policy matched the enum but never called the
arithmetic (it carried its own saturating_sub split). 'auto' and '-1' stay
accepted and now mean fully resident, which is what the engine did with them
anyway. The qwen35 Auto arms become unreachable and go with the variant.
Avery Drouillard added 16 commits September 30, 2026 09:47
HfqBackend carried two projection readers (device + host) behind a host_local
flag, and qwen35 carried its own copy of the reader with a private quant-type
match that had drifted six formats ahead of the runtime's codec registry
(qt 31/32/33/34/36/37 — MQ5G256 and the MFP4/MFP3/MFP2 E8+Lloyd family).

Now: RAW_CODECS covers those six, DType::requires_k_mod_256 carries their K
constraint, expected_payload_bytes replaces six hand-copied length checks, and
one reader (hfq::load_weight_tensor + decode_weight_bytes) takes a Residency
instead of a second entry point. The host-decode fallback (qt 1/2/16) got a
host upload, so an offloaded layer can no longer silently land on the device.

qwen2's own reader is a strict subset of the shared one and went with it. The
kernel.lm_head_f16 policy moved to hipfire-config (its only consumer, qwen35's
private match, is gone) and now applies at the LM-head read only: a qt-1
projection is storage, not policy, and llama's fixture — every projection qt 1 —
is what proved it.
qwen35 derived its own i_gpu_start from memory.gpu_layer_budget at config
construction, so only qwen35 could spill and the decision had no relationship to
what the load actually did. Placement is now one decision in one place:

  Layout::spill_count(n_layers, requested) -> spilled, applied by
  Layout::resolve_residency, read back per layer by WeightSource::read_layer.

Every arch holding an HfqBackend therefore inherits a spill (llama, qwen35, and
qwen2's backend), the 'partial offload: N resident / M offloaded' line is emitted
by the loader for all of them with the same wording, and qwen35's
Qwen35Config::i_gpu_start, apply_offload_policy, offload_split and
residency_report are gone.

memory.gpu_layer_budget is now Option<usize> (the requested resident count) with
the enum deleted; unset/empty/auto/-1 all resolve to 'no budget'. The TUI's -1
row says 'all on GPU' rather than 'auto (engine decides)', which is what -1 has
always meant.

Fail closed at the same seam: a source that cannot honour a spill refuses the
load before allocating (PaRoQuant, MoE preserved-experts) and so does a
unified-memory device, where spilling frees nothing.

Verified: 9B resident output byte-identical to master (greedy) and a 24/32 pcie
spill byte-identical to resident.
…e refuses a spill

qwen2 drove its own embed -> norm -> lm_head -> layer loop, so it was the one
HFQ arch that could not spill and the one with no whole-model rollback from the
shared transaction. It now has a WeightSource and goes through
model_load::load_weights like llama and qwen35, with the per-layer free the
transaction's rollback needs; its progress lines and load order are unchanged.

The manifest pilot's executor has exactly one upload (pooled device memory), so
WeightPlacement gains the residency tier the review asked for, plan_manifest
populates it from the shared split arithmetic, and fulfill_manifest_single
REFUSES an entry the tier calls host-mapped rather than allocating it in the VRAM
the spill exists to free. The legacy load_weights_hfq route is what spills llama
today; this pilot refuses until it grows a host upload.
… what the loader allocated

Three decisions that used to be re-derived per site are now recorded once:

* ExecTarget on WeightTensor/WeightRef. The loader sets it from the resolved
  residency plus memory.offload_exec; the dispatch CPU seam tests the recorded
  decision instead of re-deriving it from where the bytes happen to live. That
  distinction is real: host-mapped + pcie is a GPU-read-over-the-link mode, not a
  CPU case. cpu_exec_enabled() and its process-wide LazyLock are gone; dispatch
  asks its own DispatchCtx, which reads the flag off the Gpu.

* GraphState::cpu_exec_weights, set once per load. All four capture entries plus
  Gpu::begin_stream_capture and ReplayController::begin_capture refuse while it is
  set, so a model with a CPU-executed step cannot be captured or replayed by any
  route - including the MTP-proposal and draft-FFN wrappers, which used to call
  hip.stream_begin_capture directly and bypass the guard. Replay's reset_for_model
  deliberately preserves it: the daemon calls that AFTER a load, so clearing it
  there would blind the gate for the model that just set it.

* LoadStats, measured from the tensors the source produced (WeightTensor defines
  owned_bytes next to free_all, and qwen35 mirrors free_moe_ffn branch for
  branch), replacing both the file-size estimate and the tensor-name parsing in
  serve_engine. Admission charges the pinned spill to the host tier up front, and
  serve preflight now runs on the loader's own numbers after the load.

The Redline refusal moved here from hipfire-dispatch: it needs the resolved spill
and the replay backend, and both live in the shared loader, so all four load
entries are covered by one check rather than four.

Verified on gfx1201: resident and 24/32 pcie-spill outputs still byte-identical to
the master baseline (greedy), and HIPFIRE_GRAPH=1 + offload_exec=cpu emits exactly
one 'hipGraph capture disabled' line with no assertion and coherent text.
weight_gemv_swiglu_residual's CPU arm called run_host_mapped_gemv_residual
directly, so the one op family that fuses silu+rotate+gemv kept its own private
convention for the seam. It now builds a GemvResidual step and calls
execute_steps, so the CPU decision, the rotation disposition and the launch all
come from dispatch and the recorded WeightRef.exec.

The GPU arms are deliberately untouched. Phase 5 as originally scoped replaced
them with two generic steps, which would have deleted the fused
silu_mul_rotate_mq / _awq arm that every resident MQ down-projection takes —
FUSED_TABLE carries no silu+gemv_residual key, so the pipeline does not re-fuse
it — and that arm also owns the AWQ divide-then-rotate ordering. That is the hot
path V1 certifies byte-identical, so the change is reduced to the arm the review
was actually about.

Verified: resident output still byte-identical to the master baseline, and the
cpu-execized output is identical to the pre-change run.
…_weight

LayerWeights::owned_bytes claims to mirror free_gpu branch for branch, and its
dense DeltaNet arm listed only attn_norm/ffn_norm/a_log/dt_bias — conv_weight and
norm_weight are released by the teardown and were never charged. Small in bytes,
but it is a silent under-count in the accounting admission reads, which is the
failure mode the function exists to prevent.

The helper now takes slices instead of a fixed 4-element array, so an arm cannot
drop a tensor by arity again, and the MoE arm's comment (which claimed the dense
arm carried the two tensors separately) is corrected.

Cross-checked as a difference: with 8 LinearAttention layers spilled, the fix
moves host_pinned_bytes by exactly (98_176 + 512) x 8 = 789_504 B, the conv/norm
bytes of those layers.
…is rework

hipfire_cpu::quant::dequant_group is the canonical CPU decoder — it is what the
CPU-exec seam runs — and weight_backend::dequantize_to_f32 now delegates for every
quant_type it covers, keeping its inline arms only for the formats the crate does
not name (so the FWHT un-rotation for those is untouched). With one implementation
the cross-check that held the two transcriptions together is tautological and is
deleted with its only consumer, dequantize_weight_to_f32.

That delegation is also a behaviour fix, found by the check itself: the canonical
decoder PANICKED on qt 49 (MQ3G256V2) while hipfire_cpu decodes it, so
dequant_f32 used to abort on a format the seam handles fine.

Evidence, in order: the cross-check was green before any change (1879 tensors,
qts {1,3,8,13,15,20}); with the delegation disabled and the fixture set widened it
compared the old inline arms against hipfire_cpu bit-for-bit over 3130 tensors
including qt 44 (qt 47 is covered only by the delegated run — its fixtures also
carry the qt-49 tensor that panics the old decoder). Then the delegation went in
and the file went away.

Qwen35Config::spill_refusal carries the MoE refusal so it is testable without an
HFQ fixture, and load_weights_inner prints the measured LoadStats once under
HIPFIRE_OFFLOAD_DEBUG=1.

docs/plans/partial-gpu-offload-design.md gains a post-review revision note and
inline retirements, because it is the design record for APIs this rework deletes.
Mechanical only: the repo formats the files a branch changes
(scripts/ci-rustfmt-changed.sh), and the exec-field sweep left ~40 insertions at
the wrong indent along with some re-wraps. Split from the semantic commits so a
reviewer can skip it.
…md5s

The battery+chain checkbox was ticked against the pre-rework implementation, while
the capture gate, the recorded exec target and the loader-side refusals all changed
since. Re-ran both modes plus the redline pcie-spill arm against the current code
and recorded the numbers with the model/hipfire/daemon md5s the reporting rule
requires (the local 9B is 296092bf…, not the pinned fixture).
…references

`dequantize_to_f32` delegates here now; the deleted
`cpu_quant_cross_check` test no longer exists to keep two copies honest.
…offload fields

The Settings rows for `memory.gpu_layer_budget` / `memory.offload_exec` stay,
but the writer no longer changes behaviour for any other setting:

- `write_value` keeps master's rule (an empty buffer is invalid); the generic
  empty-means-unset retry and the `IntOrUnset` field kind it needed are gone.
- Clearing a set budget is the row's Delete/Backspace
  (`reset_selected_setting` -> `writer::delete_key`), the same path every
  other row uses, so `persist_setting`'s `Null`/override bookkeeping reverts.

The only removals left in `crates/hipfire-tui` are the generated crate-map rows
and five lines of pre-existing rustfmt debt that the changed-file fmt gate fixes.

Tests: the two tests that encoded the empty-buffer clear now pin parity — an
empty buffer is invalid for the offload row exactly as it is for `max_tokens`,
`temperature`, `mmq_screen` and the pre-existing nullable
`deepseek4_experts_per_token` — plus Delete-clears and the `null` spelling.
…ling as unset

With `persist_setting` back on master's body, a write of the literal `null` (the
spelling `hipfire config set` uses to clear this nullable key) leaves that string in
`values` until the next reload, so the row read "null on GPU" even though the key
is off disk. Handled in the new row's own label (`"" | "null"` -> "all on GPU")
rather than by special-casing the shared write path, so no mainline line changes.
The arch rework moved six formats out of qwen35's private quant match and into
RAW_CODECS (qt 31/32/33/34/36/37), but the register still declared them
`arch-loaded`, which its own rule forbids for a format with a table row:
`scripts/check-quant-registry.py` reported qt_disposition_mismatch = 6 and
`scripts/leanup-ratchets.sh` failed on it. qt 35 (MFP4G32E8SOA) is untouched —
it has no RAW_CODECS row and qwen35 still owns its upload.

Also refreshes the generated crate-map blocks for the 15 crates whose counts
drifted across the rework, so `scripts/check-crate-maps.py --check` matches the
tree again.
…pu-offload

# Conflicts:
#	registry/v1.json
…tream

Merging upstream/master (66e825b "registry: add qwen3.8:27b-mq4-xts (H2) and
re-pin qwen3.8:27b-mq4-xt on master") reproduced the gate failure this PR kept
inheriting: that commit grew crates/hipfire-registry/src/lib.rs by 22 lines
(2,197 -> 2,217) without regenerating crates/hipfire-registry/map.md, so
master's own gates job fails on 'hipfire-registry: generated counts/content are
stale' — and any PR that merges master inherits it. That is why the failure was
unreproducible until the upstream merge was done here.

Also in this merge: registry/v1.json conflicted between the fork's registry-bot
commit (2f37c43, pulled in by our earlier fork-master merge) and upstream's
newer registry commit; resolved to upstream's copy. This branch never edited
that file.

Verified on the merged tree: check-crate-maps 44/44 matching, leanup-ratchets OK
(21 metrics, 0 violations), ratchet-diff vs upstream/master OK.
@aldrouil

Copy link
Copy Markdown
Contributor Author

1. Summary

I had an agent go over the critiques this morning. After a fair amount of work, its results are below.

It removed a few of the commits that were requested to be removed, I will submit bugfix PRs for those later. Including the TUI sanitation. I had the agent restore the existing behavior while still adding in the TUI fields.

As a side effect of the requests you asked for, qwen2 dense models should work with CPU offload. This has not been tested. There are no qwen2 models in the hipfire registry.


2. The review, answered

# review ask what shipped
1 placement belongs in the loader, not Qwen35Config Layout gained layer_residency, and spill_count / resolve_residency resolve memory.gpu_layer_budget once in load_weights_inner; WeightSource::read_layer(gpu, layer, residency) carries it into HfqBackend.residency. Qwen35Config::i_gpu_start and apply_offload_policy are gone. llama inherits placement, and qwen2 was moved onto the shared orchestrator so it inherits too — the one arch that was still bespoke. Manifest got the tier as asked: WeightPlacement.residency, populated in plan_manifest, and fulfill_manifest_single refuses a host-mapped placement rather than allocating it in the VRAM the spill exists to free
2 one upload path, not a qwen35 twin The registry absorbed the six formats qwen35 decoded privately (qt 31/32/33/34/36/37), decode_raw_codec takes Residency, and read_proj takes &mut Gpu — which is what removed the need for the second entry point. read_proj_host, host_local and qwen35's ~700-line private quant match are deleted: qwen35/load.rs −938/+76. expected_payload_bytes replaced six hand-copied length checks
3 downstream decisions read residency off the loaded weights WeightTensor::exec / WeightRef::exec are set once by the loader from residency × offload_exec, and the dispatch seam tests that recorded target instead of re-deriving locality. cpu_exec_enabled()'s process-global LazyLock and cpu_offload_active(i_gpu_start) are deleted, so no step re-derives the decision — and host-mapped weights under pcie correctly stay a GPU path
4 gate capture centrally GraphState::cpu_exec_weights is set once per load and read by all four begin_*_capture entries, Gpu::begin_stream_capture and ReplayController::begin_capture. The two bypasses are closed: mtp_spec.rs and dflash.rs used to call hip.stream_begin_capture directly and now go through begin_graph_capture. The Redline refusal lives in the loader's single seam. The review's own caveat — "spec decode + offload_exec=cpu is currently unvalidated" — is now exercised (§3). One item is still config-adjacent: ensure_fp16_shadow remains dtype-derived, so flipping the exec target at a fixed budget needs a fresh load
5 admission from what the loader allocated The loader reports LoadStats { device_bytes, host_pinned_bytes }, summed per tensor as it builds them. Those feed ModelFootprint directly, pinned host bytes are charged into the existing host tier, and serve_engine's tensor_layer_index + resident_weight_bytes are deleted — the preflight now runs on the loader's own numbers
6 fail closed on sources that cannot spill WeightSource::spill_refusal refuses a budgeted load for PaRoQuant and for any num_experts > 0 model before a byte is allocated, and a unified-memory device is refused on the same seam. The VL arm is a documented non-refusal rather than a refusal: this source is the text tower and the vision sidecar loads off the layer path, so there is no half-applied branch inside it to refuse
7 one CPU decoder hipfire-cpu is now a hipfire-runtime dependency and dequantize_to_f32 delegates to hipfire_cpu::quant::dequant_group for every format it covers, keeping inline arms only for the formats the crate does not name; the cross-check test and dequantize_weight_to_f32 went with the duplication. Making it canonical found a latent bug: the canonical decoder panicked on qt 49 while hipfire_cpu decodes it correctly. One table remains, argued in the review's own terms: cpu_quant_for is keyed by DType while the registry is keyed by quant_type, and several DTypes have no qt row, so folding the decoder into the registry row would make that mapping lossy — cpu_decoder_tables_agree_per_format pins the two in step
8 use Steps for the fused FFN The CPU arm now builds a GemvResidual step and calls execute_steps, so the seam owns the decision, the rotation disposition and the launch. The GPU arms are deliberately untouched: FUSED_TABLE has no silu+gemv_residual key, so two generic steps do not re-fuse — substituting them deletes the fused_silu_mul_rotate_mq* arm and its AWQ divide-then-rotate ordering, on the path the zero-diff guard certifies
9 trim scope Three of the four splits are their own branches cut from master: fix/gemv-mq4g256-kernargs (45e70134), chore/rdna-compute-gpu-test-serialization (b6c39a9d), docs/offload-perf-records (2bae014f). largest_fitting_tail, resolve_i_gpu_start and GpuLayerBudget::Auto are dropped. The TUI rows stay, and the TUI diff removes no behaviour-bearing line: 620 insertions / 11 deletions over 6 files, where the 11 are the generated crate-map rows (6) and five fmt-mandated reflows in app.rs (5) — moves, not deletions (the assertions they carry are all still present), and required because master's own app.rs fails rustfmt --check and scripts/ci-rustfmt-changed.sh checks whole changed files. write_value and persist_setting are byte-identical to master: an empty buffer is invalid for every numeric row, and clearing a set row is the row's Delete/Backspace (reset_selected_setting → writer::delete_key), the path every other row already uses. A parity test pins that over max_tokens, temperature, mmq_screen and the pre-existing nullable deepseek4_experts_per_token; the only new-code accommodation is that the new row's label renders the schema's null spelling as unset. The rows themselves are the keys' discoverability (both are settable from the CLI and the config file), and a TUI-only branch off master does not compile, because master carries neither key, which makes it a stacked PR that must merge with or after this one

The umbrella concern — arch crates holding as little as possible — is where the diff goes:
qwen35/load.rs lost 938 lines net of 76 in one commit, and placement, the reader, the
execution target, the capture gate, the refusal and the admission accounting now live in
hipfire-runtime / hipfire-dispatch.

3. Evidence, all of it from this machine (gfx1201, ROCm 7.2)

  • Zero-diff guard, two fixtures. Greedy (-t 0), master's binary against this branch:
    qwen3.5-2b.mq4 (md5 9ed6628f2df83ef4b1c062afd4a85bfb) — stock ≡ branch ≡ 3/24 pcie
    spill, byte-identical text; qwen3.5-9b.mq4 (md5 296092bf1e6a45d78c1acf815eb93366) —
    stock ≡ branch resident, and 24/32 pcie spill ≡ resident.
  • Capture gate. HIPFIRE_GRAPH=1 + offload_exec=cpu + spill: exactly one
    cpu exec: hipGraph capture disabled line, no assertion, coherent text — and the same with
    DFlash active, which is the spec-decode path the review flagged. Under pcie instead,
    capture stays on: zero disable lines, and redline_daemon_harness.py with a 24/32 spill
    returns pass=True — capture stable at 128/512/decode, 19 AQL contracts, shadow parity
    exact=True.
  • Spec decode with a spill, both exec targets. DFlash + cpu: draft loaded, one disable
    line, no assertion. DFlash + pcie + spill + HIPFIRE_GRAPH=1: draft loaded, zero disable
    lines, and the output is byte-identical to the resident run.
  • Redline refusal. Spill + offload_exec=cpu + replay.backend=redline: the load is
    refused, the message names both keys, no text is produced — the intended failure rather than
    a replayed stale route.
  • Serve path. serve_harness.py battery and chain on the cpu-spill arm: 5/5 turns
    each, runaway=0 empty=0 attractor=0 retrieval_miss=0, decoded text read (a real
    merge_sorted, the correct axial-tilt answer, on-topic prose and instruct turns), and
    chain shows the prefix cache working (cached=431/682/797/943).
  • Accounting against the device and the OS. LoadStats' device+host total is bit-invariant
    across budgets (5 300 590 592 B). The marginal anonymous-RSS growth of the spill matches
    host_pinned_bytes plus the 106 × 1 MiB host-tail pads to −0.01 % (−119 296 B on
    1 031 070 208 B). Measured VRAM sits a fixed ~300 MB above device_bytes while 920 MB of
    weights move between budgets, i.e. allocator granularity rather than a scaling error.
  • Workspace gates. cargo build --release clean; cargo test --lib --workspace 2991
    passed / 0 failed
    (42 lib targets, rc=0, re-measured on the final commit); check-layering
    0/0; check-env-docs clean; rustfmt clean on all 145 changed .rs files; no-gpu-ci.sh Rust
    phases green. Test coverage is additive too, since the review flagged removals: the diff adds
    57 #[test]/#[tokio::test] attributes and removes 3 (net +54; 4092 → 4146 over
    crates/**/*.rs), and those three are the old f16_lm_head_mode_* shape — that policy now
    lives in hipfire-config::kernel::parse_lm_head_f16_native, which carries its own tests.
    Its Python phase fails five mq4c_repack tests for a reason independent of this diff: the
    module never defines parse_hfqm_index, HfqmError or main, this branch touches no
    Python, and both files were last modified 2026-08-18.

4. Replication against the pinned 2B, using the record's recipe

hipfire bench <2B> --spec off --runs 5 --warmups 2 --max-tokens 64 --backend noslots --workload stateless --prompt-file benchmarks/prompts/gpu_offload_probe.txt --json, one fresh
daemon per point, prompt md5 5835c71e471849b4a72e1dc8e39695e7 and model md5
9ed6628f2df83ef4b1c062afd4a85bfb — the same artifacts 2026-09-27-cpu-exec-spill-sweep-2b-9b-27b.md
pins.

budget (spilled) pcie cpu ratio record pcie record cpu record ratio
24 (0) 269.2 269.2 1.00 258.5 258.4 1.00
21 (3) 138.2 103.5 1.34 131.1 99.3 1.32
15 (9) 71.8 53.2 1.35 67.5 48.5 1.39
12 (12) 57.6 41.5 1.39 54.3 38.7 1.40
18 (6) 94.2 68.8 1.37 62.6 65.5 0.96

The pcie-vs-cpu ratio reproduces (1.34–1.39 against 1.32–1.40), and the record's observable
gates reproduce exactly: b24 cpu prints offload_exec=cpu but nothing is spilled with zero
capture-disabled lines, b12 cpu prints 12/12 spilled layers fully covered; uncovered quants: none with one, and pcie points print none. The ~4 % absolute offset is the host, not the
code — the no-spill control shows the same +4.2 % (269.2 vs 258.5), and scaling by it puts
every point within 1–3 %. Only b18 differs, and its recorded pcie median sits inside its own
round spread ([62.6, 62.6, 89.3], against our 94.2), so it is one depressed round in the
record rather than a disagreement; the other four points replicate.

The cpu arm carries no byte-identity claim by design: its contract is llama.cpp-level —
coherence plus a measured divergence — which is why only pcie is byte-checked.

Branch binaries for every number above: hipfire
2b3af98c16bf79d1e4b1bc1e6aa37e31, daemon 3cd0af9390cfaebeb2de099abd60bcb2.

Kaden-Schutt pushed a commit that referenced this pull request Sep 30, 2026
…onto land/040)

Squash of PR #793 (warpfront/hipfire, head 307009e,
31 commits by Avery Drouillard) onto land/040-fixes 9eb9e1d.

memory.gpu_layer_budget / HIPFIRE_GPU_LAYER_BUDGET spills a prefix of dense
Qwen3.5 layers to hipHostMalloc-mapped host RAM; memory.offload_exec=cpu /
HIPFIRE_OFFLOAD_EXEC runs those layers' weight GEMVs on the new hipfire-cpu
crate. Unset keeps every layer resident.

Rebase resolution (fold/serve):
- CHANGELOG: the PR's entries are replaced by one consolidated entry in the
  follow-up commit.
- crate maps: generated blocks regenerated (no hand-written map edits in the PR).
- load.rs: land's qt=52 (MQ4-G256 v2 Lloyd) arm uploads through the injected
  uploader like every other raw arm, so a spilled Lloyd layer lands in host RAM.
- prefill.rs: land's widened_test_config literal gains i_gpu_start: 0.
- Cargo.lock: hipfire-cpu at the workspace version 0.4.0.
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
… retained Redline default for CPU-executed offload

Follow-up to the #793 squash for 0.4.0 (serve audit O3).

- hipfire-loader admit_source refuses a memory.gpu_layer_budget that would
  spill layers under tp>1 (dense-TP and EP rank loaders keep every layer
  resident), pp>1 (the first band would spill into host RAM mapped by device
  0) or on a MoE model (the expert stacks stay in VRAM), before any teardown.
  An unset budget skips the check and the extra config parse.
- HfqSource::prepare backstops the same refusal (MoE, n_devices>1) for every
  other qwen35 load path and prints the residency line there, once the
  placement is applied; config parsing no longer prints it.
- retained_redline_default: a qwen3_5 process configured with
  offload_exec=cpu and a layer budget gets no retained default. The tape
  records GPU launches only and would skip the host GEMVs.
- append_betaalpha_to_z keeps the folded Z|beta|alpha rows of a spilled
  layer in host RAM (upload_raw_host + free_tensor) instead of moving them to
  VRAM and passing a host-mapped owner to release_tensor_immediate.
- CHANGELOG: one consolidated #793 entry in a new "folded PRs" subsection.

Co-authored-by: Avery Drouillard <avery@averynet.xyz>
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
…XPERT_VRAM_LAYERS) + routing trace

Experiment branch for Flash-Next on gfx1201 (32 GB). Trunk layers 0..N keep
their routed experts in VRAM; later layers' and the MTP layer's experts are
fulfilled into pinned device-mapped host RAM (hipHostMalloc Mapped, ported
from #793's hip-bridge/rdna-compute owner) and read zero-copy over PCIe by the
unchanged sealed MoE kernels through their per-layer pointer tables.
New WeightResidency::HostMapped; HIPFIRE_QWEN4_ROUTE_TRACE dumps per-layer
routed expert ids for cache sizing.
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
# Conflicts:
#	crates/hip-bridge/map.md
#	crates/hipfire-arch-qwen35/map.md
#	crates/hipfire-config/map.md
#	crates/hipfire-dispatch/map.md
#	crates/hipfire-generate/map.md
#	crates/hipfire-loader/map.md
#	crates/hipfire-runtime/map.md
#	crates/rdna-compute/map.md
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
…t serve, opt-in)

Conflicts vs fold/serve (#793 offload, #787): Cargo.toml features (sha2 non-optional
per SCS; keep #793's lab feature), serve_engine planned/weights bytes (SCS paged KV
formula + #793 resident_weight_bytes), config.rs tests (both), hfq.rs (keep #787
probe_arch_id), cli local_model_paths (#787 registry arg + SCS catalog tags).
vs fold/kernels: attention.rs (both helpers). Semantic: #793's offload guard in
forward_batch_slots_graphed_opts now falls back through SCS's
forward_batch_slots_opts (kv tier, lm_head_skip, spec_capture). CHANGELOG entry
for the stack added.
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
User decision: hipfire quantizes from BF16 sources; GGUF -> mq4 double-quantizes,
and GGUF users should use llama.cpp. Reverts 02e60cd (the only #785 commit in
fold/serve); CHANGELOG entry removed and the folded-PRs subsection retitled;
hipfire-quantize map regenerated. The rest of fold/serve (#793, #787, #595) stays.
Kaden-Schutt added a commit that referenced this pull request Sep 30, 2026
…outed experts on discrete cards, auto expert VRAM layers, host-RAM preflight)

Conflicts: hip-bridge ffi.rs/lib.rs and rdna-compute dispatch.rs. The branch
adds the same hipHostMalloc(Mapped) owner class (HIP_HOST_MALLOC_MAPPED,
host_malloc/host_get_device_pointer/host_free, DeviceBuffer::HostMapped,
Gpu::host_mapped, the free_tensor/ensure_vmm_cleaned arms,
HOST_TAIL_PAD_BYTES) that land already carries from #793's partial offload;
kept land's copies. The auto-merged duplicates (a second
HIP_HOST_MALLOC_MAPPED, release_registered_host_mapped, host_mapped_count)
are dropped. `Gpu::upload_raw_host_mapped` (the Qwen4 expert upload) keeps
the branch's behaviour, a CPU copy through the host pointer plus a zeroed
tail pad, on land's `alloc_host_mapped_tensor`. Maps regenerated
(hipfire-arch-qwen4 gains expert_residency.rs, loader, runtime,
rdna-compute); env-vars inventory regenerated (HIPFIRE_QWEN4_EXPERT_VRAM_LAYERS,
HIPFIRE_QWEN4_ROUTE_TRACE).
aldrouil pushed a commit to aldrouil/hipfire that referenced this pull request Oct 3, 2026
… of config

One home for "where do this model's weights live", so a second arch adopts a
placement policy instead of copying one (PR warpfront#793 review points 1/3/5/6/9).

`hipfire_runtime::offload`:
* `Placement` — per-layer `WeightResidency` for a layer's own weights and
  `ExpertResidency` for its routed experts, kept as distinct types so the two
  axes cannot be swapped silently.
* `LayerBytes` — the arch's byte census (non-expert / expert / always-resident),
  plus `device_bytes` and `host_expert_bytes`.
* `plan(layers, budgets, capacity, orientation)` — llama.cpp's tier order:
  experts first, whole layers only when the expert tier cannot free enough, stop
  at the first fit. `Orientation` expresses both the dense reading (the last N
  layers resident) and Flash-Next's (the first N), so neither arch forks the
  search. An explicit `Layers(k)` pin is a hard constraint: a host-placed layer
  would take its pinned experts with it, so the layer tier is capped by it.
* `Capacity` / `KvReserve` / `kv_reserve` (checked multiply).
* Host admission lifted from `hipfire-arch-qwen4::expert_residency` with its
  unit tests: `check_host_ram` (MemAvailable + the TTM page-pool estimate),
  `gtt_budget` / `check_gtt_cap` (TTM pages_limit), `ttm_pool_estimate{,_from}`,
  `admit_host_placement`. No fixed host budget anywhere.
* `report` — the residency line, including the per-layer routed-expert MiB an
  operator needs to choose a count.
* `largest_fitting_tail`, moved out of `hipfire-config`: placement arithmetic is
  not configuration vocabulary, and it needs a measurement config cannot take.

`OffloadBudget` (renamed from `GpuLayerBudget`) stays in `hipfire_config::memory`
— config has no intra-workspace dependencies, so it owns the vocabulary while
`offload` owns the arithmetic. One enum now covers both knobs.

Tests: `cargo test -p hipfire-runtime --lib offload::` 16 passed;
`-p hipfire-config`, `-p hipfire-arch-qwen4`, `-p hipfire-loader` green; the whole
workspace still checks.

MoE-offload plan, stage 1 step 4 (first half; the qwen4 port follows).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants