Conversation
…_gpu_start Squash of 1f75880, efa8e6e, 60a56f8, ec83496. 1f75880 added the host-located VMM allocation primitive (backing pages in system RAM, mapped into GPU VA). Root-cause fix it hinged on: HIP_MEM_LOCATION_TYPE_HOST was defined as 0, which ROCm reads as Invalid/None; corrected to 2 (Device=1, Host=2 per driver_types.h), and the C VMM probe confirms gfx1201 host-location granularity 4096 once correct. New surface: HipMemAllocationProp::host_pinned(), HipMemLocation::host(), mem_get_handle_properties() fail-closed placement check, MemoryLocality enum + allocation_prop router + reserve_host(), Gpu::alloc_vmm_tensor_host(). Verified gfx1201: HOST_OFFLOAD PASS smoke, hip-bridge vmm unit tests 6/6. NOTE: this VMM mechanism is superseded later in the stack by hipHostMalloc (95c7b6a) — the primitive, not the enum fix, was reverted. efa8e6e added memory.gpu_layer_budget (Full | auto/-1 | N resident) plus largest_fitting_tail() pure admission fn (refuse-cleanly None when nothing fits). Tests: 9 memory module tests; hipfire-config suite 83 passed / 0 failed. Fixture pin: qwen3.6-27b.mq4 size 14984158208 (sha256 86a5f80f...). 60a56f8 added resolve_i_gpu_start (Full->Some(0), Auto->tail, Layers(n) ->n_layers-min(n,n_layers)); pure, unit-testable, no behavior change. ec83496 added Qwen35Config::i_gpu_start (layers [0..i_gpu_start) spill); data-only, default 0 fully resident, reversible.
Squash of 88cc8ff, defd51a, 2ef2dfb, a545410, 0455781. 88cc8ff: upload_f32_host / upload_raw_host (+upload_raw_with_copy_host with arena-release-on-copy-fail). Allocation half of offload; no callers yet, behavior unchanged. defd51a: HfqBackend.host_local + read_proj_host: Option<fn(..., &mut Gpu)> twin seam (&Gpu device reader cannot host-allocate since arena registration needs &mut Gpu). None = hard error on offloaded layer, never silent device fallback. CPU dequant factored into dequantize_norm/dequantize_to_f32 so both paths share the attractor-critical math. All other construction sites set host_local:false, read_proj_host:None — neutral. 2ef2dfb: load_weight_tensor_host via shared load_weight_tensor_raw_with (one copy of 36 upload sites + K%256/blob-length guards; qt 32/33 guards kept). [m,k]-shaped arms (qt 1/2/16) host-localized too. f32-dequant fallback quants REFUSED on offloaded layers (fail-closed). AWQ scale stays device-side deliberately (1-D f16 vector of len K). a545410 + 0455781: host_offload_smoke proves byte parity on real data (17825792 bytes identical, gfx1201 qwen3.5-9b.mq4, q_proj qt=13 [8192,4096]) AND locality (vmm_host_located asserts both directions + blob-size check), with a fixture picker (2-D raw-code tensor, dims>=1024) instead of a hardcoded name. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…cope Squash of 21ff5c9, 95b54a1, 89a8567. Docs only; no code change. 21ff5c9: four settled decisions — proj needs the read_proj_host twin seam; f32-dequant fallback fail-closed for offload; byte-identity claims require greedy decode (default-temp engine nondeterministic, measured on branch and stock); target fixture qwen3.6-27b.mq4 (14984158208 B, sha256 86a5f80f...). Latent MoE hazard noted (load_moe_ffn never sees HfqBackend). 95b54a1: hipfire bench pins greedy itself (temp 0.0/top_p 1.0), so the -t 0 rule is run-only; arms must interleave across fresh processes (consecutive same-arm runs drift ~2x — measured branch 38.5 then 22.8 vs stock 23.0 then 40.4; interleaved medians 40.25 vs 40.35 tok/s, 0.25% apart). Resident VRAM control ~13.7 GiB (86% of 16304 MB) via rocm-smi card0 CSV field. 89a8567: scopes -t 0 to run; bench output not for eyeballing (empty-think template); VRAM control from bench JSON (13042 MB branch vs 13098 MB stock, ~13070 MB) — offload proven only if vram_free_mb lands well above ~3100 MB. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ng, repro Squash of b8b6520, 71e81c5, 63224ba, f51fd44, 7364371, ea61cc7. Intermediate diagnosis on the VMM path; mechanism superseded by hipHostMalloc (95c7b6a), probes retained as regression coverage. b8b6520: hipMemMap needs granularity-aligned size; both alloc_vmm_tensor paths passed caller bytes through. Real MQ3 blob (99840 B vs 4096 gran) died "must be a non-zero multiple of granularity 4096". Fixed on BOTH paths (device twin had the identical latent defect). Found by the first real offloaded load, not smokes (aligned F32 only). SURVIVES in final tree. 71e81c5: host_kernel_access probe — real gemv_f32 over host VMM vs device twin at the exact faulting 99840 B size. HOST_KERNEL_ACCESS PASS (max diff 0): platform is NOT the limitation. 63224ba: wired apply_offload_policy -> i_gpu_start -> per-layer host_local. Default unset = fully resident, byte-identical to stock (verified vs hipfire3 on qwen3.8-27b.mq3-xt). Auto(-1) failed closed (no capacity measurement at config build). EP/MoE pinned resident. Offloaded path FAULTED ("Page not present"); ruled out platform/graph/scale — remaining suspect: quantized GEMV over host DType::Raw blobs. f51fd44 + 7364371: two-arm repro (F32 PASS max diff 0; MQ4 FAULT on host VA) with non-vacuity guards (device ref must be non-zero, exact match) and layout asserts (row_bytes=(K/256)*136, K%256==0). Fixture bug fixed en route (f32 scale/zero headers, not arbitrary bytes). ea61cc7: root cause — forward kernels overread the tensor tail; exact-fit VMM (every observed MQ3 blob an exact 4096 multiple) made the overread the first unmapped byte. Fix: pad reservation + map whole, owner_buffer still exact prefix. Verified: BUDGET=62 run -t 0 decodes, md5 8826a76be712ad4b9d74c96021f18bda byte-identical to resident. +OFFLOAD_DEBUG VA logging. NOTE: ea61cc7's TODO.md blocker log was deleted by 998e55a (correctly — resolved); f51fd44's "MQ4-specific" narrowing was retracted in that log (device ref all-zero) and is not re-asserted here. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Squash of ef4aa23 and 998e55a. The load-bearing pair of the stack. ef4aa23: offload was net-zero for VRAM — host-located VMM arena charged the device heap 1:1. Measured gfx1201/ROCm 7.2 via 1 GiB hipMalloc ladder with 4096 MB held: control 15360 MB, VMM arena 11264 MB (cost 4096 MB), hipHostMalloc(mapped) 15360 MB (cost 0 MB). Production uploads rerouted to new Gpu::alloc_host_mapped_tensor (tail pad now costs host RAM only). Also fixed latent gemv_mq4g256 bug: 7 kernargs to a 5-param kernel (FWHT tables belong to mq_rotate_x), m/k misread, all-zero result — no production callers, confined to the diagnostic. Verified greedy -t 0 (code_edit_rewrite_copy.txt md5 80b910784c456bbb5532eb0493a497b9, qwen3.8-27b.mq3-xt): byte identity md5 8826a76be712ad4b9d74c96021f18bda (resident, BUDGET=62, BUDGET=32); footprint resident 13042 MB vs budget-32 7992 MB (-5050 MB for 4946 MB spilled; pre-fix arms read 13042 vs 13480). headroom PASS (cost 0 MB), kernel_access ARM1_F32 + ARM2_MQ4 PASS (bitwise-equal, non-vacuous, via production path). Decode at budget 32 PCIe-bound by design (4.4 vs 39 tok/s) — capacity feature, not speed. Follow-up noted: pinned host RAM stalls under memory pressure (layer 38 stall at 0 MB free) need watching. 998e55a: deleted the dead host-VMM allocator (MemoryLocality/HostPinned, reserve_host, allocation_prop router, host_pinned(), alloc_vmm_tensor_host; vmm_host_located -> host_located ownership check) — branch-only code, trap for re-adoption. Migrated smokes to production upload_f32_host (+ no-arena + owner-count asserts). Also fixed pre-existing vmm_tensor_smoke failure (non-granular map rejection contradicted b8b6520's round-up since ea61cc7). Verified gfx1201: build clean, all smokes PASS, host_offload_smoke qwen3.6-27b.mq4 PARITY PASS (27852800 B identical).
…g sweep Squash of b2057dd, 2e7b65e, feb4bec, 50d0822, 7611a45. b2057dd: GPU tests in dispatch::tests serialized on a module lock (try_gpu returns Gpu + guard). upload_raw_copy_failure asserted byte-exact hipMemGetInfo free while siblings shared the device — failed ~1 in 3 runs with no source change. No assertion weakened. Verified 6/6 parallel runs + four-package gate. 2e7b65e: SlotEngine gate charged the whole model file; now charges the resident share summed from the HFQ index (layers non-uniform; non-layer tensors always resident). Off == file size exactly (no-op for existing configs). Pays off: max_seq=262144 un-loadable resident (OOM, 330 MB free) loads at BUDGET=32 and emits resident-identical md5 8826a76be712ad4b9d74c96021f18bda. Coherence: battery+chain coherent, attractor=0; unclosed-think turns fail identically on resident control (pre-existing). Offloaded decode 4.5 vs 40.0 tok/s (8.9x PCIe penalty). feb4bec: copy-fail assert 64 bytes exact -> 64 MiB signal / 8 MiB tolerance (leak of this owner is 64 MiB; kills cross-process settling flakes). Pool-counter equality unchanged. NOTE: superseded twice more in stack (1c2fa50: 512 MiB/64 MiB on measured 28-56 MiB never-returned driver overhead; 9bb5992: 500 ms poll for late-vs-lost reclaim) — see commit 8 for final form. 50d0822: prose sweep host-VMM -> host-mapped system RAM (i_gpu_start doc, host_located doc); 5 deliberate "rejected alternative" mentions kept. Verified smokes PASS. 7611a45: release_registered_host_mapped removes only-if-released (was unconditional drain -> invisible leak on hipHostFree failure); ensure_vmm_cleaned refuses on live host_mapped owners too (+ naming test). Changelog: weight sweep 1814 ms resident vs 57592 ms at =32; offloaded Redline parity unverified (resident PASS exact; offloaded load >900 s timeout, 1073.5 s sweep under harness vs 57.6 s plain).
…ding Squash of 3018ddf, 12fa133, e4647e2, c718d8d. 3018ddf: gpu_layer_budget help stopped promising auto-fit (-1 fails closed; largest_fitting_tail unwired to a device measurement). Changelog: offloaded arm passes routed Redline at completable budget (=62, identical kernel hashes); =32 load exceeds 900 s under harness. 12fa133: Offload row added to easy-mode lists (all four positional lists + knobs::KNOBS explainer). Unset renders "all resident" (label never round-trips into config). Explainer: counts RESIDENT layers (32/64 spills 32); sharp two-sided cost (per-step PCIe + slow load); -1 unimplemented. Verified on rendered surface; tui tests 153 pass. e4647e2: render-level guard (label+value pair + explainer title via TestBackend). First draft asserting contains("Offload") was near-vacuous (passes on "OffloadXX"); now whitespace-collapsed "Offload all resident", sensitivity checked both ways. 154 pass. c718d8d: row editable (EDITABLE_FIELDS spec; reasoning_effort same fix; REASONING_EFFORTS exported from hipfire-config; -1..65536 bounds, null clears) + wording un-inverted (Offload/all-resident -> GPU layers/all on GPU; direction with concrete example; auto as unavailable not failure). Three guards proven to fail without fix (class guard, typed-value end-to-end, help coverage). Arrows-on-integers left as-is (tab-uniform). Verified: tui 156 + config 88 pass.
… hardening Squash of 03e963c, 7b5eae6, 1c2fa50, 9bb5992. 03e963c: auto(-1) no longer fails the load (unresolvable at config build: needs measured capacity + per-layer bytes, no Gpu in hand) — keeps every layer on GPU with a log note. NOTE: framed as "placeholder" here; 7b5eae6 reframes as engine-decides (below) — latter reading stands. No-op budgets reported (Layers(200)/64-layer saturates silent -> now explicit). Render guard made hermetic (both renderings pinned, override marker handled). Placement arithmetic factored to pure offload_split + residency_report (apply_offload_policy infallible; silence still means unset-default). Tests pin direction/saturation/degenerate; verified 23 ok-blocks across 6 crates. 7b5eae6: auto = "the engine decides" like every other key (kv_cache, dflash/mtp/flash/mmap_screen); currently decides resident until capacity measurement feeds largest_fitting_tail. Behavior unchanged (Auto=>0); vocabulary + docs + CHANGELOG only. 1c2fa50: feb4bec's guessed 64 MiB/8 MiB still failed ~2/5 (28 MiB retained exactly). Measured via free_retention probe: owner bytes always return exactly; driver intermittently charges 0/28 MiB (64 MiB req, up to 56 MiB at 2 GiB) hipFree never returns. Now 512 MiB signal / 64 MiB tolerance (no observed overhead reaches it; leak drops 512 MiB). Proven: skip-hip.free reddens with "536870912 bytes came back short"; 8+3 runs + full gate pass with user serve live on same GPU. 9bb5992: poll get_vram_info up to 500 ms — probe showed reclaim LATE not lost (11/12 exact, missing 28 MiB on later pass). Tolerance covers never-returned overhead; poll covers lag. Proven: leak still reddens ("never came back ... after 500ms"); 6 runs + six-package gate pass.
…4 path Squash of 1d40e51, 6e6a6aa, c83f64f, ea74e1d. 1d40e51 (S1): hipfire-cpu leaf — decode_group_codes (affine/codebook only, what gemv dots against forward-rotated x) vs dequant_group (+inverse FWHT, canonical basis); mixing them is silent R^-2 not a crash. 10 formats (qt 13/44/15/8/17/20/1/2/16/3); qt 20 measured from registry qwen3.5:2b-mq3 (plan's "mq3=qt17" would miss every shipping -mq3 artifact). Evidence: cpu lib 19 tests (canonical-generated tables, WH oracle); cross_check 1695 real tensors bit-identical; gpu_gemv_parity 8 formats worst 6.7e-7 vs 1e-4 (real qt13/20/15/8 rows + synthetic bit-exacts). HIPFIRE_OFFLOAD_DEBUG via developer_var (fixes env-docs check). 6e6a6aa (S2 seam): memory.offload_exec=pcie(default, byte-identical today)|cpu; fusion skipped on CPU windows; Step::Gemv/GemvResidual over host-mapped weights run D2H->[AWQ divide]->rotate->hipfire_cpu::gemv->H2D (residual accumulates in place); swiglu-down splits (GPU silu_mul_f32 + CPU GEMV). host_bytes registry resolve; memcpy_dtoh_auto refuses under capture; graphs disabled for model lifetime when spill possible. Two real bugs fixed (invisible to parity): AWQ sidecar divide (138 fixtures) + double rotation of Prerotated inputs. Contract: llama.cpp-level coherence + measured divergence, not byte-identity (2B: 1343-char pcie md5 5722aca7..., cpu 652-char agreement then diverges; 9B likewise 723-char). Boundaries: slots/serve GEMMs + prefill stay GPU-side; lm_head always-resident; Q8HFQ/Givens/Paro refused. Tests: 21 cpu / 89 config / 272 dispatch / 251 rdna-compute. c83f64f (S3): coverage table (7 shapes, 10 roles, 2B dim 2048; trace "0 host-mapped steps still on GPU", rotation tags real). Boundaries verified: slots/serve + prefill uncovered; lm_head never host-mapped; cpu+spill+ Redline = load error (verified both ways); exec=cpu with no spill = one info line, no-op. ea74e1d (S4): AVX2+FMA Mq4G256 row dot (32 nibbles/16 B, scale·Σ+zero·Σx per group; 1e-5 vs scalar, scalar pinned to S1 arithmetic) + per-step D2H/GEMV/H2D trace + cpu_exec_time_ns. Measured gfx1201/7800X3D 9B mq4 (binaries+prompt md5s in orig): spill-8 cpu 24.1 vs pcie 21.6 tok/s (+11.6%), spill-16 +27%; vram_free identical (capacity unchanged); ~23 vs ~17 GB/s effective. Ceiling 44.5 GB/s host-mapped (44.8 heap), limit is per-step serialization not codegen. Routes: cpu battery+chain 5/5 coherent; pcie redline PASS exact; cpu redline = intended refusal; no-gpu-ci green (modulo pre-existing mq4c_repack failure).
Squash of 0ffb07d and 733f350. Docs only; no code change. 0ffb07d: dated historical record (docs/perf-checkpoints/2026-09-27-gfx1201-cpu-exec-offload.md) — S4 numbers previously lived only in a commit message and /tmp. Supersedes the two 3-6x-slower readings (contaminated: concurrent release build + second sweep on same 16 threads). Clean interleaved: 24.1 vs 21.6 tok/s cpu at 8/32 spilled (+11.6%; +27% at 16/32), vram_free_mb identical. Beyond medians: 72 steps/token x ~0.53 ms ~= 38 ms vs 41.5 ms/token (no hidden serialization; copies ~18%); kernel ceiling 44.5 GB/s host-mapped (44.8 heap), 84 GMAC/s x16 / ~11 x1 (Zen 3, no AVX-512) — cpu arm at ~80% of own ceiling; determinism (two processes + x16 vs x1 byte-identical); divergence = single whitespace token at ~190 tokens then drift (near-tie signature). 733f350: shared-prefix length = argmax-survival count under 6.7e-7 per-step diff (sensitivity, not accuracy) — stated in 6.2.1 so it is not read as an error rate.
…nt record Squash of 853a386, 74b0886, 924d7fd. 853a386: CpuQuant 10 -> 28 formats (every dense-capable qt-register entry). qt 1/2/3/16, 6-12, 13/15/17/18/31, 19/20/30/51 Lloyd, 40/41 ternary/binary, 44/45/47/48/49/50 V2+MQ4C. Packings from canonical decoder (bit-exact tables for 7/11/12/18/19/30/40/41) or mq-v2-family.md wire spec (qt 47/48/49/50, 45 — transcribed from gemv_mq*g256v2 kernels, byte-agree). Evidence: cpu lib 24 (layout table all 28, f64-oracle GEMV, 19 expectation tables); cross_check 1695 real tensors bit-identical; gpu_gemv_parity worst per format (Mq6G256V2 1.9e-7, Mq5G256V2 2.9e-7, Mq3G256V2 7.2e-8, MQ4C/MQ2-fam/HFQ* bit-exact); qt 49 real-27B q_proj (12288x5120) 2.4e-7 / 3.5e-7 — the row making all-qt49 27B CPU-executable. Fixes: coverage line now resolves via loader dtype_from_quant_type (was RAW_CODECS passthrough, blind to arch-loaded qt 31/MFP4/PARO); qt 11/12 reported arch-ineligible (no dense GEMV kernel on gfx12) not failed. Structural exclusions named (qt 5 stride, 14 Mq8Internal, 21/24/32-37 HFP4/MFP4 layout, 28/29 PARO, 38/39/22 MoE/non-weight). E2E 2B spill-12: 12/12 covered, 0 on GPU. 74b0886: PERF BUG (shipped in prior two commits): V2 family + qt 44 widened per-128 fp16 [s0 z0 s1 z1] per element (branchy half->f32 => conversion-bound). Hoisted to v2_halves/half_of: 27B qt49 m=12288 k=5120 19.95 -> 5.70 ms/step (3.5x; run 286 s -> 163 s). Arithmetic order unchanged (24/24 tables bit-exact, cross_check still 1695 green). qt 44 = current --format mq4 output, so all mq4v2 bodies were affected (2B/9B qt 13 flat-f32-header hid it). Parity now covers BOTH rotation arms (check_prerotated all formats: qt 45 bit-exact both arms despite "HFQ4-G256" header comment). Record corrections: Bf16 from_quant_type(16) row was missing (S1-S3 claimed coverage wrongly — fixtures carry none, so their lines stayed true); 6.2.1 coverage bullet updated to 28-format set. Remaining lever: only Mq4G256 vectorized; qt 49 scalar ~11 GMAC/s vs qt 13 30-60 GB/s — AVX2-for-V2 next, not correctness. 924d7fd: dated amendment for qwen3.8-27b.mq3-xt (497 projections all qt 49): budget-56 (8 spilled) loads 8/8 covered; budget-48 (16 spilled, ~2.5 GB pinned) stalled 2x at layer 62/64 with 0 GB free — pinned RAM is a load-time capacity knob. qt 49 post-hoist real-tensor rows re-measured unchanged (49 7.19e-8; 44/47/48 + 45/50 rows). STATED GAP: no eyeballed completion on this fixture (160-token run hit reasoning-budget gate; 2 loads stalled; 600-token run same gate) — qt 49 rests on parity + coverage/trace, not decoded text.
Squash of ab090d7, 7133b3c, 8b2c3f0. ab090d7: self-contained benchmark handoff runbook (fixture digests for 2B/9B/27B + prompts + binaries; knobs + split assertion; interleaved fresh-process loop; five contamination modes that each cost a sweep leg; run matrix; output-reading guide; sanity anchors + expected crossover; 27B footprint recipe; report requirements). Indexed in docs/INDEX.md; HIPFIRE_CPU_EXEC_TRACE added to env inventory (env-docs green). 7133b3c: identity table gains git rev-parse HEAD beside binary md5s; contamination table (symptom->cause->check: concurrent build ~300% on cpu arm, surviving daemon, reused-pid kill, steps:0 = trace unset, uncovered quant, prefill delta impossible, vram delta = pass-condition failure, attractor-as-tight-stddev); knobs frozen between arms (leave HIPFIRE_VERIFY_GRAPH alone). 8b2c3f0: simd::mq3g256v2_row_dot_avx2 (AVX2+FMA+F16C: one header widen, 3-byte chunks via broadcast+vpsrlvd+mask+convert+FMA, 8 contiguous 3-bit codes per 24-bit LE field, no f32 materialization; affine header = 2 FMAs per half). row_dot_enabled() per-format gating (qt 13 AVX2+FMA, qt 49 +F16C; AVX-512 needs own gate — Zen 2 k9lin in host set). Trace now per-shape means (was process-wide mean misattributed per shape: "m=48 5.31 ms" was first-3-steps mean; real m=48 = 0.048 ms scalar / 0.024 AVX2). Measured 27B qt49 BUDGET=56 (8/64 spilled), round-robin {cpu,pcie} x {pre,post} x 3 fresh processes: cpu 3.10 -> 15.20 tok/s (+390%), pcie control 13.20 both binaries (delta is the arm; -76.5% loss -> +15.2% win). Steady state 0.05-0.96 ms GEMV vs 3.4-9.6 scalar; 7.48 s GEMV + 1.42 s copies over 12800/15872 steps; non-uniform mix (16:8:8:8:4:4:1:1) so per-token quoted as ranges. Evidence: 27b amendment + raw data/ kept; crate tests 25/25 dbg+rel (f64 oracle + AVX2-vs-scalar with non-pow2 headers); qt 49 re-parity 5.364e-7/9.537e-7 vs 1e-4, others unchanged; greedy run cpu==pcie 610 B. Docs: env inventory, runbook, design note, CHANGELOG. Raw bench data (bench JSON, traces, drivers, kernel probe) kept byte-identical per tree-parity requirement. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The coverage map (docs/quant-formats/cpu-simd-coverage.md) lists the V2 family as scalar-only, ranked "warpfront#1 highest" for reuse x shipment: qt 44 (Mq4G256V2) is what the mq4v2 bodies and qwen3.8-27b.mq4-xt actually carry, and qt 47/48/50 are the same header over different payload widths. All four share qt 49's fp16 per-128 header, so they were one kernel each away. x86.rs becomes a table of layouts over one decode core instead of a kernel per format: - codes8<BITS> — eight BITS-wide codes as i32 lanes, in element order. An 8-code chunk is exactly BITS bytes, so nibbles, the 2/3/5/6-bit cross-byte packs and (next) BQ1's sign bits are all one broadcast plus one variable shift: <=4-bit in 32-bit lanes (one vpsrlvd for all eight), 5/6-bit in 64-bit lanes with the four low dwords compacted per half. - codes_dot_sum — (sum c*x, sum x) over a run of chunks, so a group pays two horizontal sums, not one per chunk. - f16_quad — the V2 header widened with one vcvtph2ps; v2_group_dot scores each 128-element half and then affines it. - row_dot! — the row driver (sum over groups), monomorphized per format because a #[target_feature] body's features do not propagate into a closure or through a function pointer. load_le reads *exactly* the chunk's bytes (the 3-bit width reuses qt 49's one-preceding-byte 32-bit load, which stays in bounds because a payload always starts after its header). A wider load is one instruction cheaper and runs up to three bytes past the group on the last chunk, which the row slice licenses only for rows that are not the tensor's last. qt 49 moves onto the shared path; its arithmetic is unchanged (same extraction, same per-half dot/sum, same combining order), so its recorded numbers still hold. qt 13 keeps its hand-unrolled 32-codes-per-16-byte nibble group — see the module docs for why. Dispatch: row_dot_enabled(q, req) is now "the features q's kernel needs are present" (features_for: F16C for fp16-header kernels, AVX2+FMA otherwise), and row_dot_avx2 is the single format -> kernel map, returning None for formats not ported yet so gemv's per-row match disappears. Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C): - crate tests 26/26 green, including gemv_matches_the_f64_reference, which now runs the 5 V2 formats through the SIMD path against an independent f64 oracle (AVX2 is active by default on this host). - avx2_and_scalar_agree_within_tolerance now asserts the dispatcher hands the row to a kernel at all: a None would leave gemv on the scalar path, where the two outputs agree by construction and the comparison would pass vacuously. Headers are rewritten non-power-of-two so a mixed-up half or a swapped scale/zero is O(1) relative. - kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best of 3, not a serve-path measurement): qt44 4468 -> 281 us (15.9x), qt47 7025 -> 487 (14.4x), qt48 8863 -> 439 (20.2x), qt50 4132 -> 310 (13.3x), qt49 2520 -> 327 (7.7x), qt13 unchanged 3965 -> 297 (13.4x). Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Tier 1 (qt 15 Mq6G256, qt 31 Mq5G256), Tier 3 (qt 45 Mq4CG256) and the rest of the f32/fp16 single-header family, all of which the coverage map lists as scalar-only. Two observations collapse most of that list into rows of a table: - The kernel only sees bytes, and rotation is applied to the *activation* by the caller (crate::gemv / quant::is_fwht_g256). So an unrotated format and its rotated twin share a kernel: qt 6 has qt 13's geometry, qt 8 -> 15, qt 11 -> 17, qt 9 -> 18, and the G128 blocking formats (qt 7/10/12) are the same shape at 16 chunks per group. - TQ2 (qt 40: (code-1)*d) and BQ1 (qt 41: bit ? +d : -d) are affine once rewritten as scale*d + zero — (d, -d) and (2d, -d) respectively — so they are one uniform_group_dot each with a derived header, not two kernels. x86.rs gains uniform_group_dot::<BITS, PAYLOAD, CHUNKS> (the group body the V2 family already used for one half) and an affine_row_dot! macro that instantiates the header reader plus the row driver together, which is what keeps 15 formats to one macro call each. qt 41 is the first user of codes8::<1>: one byte of sign bits is eight one-bit codes, i.e. the same broadcast-and-shift as every other width. Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C): - crate tests 26/26 green; gemv_matches_the_f64_reference covers all 15 new formats through the SIMD path against an independent f64 oracle. - avx2_and_scalar_agree_within_tolerance extended to every new format, with awkward_headers rewriting the header per family (f32 pair, fp16 pair, fp16 d) so a swapped scale/zero is O(1) relative. - kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best of 3, not a serve-path measurement), speedup vs the scalar decode: qt6 13.6x, qt7 13.9x, qt8 6.6x, qt9 14.0x, qt10 8.8x, qt11 7.7x, qt12 6.9x, qt15 6.6x, qt17 7.2x, qt18 12.8x, qt31 19.8x, qt40 10.7x, qt41 10.2x, qt45 13.0x. No format regresses. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Tier 2 of the coverage map: qt 20 (Mq3G256Lloyd, what the registry's -mq3 tags actually carry), qt 19 (Mq2G256Lloyd) and qt 30 (Mq4G256Lloyd), plus qt 51 — byte-identical to qt 19 and unrotated by design, which is a caller-side difference, not a kernel one, so it is a second row_dot! over the same body. The Lloyd tier has no affine header: the group's CB fp16 entries *are* the decode and the payload indexes them, so codebook_group_dot needs one accumulator instead of two and no scale/zero. The lookup is vpermd, not vgatherdps: a 4- or 8-entry book is one permute, and a 16-entry book is two plus a blend on the index's top bit — which works precisely because vpermd reads only an index's low three bits, so both halves' lookups are valid and the blend selects. For a table this small the gather is the slower instruction. Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C): - crate tests 26/26 green; gemv_matches_the_f64_reference covers all four formats through the SIMD path against an independent f64 oracle. - The tolerance test's fixture now rewrites the *codebook* for this tier, not a header: the shipped fixture books are powers of two and qt 30's repeats an eight-entry book into sixteen lanes, which would hide an off-by-eight table error. write_cb installs entries with alternating signs and a distinct magnitude per lane across four exponent classes. - Mutation-checked that the new coverage bites: dropping the 16-entry blend fails qt 30 at rel 7.08e-1, and moving qt 20's payload offset by one chunk width fails at rel 3.50e+1. Both revert clean. - kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best of 3, not a serve-path measurement): qt19 21.6x, qt20 8.6x, qt30 16.5x, qt51 21.1x over the scalar decode. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The last four formats without a kernel (qt 1 F16, qt 2 F32, qt 16 Bf16, qt 3 Q8F16) close the table, so row_dot_avx2 no longer has a fallback arm and row_dot_enabled is now purely a CPU-feature question. Both matches are exhaustive: adding a CpuQuant variant fails to compile until its kernel and its gate are stated, which is what keeps the two tables from drifting apart. F16/Bf16/F32 are element formats — no codes, no header, just weights to widen — so dense_row_dot! runs them: 256-element groups, eight elements per lane group, and the only per-format variable is the element width. Bf16 is the one kernel in the crate that needs no F16C (a 16-bit shift is the whole conversion), and F32 needs no conversion at all. qt 3 Q8F16 is the last hand-written group: a 32-element block of one fp16 scale then 32 signed bytes, decoded with vpmovsxbd. docs/quant-formats/cpu-simd-coverage.md arrives with this change (it was the plan this series executed, and lived only in a sibling checkout): the per-format kernel table with its ISA gate, how the kernels are organised, what is still not covered (no CpuQuant variant, non-x86_64, no NEON), and the verification story. Indexed in docs/INDEX.md; the CHANGELOG entry covers the whole series (and the qt-49 entry's now-renamed kernel symbol is corrected there). Verification (host: Zen 4 7800X3D, AVX2+FMA+F16C; GPU gfx1201 / RX 9070 XT): - crate tests 26/26 green. every_cpu_decodable_format_has_a_kernel now takes its format list from from_quant_type over the whole u8 space, so it fails if the kernel table and the qt map ever disagree. - gpu_gemv_parity (pre-existing, #[ignore]d, real fixtures + GPU): the production GPU launcher against this crate's GEMV with the vector path active, driving the real cpu_exec seam with a forward-rotated activation. Real tensors of qwen3.5-2b.mq4/mq6/mq3 and qwen3.5-9b.mq4 — qt 13/15/20, with AWQ and prerotated arms (18 checks) — plus synthetic buffers for the rest: 23 formats in its per-format summary. Worst relative error 9.3e-7 (Mq4G256, the +awq real-tensor arm at m=6144 k=2048); then Mq5G256V2 2.9e-7, Hfq6G256 2.3e-7, Mq6G256 2.3e-7, Mq3G256Lloyd 2.2e-7, Mq6G256V2 1.9e-7; then Mq4G256V2 8.9e-8 and Mq3G256V2 7.2e-8; the remaining 15 of the 23 report exactly 0. Two honest gaps: qt 11/12 report SKIP (no production dense GEMV on gfx1201, so the comparison could not run for them), and F16/F32/Bf16 are absent from the test's synthesised table, so those three rest on the crate tests alone. - The tolerance test's fixture rewrites the *elements* (and, for Q8F16, the block scale) for these formats: the shipped fixture values are small powers of two, which would make every product exact. With them awkward, the measured worst case across every format and shape is 2.3e-5 relative, so the bound there is now 1e-4 — documented as f32 accumulation noise rather than a decode bound, against the 7.1e-1 / 3.5e+1 a real decode error produces. - Mutation-checked: dropping the bf16 shift fails at rel 1.0e0, moving qt 3's payload offset fails at rel 2.0e+1. Both revert clean. - kernel-level microbench (m=4096 k=5120, same rayon pool both arms, best of 3, not a serve-path measurement): qt1 20.2x, qt2 13.2x, qt16 15.4x, qt3 7.9x. Nothing in the table regresses; the scalar decode stays the ARM/non-AVX2 path and the test reference. No serve/bench run: per AGENTS.md the parity route is the claim-scoped gate for a kernel change, and this makes no model-level perf claim. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Two defects in the Settings surface the layer-offload work added
(`feat(tui): offload row in Settings easy list`), both on the
`GPU layers` row (`memory.gpu_layer_budget`) and both invisible to its
guards — those assert that a row has a spec, an explainer and draws, not
what a value means.
1. Clearing the row stored the literal string "null". `writer::write_value`
returns `Value::Null` for the documented clear spelling (`null`), and
`persist_setting` reflected it as `Value::Null.to_string()` — the string
"null". The row's label then read "null on GPU" (its unset test is
`is_empty()`), `easy_override_state` claimed a change for a key that
was not in config.toml, and the next edit seeded its buffer from
"null", so typing a count produced "null32" — rejected by the schema,
which made a cleared row unsettable without hand-erasing the buffer.
A cleared value now stores the unset spelling and drops the override
marker (which has always meant "this key is explicitly on disk").
2. Enter on a row the user never typed into was an error. The editor
opens empty exactly when the stored value is null
(`current_setting_value` returns ""), and `parse_cli("")` on a
nullable integer is `could not parse ''`, so Enter answered
`rejected invalid value for gpu_layer_budget: ""` and stayed in edit
mode. An empty buffer on a field that cannot hold the empty string IS
the unset spelling, so `write_value` now retries it as `null` — which
clears, and clearing an already-unset key is a no-op. A `FreeStr`
field accepts "" directly and never reaches that arm, so
`prefill_drafter` still means the empty string.
Also, the same row rendered its reserved auto spelling as "-1 on GPU",
claiming minus one layers were resident; it now reads
"auto (engine decides)", matching the schema's own wording ("defers
placement to the engine").
Verification (host: Zen 4 7800X3D; no GPU needed):
- tui 160 tests + config 89 tests pass.
- Each fix is mutation-checked, i.e. the new tests fail with the fix
reverted: restoring `Value::Null.to_string()` fails
`clearing_the_offload_row_returns_it_to_unset` with
`left: Some("null") right: Some("")`; removing the empty-means-clear
retry fails
`empty_input_clears_a_nullable_field_but_stays_empty_for_a_string` with
`Invalid { key: "gpu_layer_budget", value: "" }`.
- The clear path is driven through the real key handler, not the writer
alone: Enter → type 32 → Enter (renders "32 on GPU"), then erase the
seeded buffer → Enter (renders "all on GPU", no override marker, key
gone from config.toml), then `null` → Enter (same state), then type 8 →
Enter (renders "8 on GPU") — the last step is the one that used to
produce "null8".
- The changelog line here is the clear-semantics fix; the row's own entry
(the `Offload exec` setting) lands with the row's change.
Signed-off-by: Avery Drouillard <avery@averynet.xyz>
… exec` `memory.offload_exec` (compat env `HIPFIRE_OFFLOAD_EXEC`, default `pcie`) decides whether a spilled layer's weight-reading GEMVs run on the GPU over the link or on the CPU. It was reachable from the config file and the env var and from nothing in the TUI: the key existed, the engine read it, and no Settings row offered it — the same gap the layer-offload row closed for placement. This adds the second half of that decision next to the first. - `Offload exec` easy row, right after `GPU layers`, in all four positional lists (`easy_keys`, `easy_override_state`, `easy_help_keys`, `easy_rows`) with a `knobs::KNOBS` explainer and per-option help for both arms. Cells read `over PCIe` / `on CPU`, in the same grammar as their sibling (`all on GPU` / `32 on GPU`). - An enum row, so Left/Right/Space stage a preview and Enter commits — the path the existing `EDITABLE_FIELDS` spec mechanism already handles; no numeric buffer, so none of the empty-buffer semantics apply. - Its option list is `hipfire_config::OFFLOAD_EXECS`, and that same const is now what the schema field validates against, so the two cannot drift. A local copy would have been the class that left `reasoning_effort` uneditable and still leaves the TUI's `thinking_budget` list without `off` (a value the schema accepts, so cycling from it jumps to the wrong option). Exporting the list for a two-value enum is the precedent set by `REASONING_EFFORTS` in the same file. Guards (each exercised, not just asserted): - `every_inline_editable_easy_row_has_a_field_spec` in the config crate now covers the new row by construction (it walks `easy_keys`), and `offload_exec_options_are_the_schema_values` pins the option list, writes every arm through the shared validator, and rejects a value outside it. - `settings_easy_draws_the_offload_exec_row_and_its_explainer` renders the real Settings surface (ratatui TestBackend, both arms plus the explainer) with the values pinned in-test, because `render_with` builds on `App::load()`, which reads the developer's real `~/.hipfire/config.toml`. - `easy_offload_exec_row_cycles_and_commits_every_schema_value` drives the real key handler end to end: cycle stages a preview without writing, Enter commits, the row re-renders, the key lands in config.toml, and cycling wraps. - `offload_exec_row_renders_both_arms` pins both cell values. Verification (host: Zen 4 7800X3D; no GPU needed): - tui 163 tests + config 89 tests pass; `cargo run -p hipfire-tui` starts on a pty and runs without panicking. - Not verified here: a real terminal session driving the row (the render tests are the crate's own surface harness, and the pty run was start-only); and the engine-side effect of flipping to `cpu` (that is the offload path's own measurement, docs/perf-checkpoints/2026-09-27-*). Signed-off-by: Avery Drouillard <avery@averynet.xyz>
… exec spellings
Follow-up corrections to the `Offload exec` row and to the clear semantics
around it — the sentences a reviewer needs stated exactly rather than
approximately:
- `writer.rs`: the `gpu_layer_budget` spec comment claimed "typing `null`
clears the override". The editor cannot do that: its buffer seeds from the
current value, so typing appends ("32" + "null" = "32null", which the
schema rejects) — the end-to-end test in the previous commit had to be
rewritten for exactly this reason. The comment now says what the code
does: an *emptied* buffer is the unset spelling (`write_value` maps empty
input on a nullable field to clear), and `null` is the config/CLI spelling
rather than something the buffer can reach once a value exists.
- `knobs.rs`: the exec explainer now names both spellings the user sees —
"over PCIe" is `memory.offload_exec = pcie`, "on CPU" is `= cpu` — so the
row cannot reintroduce the un-inverted-wording problem the GPU-layers row
was fixed for; and it carries the timing ("read once at load — changing it
takes effect on the next serve/restart"), because the key is process-scoped
and snapshotted at startup. Without it the row would promise an immediate
effect it cannot deliver.
- `writer.rs` test: `empty_input_clears_a_nullable_field_but_stays_empty_for_a_string`
now also pins that the empty-means-clear retry cannot hand a *non-nullable*
enum a clear spelling its schema does not have:
`write_value(path, "offload_exec", "")` must still error.
Scope note, since it is easy to misread the previous commit: the
empty-means-clear retry lives in `hipfire-tui`'s own writer, which has a
single non-test caller (`app.rs` `persist_setting`). The CLI does not go
through it — `hipfire config set` uses `hipfire_config`'s
`ConfigLayer::set_cli` → `field.parse_cli` — so
`hipfire config set <nullable key> ""` still answers `could not parse ''`.
This commit changes no CLI behaviour, and neither does the previous one.
Verification: tui 164 + config 89 pass; the pinned assertion above is
asserted against the shared validator, not assumed.
Signed-off-by: Avery Drouillard <avery@averynet.xyz>
The knob that drives the whole feature had no env-var reference row: docs/env-vars.md gained HIPFIRE_OFFLOAD_EXEC and HIPFIRE_CPU_EXEC_TRACE but not the budget, and the AGENTS.md §7 table gained only the exec row. Adds both (typed-key table + compat-env table + the §7 quick-reference row). The generated-drift check passes because the field!() macro already declared the alias; the human-facing tables were the gap. Also asserts the plan's shipped status and replaces §11's unchecked acceptance list with the evidence of record: the byte-identity and zero-diff criteria are ticked against the fixtures that evidenced them (2B/9B), the 27B tier is marked not-evidenced on MQ4XT (the run of record is qwen3.8-27b.mq3-xt), and the A3B criterion is marked half-evidenced (structural exclusion, no sparse trace compared).
…in the reader seam Three doc comments on the projection reader still said the host path 'registers a VMM arena'. That mechanism was deleted in favour of hipHostMalloc(hipHostMallocMapped), which registers a host pointer the registry frees with hipHostFree. Comment-only.
…pins The mq3-xt amendment cites benchmarks/prompts/gpu_offload_probe.txt (md5 5835c71e471849b4a72e1dc8e39695e7, 59 tokens) as the measured fixture; the 27B numbers are not reproducible without it, and AGENTS.md rule 4 requires canonical bench prompts to live here rather than at a scratch path. Byte-identical to the pinned md5.
…rds and raw captures The records and their raw captures named this machine's layout: the signed-in user in /home/<user>/.hipfire_kernels and a dev-checkout path, and a host mount for the 27B artifact. Normalized to ~/.hipfire_kernels, <repo-root> and <models-dir>/qwen3.8-27b.mq3-xt. No number, digest, size or line count changed; 25 files, 38 lines. The new placeholders are angle-bracket substitutions, so the shell drivers in the data dir stay obviously non-executable as-is.
Four records were referenced but not present, so the links from the committed records, the benchmark handoff and three Rust doc comments did not resolve: the llama.cpp per-spilled-layer baseline (the source of the 27.1 GB/s link rate), its hipfire counterpart, the independent CPU-exec reproduction, and the 2B/9B/27B spill sweep the mq3-xt amendment corrects. Added as historical records, paths normalized like the rest of the family. The dated-reference graph now closes.
The repo's changed-file rustfmt gate flags 27 of this branch's own files, and three of the hunks are new lines rather than historical debt: an ungrouped import and two unwrapped calls in the qwen35 load path, plus the offload module's indentation in hipfire-config. Mechanical, produced with scripts/fmt-changed.sh against the ref the branch was cut from; `cargo check --workspace --examples` still passes and rustfmt --check --config skip_children=true is clean on all 27. The CI job is advisory, but this is what the tracked pre-commit hook would demand of any later commit touching these files.
scripts/verify-bind-thread.sh reports 12 of 94 pub fn in impl Gpu missing the bind_thread contract on this branch, against 8 of 86 on the ref it was cut from: the four new host-mapped entry points were the delta. Resolved per function rather than blanket-annotated. - upload_raw_host / upload_f32_host issue device work (hipHostMalloc through alloc_host_mapped_tensor, then a raw memcpy_htod), so both now bind the calling thread first. - host_located / vmm_handle are pure state reads — the tensor's own ownership tag and the arena registry — and take the gate's documented skip marker with that reason. Remaining: 8 of 94, byte-identical to the pre-existing set on the base revision (last_launched_kernel, deadline_exceeded, prepare_mq4v2_fp8_x, upload_f32, upload_f16_bits, upload_raw, htod_uploads, pool_stats). Untouched: pre-existing debt, and the script is a pre-commit hook rather than part of CI's required gates job. cargo test -p rdna-compute --lib: 251 passed.
…ed placeholders
Three follow-ups to the path sanitization:
1. The A/B drivers were left non-runnable by it: '<repo-root>' and '<models-dir>' are redirection operators in shell, not placeholders. They now derive REPO_ROOT from git and take MODEL from the environment (MODEL=... ${MODEL:?}). 'bash -n' is clean on both.
2. Two statements denied the files' existence: the drivers' own headers said 'not committed', and the base record's reproduction notes said the arm driver and microbenchmark 'were scratch and are not committed'. Both are committed in this record's data directory; the text now says so.
3. '<models-dir>' -> '/path/to/models' across the records and their JSON/err captures: the angle-bracket form is a valid redaction but is also repo house style for unset variables in fenced commands, and the path form reads unambiguously as a placeholder in both prose and payload strings. No digest, size, count or number changed.
…wrappers
`clippy::not_unsafe_ptr_arg_deref` is deny-by-default, so `cargo clippy -p hip-bridge`
failed to compile the crate (3 errors, exit 101) and took the whole workspace run down
with it — hipfire-dispatch, hipfire-arch-qwen35, hipfire-runtime and rdna-compute only
ever reported that cascade, which is why four crates appeared to have errors of their
own. The same command on the pre-change tree is clean (5 warnings, 0 errors), so all
three are new: `host_get_device_pointer`, `host_free`, `mem_get_handle_properties`.
The file already carries this allow, with the same rationale ("HIP treats `device_ptr`
as an opaque GPU address; Rust never dereferences it"), on `mem_get_address_range`, so
these mirror the established pattern rather than turning three public signatures into
`unsafe fn` and rippling that to every caller. Regenerates hip-bridge's map for the
seven added lines.
fivetide
left a comment
There was a problem hiding this comment.
Thanks for tackling this. Capacity offload is something we want, and the measurement discipline here is excellent. The mechanism layer is in good shape: the host-mapped primitive with an ownership registry and cleanup check in rdna-compute, the DeviceBuffer::is_host_mapped tag, the HfqBackend reader seam, and the CPU seam in execute_steps are all in the right places.
Our main concern is architectural. Hipfire's direction is that arch crates should contain as little code as possible, ideally typed declarations, with loading, placement and dispatch handled centrally. In this PR, the decision about which layers spill (and everything that follows from it) lives in hipfire-arch-qwen35. As a result, every new arch would have to copy it, and even qwen35's own PARO and MoE paths already fall through the gaps. Concrete suggestions:
- Placement belongs in the loader, not in
Qwen35Config.model_load::Layoutis already the arch-generic "where does each layer land" object. Give it a per-layerResidency { Device, HostMapped }. Resolve it once frommemory.gpu_layer_budgetin the shared load path, and pass it throughWeightSource::read_layer(gpu, layer, residency)intoHfqBackend. llama and qwen2 already buildHfqBackend, so they would get offload almost for free. For the manifest route (weight_manifest::WeightPlacement/fulfill_manifest, currently piloted on llama), the same tier becomes a field on the placement. A contiguous spilled prefix is really just a pipeline split whose first stage is the host tier (DeviceMesh::stage_for_layeralready bands layers). - One upload path, not a qwen35 twin. Rather than parameterizing qwen35's private quant-type match (whose doc comment already says it should be pulled into runtime), consolidate it into
weight_backend::decode_raw_codec/RAW_CODECSand give that function a residency/target parameter. Change theread_projsignature to take&mut Gpu. The "&Gpu-only fn-pointer contract" is our own choice and can change, which removes the need forread_proj_hostas a second entry point. - Downstream decisions should read residency from loaded weights, not config. Dispatch already asks
gpu.host_located(buf), which is the right shape. Record the execution target (GPU over PCIe vs CPU) on the weight at load time. Then graph eligibility, the Redline refusal, admission, and device-side derived caches (e.g.ensure_fp16_shadow) can all ask theGpu/weights "do you own host-executed weights?" instead ofconfig.i_gpu_startplus a process-globalLazyLock. - Gate capture centrally. Put the refusal in
GraphState::begin_graph_capture*and the replay controller'sbegin_capture, instead of two qwen35 call sites (forward.rs:1694,forward_slots.rs:2518). MTP proposal, DFlash draft, dense-TP and verify graphs aren't covered today, andmtp_spec.rs/dflash.rscallhip.stream_begin_capturedirectly, which bypasses thecapture_modeguard inmemcpy_dtoh_auto. The harness runs used--speculation off, so spec decode +offload_exec=cpuis currently unvalidated. The Redline refusal should live inhipfire-loader, where the replay backend is chosen, not in qwen35arch.rs. - Admission from what the loader actually allocated. Have the loader report device bytes and pinned-host bytes. Feed
ModelFootprintfrom that, and charge pinned host memory to the existing host tier inAdmissionController(unreclaimablehipHostMallocis exactly what that tier guards against — the PR notes 16/64 layers stalling with 0 GB free). That removes the tensor-name parsing (tensor_layer_index) inserve_engine. - Fail closed on sources that can't spill yet.
i_gpu_startis set from config regardless of source, butParoSource::read_layernever sees it andload_moe_ffnplaces experts in VRAM unconditionally. Those loads log "N offloaded" whileresident_weight_bytessubtracts the full layer bytes (experts included) from the admission charge, so a 16 GB card can be over-admitted into OOM. Until residency flows through PARO, MoE experts and VL,gpu_layer_budgetshould fail the load for them. Same for UMA devices (e.g. Strix Halo), where spilling to host frees nothing. - One CPU decoder.
hipfire-cpuis at layer 3 andhipfire-runtimeat layer 5, so runtime can depend on it. Makehipfire-cputhe canonical CPU dequant and haveweight_backend::dequantize_to_f32call it; the cross-check test and the transcription then go away. Ideally the CPU decoder becomes part of the codec registry entry, so a new format is one row andcpu_quant_fordisappears. - Use Steps for the fused FFN. Instead of hand-splitting
weight_gemv_swiglu_residualinllama.rs, express SiLU +GemvResidualas Steps so the generic seam handles it. That's also where prefill/batched coverage would eventually plug in. - Trim scope so this can be reviewed. Please split out: the
gemv_mq4g256kernarg fix (a real bug, its own PR), therdna-computetest lock and tolerance changes, the TUI rows (~640 LOC), and the ~30 perf-checkpoint data files.largest_fitting_tail/resolve_i_gpu_startandGpuLayerBudget::Autoaren't used in production yet: drop them, or land them in runtime next toadmission.rswhenautoactually decides something (placement arithmetic doesn't belong inhipfire-config, which is config vocabulary only).
Suggested order: (2) → (1) → (3/4/5) → CPU execution. After step (1) the RAM spill works for every arch that uses HfqBackend, and your qwen35 numbers should reproduce unchanged.
GpuLayerBudget::Auto, largest_fitting_tail and resolve_i_gpu_start had no production caller: the qwen35 policy matched the enum but never called the arithmetic (it carried its own saturating_sub split). 'auto' and '-1' stay accepted and now mean fully resident, which is what the engine did with them anyway. The qwen35 Auto arms become unreachable and go with the variant.
HfqBackend carried two projection readers (device + host) behind a host_local flag, and qwen35 carried its own copy of the reader with a private quant-type match that had drifted six formats ahead of the runtime's codec registry (qt 31/32/33/34/36/37 — MQ5G256 and the MFP4/MFP3/MFP2 E8+Lloyd family). Now: RAW_CODECS covers those six, DType::requires_k_mod_256 carries their K constraint, expected_payload_bytes replaces six hand-copied length checks, and one reader (hfq::load_weight_tensor + decode_weight_bytes) takes a Residency instead of a second entry point. The host-decode fallback (qt 1/2/16) got a host upload, so an offloaded layer can no longer silently land on the device. qwen2's own reader is a strict subset of the shared one and went with it. The kernel.lm_head_f16 policy moved to hipfire-config (its only consumer, qwen35's private match, is gone) and now applies at the LM-head read only: a qt-1 projection is storage, not policy, and llama's fixture — every projection qt 1 — is what proved it.
qwen35 derived its own i_gpu_start from memory.gpu_layer_budget at config construction, so only qwen35 could spill and the decision had no relationship to what the load actually did. Placement is now one decision in one place: Layout::spill_count(n_layers, requested) -> spilled, applied by Layout::resolve_residency, read back per layer by WeightSource::read_layer. Every arch holding an HfqBackend therefore inherits a spill (llama, qwen35, and qwen2's backend), the 'partial offload: N resident / M offloaded' line is emitted by the loader for all of them with the same wording, and qwen35's Qwen35Config::i_gpu_start, apply_offload_policy, offload_split and residency_report are gone. memory.gpu_layer_budget is now Option<usize> (the requested resident count) with the enum deleted; unset/empty/auto/-1 all resolve to 'no budget'. The TUI's -1 row says 'all on GPU' rather than 'auto (engine decides)', which is what -1 has always meant. Fail closed at the same seam: a source that cannot honour a spill refuses the load before allocating (PaRoQuant, MoE preserved-experts) and so does a unified-memory device, where spilling frees nothing. Verified: 9B resident output byte-identical to master (greedy) and a 24/32 pcie spill byte-identical to resident.
…e refuses a spill qwen2 drove its own embed -> norm -> lm_head -> layer loop, so it was the one HFQ arch that could not spill and the one with no whole-model rollback from the shared transaction. It now has a WeightSource and goes through model_load::load_weights like llama and qwen35, with the per-layer free the transaction's rollback needs; its progress lines and load order are unchanged. The manifest pilot's executor has exactly one upload (pooled device memory), so WeightPlacement gains the residency tier the review asked for, plan_manifest populates it from the shared split arithmetic, and fulfill_manifest_single REFUSES an entry the tier calls host-mapped rather than allocating it in the VRAM the spill exists to free. The legacy load_weights_hfq route is what spills llama today; this pilot refuses until it grows a host upload.
… what the loader allocated Three decisions that used to be re-derived per site are now recorded once: * ExecTarget on WeightTensor/WeightRef. The loader sets it from the resolved residency plus memory.offload_exec; the dispatch CPU seam tests the recorded decision instead of re-deriving it from where the bytes happen to live. That distinction is real: host-mapped + pcie is a GPU-read-over-the-link mode, not a CPU case. cpu_exec_enabled() and its process-wide LazyLock are gone; dispatch asks its own DispatchCtx, which reads the flag off the Gpu. * GraphState::cpu_exec_weights, set once per load. All four capture entries plus Gpu::begin_stream_capture and ReplayController::begin_capture refuse while it is set, so a model with a CPU-executed step cannot be captured or replayed by any route - including the MTP-proposal and draft-FFN wrappers, which used to call hip.stream_begin_capture directly and bypass the guard. Replay's reset_for_model deliberately preserves it: the daemon calls that AFTER a load, so clearing it there would blind the gate for the model that just set it. * LoadStats, measured from the tensors the source produced (WeightTensor defines owned_bytes next to free_all, and qwen35 mirrors free_moe_ffn branch for branch), replacing both the file-size estimate and the tensor-name parsing in serve_engine. Admission charges the pinned spill to the host tier up front, and serve preflight now runs on the loader's own numbers after the load. The Redline refusal moved here from hipfire-dispatch: it needs the resolved spill and the replay backend, and both live in the shared loader, so all four load entries are covered by one check rather than four. Verified on gfx1201: resident and 24/32 pcie-spill outputs still byte-identical to the master baseline (greedy), and HIPFIRE_GRAPH=1 + offload_exec=cpu emits exactly one 'hipGraph capture disabled' line with no assertion and coherent text.
weight_gemv_swiglu_residual's CPU arm called run_host_mapped_gemv_residual directly, so the one op family that fuses silu+rotate+gemv kept its own private convention for the seam. It now builds a GemvResidual step and calls execute_steps, so the CPU decision, the rotation disposition and the launch all come from dispatch and the recorded WeightRef.exec. The GPU arms are deliberately untouched. Phase 5 as originally scoped replaced them with two generic steps, which would have deleted the fused silu_mul_rotate_mq / _awq arm that every resident MQ down-projection takes — FUSED_TABLE carries no silu+gemv_residual key, so the pipeline does not re-fuse it — and that arm also owns the AWQ divide-then-rotate ordering. That is the hot path V1 certifies byte-identical, so the change is reduced to the arm the review was actually about. Verified: resident output still byte-identical to the master baseline, and the cpu-execized output is identical to the pre-change run.
…_weight LayerWeights::owned_bytes claims to mirror free_gpu branch for branch, and its dense DeltaNet arm listed only attn_norm/ffn_norm/a_log/dt_bias — conv_weight and norm_weight are released by the teardown and were never charged. Small in bytes, but it is a silent under-count in the accounting admission reads, which is the failure mode the function exists to prevent. The helper now takes slices instead of a fixed 4-element array, so an arm cannot drop a tensor by arity again, and the MoE arm's comment (which claimed the dense arm carried the two tensors separately) is corrected. Cross-checked as a difference: with 8 LinearAttention layers spilled, the fix moves host_pinned_bytes by exactly (98_176 + 512) x 8 = 789_504 B, the conv/norm bytes of those layers.
…is rework
hipfire_cpu::quant::dequant_group is the canonical CPU decoder — it is what the
CPU-exec seam runs — and weight_backend::dequantize_to_f32 now delegates for every
quant_type it covers, keeping its inline arms only for the formats the crate does
not name (so the FWHT un-rotation for those is untouched). With one implementation
the cross-check that held the two transcriptions together is tautological and is
deleted with its only consumer, dequantize_weight_to_f32.
That delegation is also a behaviour fix, found by the check itself: the canonical
decoder PANICKED on qt 49 (MQ3G256V2) while hipfire_cpu decodes it, so
dequant_f32 used to abort on a format the seam handles fine.
Evidence, in order: the cross-check was green before any change (1879 tensors,
qts {1,3,8,13,15,20}); with the delegation disabled and the fixture set widened it
compared the old inline arms against hipfire_cpu bit-for-bit over 3130 tensors
including qt 44 (qt 47 is covered only by the delegated run — its fixtures also
carry the qt-49 tensor that panics the old decoder). Then the delegation went in
and the file went away.
Qwen35Config::spill_refusal carries the MoE refusal so it is testable without an
HFQ fixture, and load_weights_inner prints the measured LoadStats once under
HIPFIRE_OFFLOAD_DEBUG=1.
docs/plans/partial-gpu-offload-design.md gains a post-review revision note and
inline retirements, because it is the design record for APIs this rework deletes.
Mechanical only: the repo formats the files a branch changes (scripts/ci-rustfmt-changed.sh), and the exec-field sweep left ~40 insertions at the wrong indent along with some re-wraps. Split from the semantic commits so a reviewer can skip it.
…md5s The battery+chain checkbox was ticked against the pre-rework implementation, while the capture gate, the recorded exec target and the loader-side refusals all changed since. Re-ran both modes plus the redline pcie-spill arm against the current code and recorded the numbers with the model/hipfire/daemon md5s the reporting rule requires (the local 9B is 296092bf…, not the pinned fixture).
…references `dequantize_to_f32` delegates here now; the deleted `cpu_quant_cross_check` test no longer exists to keep two copies honest.
…offload fields The Settings rows for `memory.gpu_layer_budget` / `memory.offload_exec` stay, but the writer no longer changes behaviour for any other setting: - `write_value` keeps master's rule (an empty buffer is invalid); the generic empty-means-unset retry and the `IntOrUnset` field kind it needed are gone. - Clearing a set budget is the row's Delete/Backspace (`reset_selected_setting` -> `writer::delete_key`), the same path every other row uses, so `persist_setting`'s `Null`/override bookkeeping reverts. The only removals left in `crates/hipfire-tui` are the generated crate-map rows and five lines of pre-existing rustfmt debt that the changed-file fmt gate fixes. Tests: the two tests that encoded the empty-buffer clear now pin parity — an empty buffer is invalid for the offload row exactly as it is for `max_tokens`, `temperature`, `mmq_screen` and the pre-existing nullable `deepseek4_experts_per_token` — plus Delete-clears and the `null` spelling.
…ling as unset With `persist_setting` back on master's body, a write of the literal `null` (the spelling `hipfire config set` uses to clear this nullable key) leaves that string in `values` until the next reload, so the row read "null on GPU" even though the key is off disk. Handled in the new row's own label (`"" | "null"` -> "all on GPU") rather than by special-casing the shared write path, so no mainline line changes.
The arch rework moved six formats out of qwen35's private quant match and into RAW_CODECS (qt 31/32/33/34/36/37), but the register still declared them `arch-loaded`, which its own rule forbids for a format with a table row: `scripts/check-quant-registry.py` reported qt_disposition_mismatch = 6 and `scripts/leanup-ratchets.sh` failed on it. qt 35 (MFP4G32E8SOA) is untouched — it has no RAW_CODECS row and qwen35 still owns its upload. Also refreshes the generated crate-map blocks for the 15 crates whose counts drifted across the rework, so `scripts/check-crate-maps.py --check` matches the tree again.
…pu-offload # Conflicts: # registry/v1.json
…tream Merging upstream/master (66e825b "registry: add qwen3.8:27b-mq4-xts (H2) and re-pin qwen3.8:27b-mq4-xt on master") reproduced the gate failure this PR kept inheriting: that commit grew crates/hipfire-registry/src/lib.rs by 22 lines (2,197 -> 2,217) without regenerating crates/hipfire-registry/map.md, so master's own gates job fails on 'hipfire-registry: generated counts/content are stale' — and any PR that merges master inherits it. That is why the failure was unreproducible until the upstream merge was done here. Also in this merge: registry/v1.json conflicted between the fork's registry-bot commit (2f37c43, pulled in by our earlier fork-master merge) and upstream's newer registry commit; resolved to upstream's copy. This branch never edited that file. Verified on the merged tree: check-crate-maps 44/44 matching, leanup-ratchets OK (21 metrics, 0 violations), ratchet-diff vs upstream/master OK.
1. SummaryI had an agent go over the critiques this morning. After a fair amount of work, its results are below. It removed a few of the commits that were requested to be removed, I will submit bugfix PRs for those later. Including the TUI sanitation. I had the agent restore the existing behavior while still adding in the TUI fields. As a side effect of the requests you asked for, qwen2 dense models should work with CPU offload. This has not been tested. There are no qwen2 models in the hipfire registry. 2. The review, answered
The umbrella concern — arch crates holding as little as possible — is where the diff goes: 3. Evidence, all of it from this machine (gfx1201, ROCm 7.2)
4. Replication against the pinned 2B, using the record's recipe
The pcie-vs-cpu ratio reproduces (1.34–1.39 against 1.32–1.40), and the record's observable The Branch binaries for every number above: |
…onto land/040) Squash of PR #793 (warpfront/hipfire, head 307009e, 31 commits by Avery Drouillard) onto land/040-fixes 9eb9e1d. memory.gpu_layer_budget / HIPFIRE_GPU_LAYER_BUDGET spills a prefix of dense Qwen3.5 layers to hipHostMalloc-mapped host RAM; memory.offload_exec=cpu / HIPFIRE_OFFLOAD_EXEC runs those layers' weight GEMVs on the new hipfire-cpu crate. Unset keeps every layer resident. Rebase resolution (fold/serve): - CHANGELOG: the PR's entries are replaced by one consolidated entry in the follow-up commit. - crate maps: generated blocks regenerated (no hand-written map edits in the PR). - load.rs: land's qt=52 (MQ4-G256 v2 Lloyd) arm uploads through the injected uploader like every other raw arm, so a spilled Lloyd layer lands in host RAM. - prefill.rs: land's widened_test_config literal gains i_gpu_start: 0. - Cargo.lock: hipfire-cpu at the workspace version 0.4.0.
… retained Redline default for CPU-executed offload Follow-up to the #793 squash for 0.4.0 (serve audit O3). - hipfire-loader admit_source refuses a memory.gpu_layer_budget that would spill layers under tp>1 (dense-TP and EP rank loaders keep every layer resident), pp>1 (the first band would spill into host RAM mapped by device 0) or on a MoE model (the expert stacks stay in VRAM), before any teardown. An unset budget skips the check and the extra config parse. - HfqSource::prepare backstops the same refusal (MoE, n_devices>1) for every other qwen35 load path and prints the residency line there, once the placement is applied; config parsing no longer prints it. - retained_redline_default: a qwen3_5 process configured with offload_exec=cpu and a layer budget gets no retained default. The tape records GPU launches only and would skip the host GEMVs. - append_betaalpha_to_z keeps the folded Z|beta|alpha rows of a spilled layer in host RAM (upload_raw_host + free_tensor) instead of moving them to VRAM and passing a host-mapped owner to release_tensor_immediate. - CHANGELOG: one consolidated #793 entry in a new "folded PRs" subsection. Co-authored-by: Avery Drouillard <avery@averynet.xyz>
…XPERT_VRAM_LAYERS) + routing trace Experiment branch for Flash-Next on gfx1201 (32 GB). Trunk layers 0..N keep their routed experts in VRAM; later layers' and the MTP layer's experts are fulfilled into pinned device-mapped host RAM (hipHostMalloc Mapped, ported from #793's hip-bridge/rdna-compute owner) and read zero-copy over PCIe by the unchanged sealed MoE kernels through their per-layer pointer tables. New WeightResidency::HostMapped; HIPFIRE_QWEN4_ROUTE_TRACE dumps per-layer routed expert ids for cache sizing.
# Conflicts: # crates/hip-bridge/map.md # crates/hipfire-arch-qwen35/map.md # crates/hipfire-config/map.md # crates/hipfire-dispatch/map.md # crates/hipfire-generate/map.md # crates/hipfire-loader/map.md # crates/hipfire-runtime/map.md # crates/rdna-compute/map.md
…t serve, opt-in) Conflicts vs fold/serve (#793 offload, #787): Cargo.toml features (sha2 non-optional per SCS; keep #793's lab feature), serve_engine planned/weights bytes (SCS paged KV formula + #793 resident_weight_bytes), config.rs tests (both), hfq.rs (keep #787 probe_arch_id), cli local_model_paths (#787 registry arg + SCS catalog tags). vs fold/kernels: attention.rs (both helpers). Semantic: #793's offload guard in forward_batch_slots_graphed_opts now falls back through SCS's forward_batch_slots_opts (kv tier, lm_head_skip, spec_capture). CHANGELOG entry for the stack added.
User decision: hipfire quantizes from BF16 sources; GGUF -> mq4 double-quantizes, and GGUF users should use llama.cpp. Reverts 02e60cd (the only #785 commit in fold/serve); CHANGELOG entry removed and the folded-PRs subsection retitled; hipfire-quantize map regenerated. The rest of fold/serve (#793, #787, #595) stays.
…outed experts on discrete cards, auto expert VRAM layers, host-RAM preflight) Conflicts: hip-bridge ffi.rs/lib.rs and rdna-compute dispatch.rs. The branch adds the same hipHostMalloc(Mapped) owner class (HIP_HOST_MALLOC_MAPPED, host_malloc/host_get_device_pointer/host_free, DeviceBuffer::HostMapped, Gpu::host_mapped, the free_tensor/ensure_vmm_cleaned arms, HOST_TAIL_PAD_BYTES) that land already carries from #793's partial offload; kept land's copies. The auto-merged duplicates (a second HIP_HOST_MALLOC_MAPPED, release_registered_host_mapped, host_mapped_count) are dropped. `Gpu::upload_raw_host_mapped` (the Qwen4 expert upload) keeps the branch's behaviour, a CPU copy through the host pointer plus a zeroed tail pad, on land's `alloc_host_mapped_tensor`. Maps regenerated (hipfire-arch-qwen4 gains expert_residency.rs, loader, runtime, rdna-compute); env-vars inventory regenerated (HIPFIRE_QWEN4_EXPERT_VRAM_LAYERS, HIPFIRE_QWEN4_ROUTE_TRACE).
… of config One home for "where do this model's weights live", so a second arch adopts a placement policy instead of copying one (PR warpfront#793 review points 1/3/5/6/9). `hipfire_runtime::offload`: * `Placement` — per-layer `WeightResidency` for a layer's own weights and `ExpertResidency` for its routed experts, kept as distinct types so the two axes cannot be swapped silently. * `LayerBytes` — the arch's byte census (non-expert / expert / always-resident), plus `device_bytes` and `host_expert_bytes`. * `plan(layers, budgets, capacity, orientation)` — llama.cpp's tier order: experts first, whole layers only when the expert tier cannot free enough, stop at the first fit. `Orientation` expresses both the dense reading (the last N layers resident) and Flash-Next's (the first N), so neither arch forks the search. An explicit `Layers(k)` pin is a hard constraint: a host-placed layer would take its pinned experts with it, so the layer tier is capped by it. * `Capacity` / `KvReserve` / `kv_reserve` (checked multiply). * Host admission lifted from `hipfire-arch-qwen4::expert_residency` with its unit tests: `check_host_ram` (MemAvailable + the TTM page-pool estimate), `gtt_budget` / `check_gtt_cap` (TTM pages_limit), `ttm_pool_estimate{,_from}`, `admit_host_placement`. No fixed host budget anywhere. * `report` — the residency line, including the per-layer routed-expert MiB an operator needs to choose a count. * `largest_fitting_tail`, moved out of `hipfire-config`: placement arithmetic is not configuration vocabulary, and it needs a measurement config cannot take. `OffloadBudget` (renamed from `GpuLayerBudget`) stays in `hipfire_config::memory` — config has no intra-workspace dependencies, so it owns the vocabulary while `offload` owns the arithmetic. One enum now covers both knobs. Tests: `cargo test -p hipfire-runtime --lib offload::` 16 passed; `-p hipfire-config`, `-p hipfire-arch-qwen4`, `-p hipfire-loader` green; the whole workspace still checks. MoE-offload plan, stage 1 step 4 (first half; the qwen4 port follows).
Author's note
This is a rather large change, so I am going to include a small preface here explaining the goal and purpose of this PR.
First off, like a lot of people, my PC is limited to 16 GB of VRAM. I have an AMD RX 9070 XT (gfx1201), not a datacenter class GPU with more than that like a 9070 PRO, or even a 3090, 4090, or 5090 from Nvidia (which are out of scope for hipfire but are relevant for the discussion from a VRAM standpoint). Because of that, I am limited in what I can run and fit on this PC. Dense models seam to be releasing in a few different size stages - 2b/4b models, 9b/12 models, and then at least for home users - larger 27b/30b models. The problem is that the larger 27b/30b dense models are at the capacity limit of a single 16GB GPU - even before you factor in KV cache allocation.
And unfortunately, the most average large GPUs tend to max out at 16GB of VRAM. Yes I know the 20GB 7900XT and 24GB 7900XTX exist. I don't have one. And other than the 9070 PRO, the max for the current generation of AMD GPUs is 16GB of VRAM.
If someone wants to run these models on a 16GB GPU with hipfire this leaves with two options.
Or option 3 which is use a different project like llama.cpp - but when models fit hipfire is significantly faster.
The goal of this PR is to add the ability for Hipfire to offload dense qwen3.5 based model layers - at the time of writing qwen3.5/3.6 MoE and qwen flash next are NOT supported - to system RAM. This enables two things:
Once this was done it was found that there was a signficant speed penalty for doing so. Much more than llama.cpp (using it as a reference performance target). An analysis of the reason why determined that was because llama.cpp uses it's CPU inference engine on it's CPU offloaded layers to avoid passing data back to the GPU. As it turns out, despite CPUs being much slower at inference it is faster to have a CPU do inference on locally available data rather than having the GPU read objects in system RAM over the PCIe link. As a result I had a subsequent agent add a basic SIMD AVX2 accelerated code path to hipfire to use on offloaded layers. I have tested a few of the CPU SIMD paths, but not all and welcome anyone doing so before this merge request is committed.
All that in mind there are a few design decisions I made.
a) By default hipfire defaults to it's 'old' behavior and does not attempt to offload layers to system RAM.
b) Even if you offload layers to system RAM, the CPU inference engine is disabled by default. With the CPU inference engine disabled the system passes all of it's data back to the GPU for inference - on my machine with a PCIe 4.0 x16 link the penalty was significant.
The code should be extensible to other dense LLM architectures.
There is one caveat that is worth mentioning though. CPU inference is not byte identical to GPU inference. That is not to say it cannot be, but it isn't. This was a deliberate design decision - but it should be noted that llama.cpp treats CPU and GPU inference the same way. CPU inference on llama.cpp is not byte identical to its CUDA/HIP/Vulkan, etc runtimes.
Summary
Adds opt-in partial GPU offload for the dense qwen35 family:
memory.gpu_layer_budgetspills a contiguous prefix of layers to host RAM, so amodel — or a context length — that does not otherwise fit in VRAM can run; and
memory.offload_exec=cpuexecutes those spilled layers' weight-reading GEMVs onthe CPU instead of reading them over PCIe once per token. Both default to today's
behavior (every layer resident, weights read by the GPU), so an unconfigured run
is byte-identical to a resident one.
memory.gpu_layer_budgetHIPFIRE_GPU_LAYER_BUDGETNkeeps the lastNlayers on the GPU and spills the rest;auto/-1defers to the engine (today: fully resident)memory.offload_execHIPFIRE_OFFLOAD_EXECpciepcie= the existing GPU kernels read the host-mapped weights over the link,cpu= execute them on the CPUHIPFIRE_CPU_EXEC_TRACE=1prints one line per distinct CPU step shapeMotivation is capacity, not speed: offloaded layers are bandwidth-bound by design.
Unset, empty and unknown values fail closed (
pcie, fully resident), so a typo canneither move work to the CPU nor force an offload.
Which surface(s) does this touch?
crates/rdna-compute,crates/hipfire-dispatch,crates/hip-bridgehfq,weight_backend), archload*,hipfire-config, Cargo manifestshipfire-runtimeweight-GEMV seam (llama.rs,weight_backend.rs,hfq.rs); nohipfire-engine/hipfire-generatechangehipfire-arch-qwen35(placement + load path);hipfire-arch-llama,hipfire-arch-qwen2(declare no offload support)crates/hipfire-quantize/ quant formatshipfire-tui(two Settings rows)Design notes
hipHostMalloc(hipHostMallocMapped), not a host-located VMMarena. Measured on gfx1201, holding 4096 MB: a
hipMemCreatehost arena ischarged against the device heap 1:1 (device headroom 15360 → 11264 MB), while
hipHostMalloccosts 0 MB. The VMM host primitive and its locality tag weredeleted so the net-zero mechanism cannot be re-adopted; the invariant is asserted
by
crates/rdna-compute/examples/host_offload_headroom.rs, because byte-identityand coherence tests cannot see it.
dereferencing a
GpuTensor; the device pointer is the host-mapped alias.i_gpu_start): layers[0 .. i_gpu_start)spill, the contiguous tail stays resident, and
token_embd/output_norm/lm_headare always resident. The policy arithmetic
(
hipfire_config::memory::largest_fitting_tail,resolve_i_gpu_start) is pure,GPU-free and unit-tested.
autocurrently resolves to fully resident — thearithmetic is landed, the device measurement that would drive it is not.
hostflag.HfqBackend::read_proj_hostis
Option<fn(&HfqFile, &mut Gpu, …)>because the host upload needs&mut Gpuand the existing
read_projfn-pointer contract cannot hand one out. A spilledlayer whose arch supplies no host reader is a hard error, never a silent device
allocation.
qt 1
Nativemode have no host upload, so they refuse for offloaded layers ratherthan quietly consuming the VRAM the offload exists to free.
would cost PCIe bandwidth per GEMV to save kilobytes.
pciespill (host-mapped addresses arestable for the model's lifetime, replay verified byte-identical). It is disabled
for the model's lifetime only when CPU execution is active, because a CPU step is
a host sync point.
hipfire-dispatch'sexecute_steps: a fusedentry spanning a CPU step is not matched, and
weight_gemv_swiglu_residualsplitsitself (GPU
silu_mul_f32, then the CPU GEMV + residual). Attention, softmax,rmsnorm, RoPE, qk-norm, the KV write, the flash-attention families and the
DeltaNet recurrence stay on the GPU, and KV residency is unchanged either way.
memory.offload_exec=cpuwith aRedline backend fails the load naming both keys, because the tape records GPU
launches and would replay stale activations.
hipfire-cpu(depends onrayononly, GPU-free tests). Ittranscribes the canonical decoder;
crates/hipfire-runtime/tests/cpu_quant_cross_check.rsholds the two copies together bit-for-bit. All 28 dense-capable formats decode,
each with an AVX2 row dot and a scalar fallback (all-scalar on ARM). Map:
docs/quant-formats/cpu-simd-coverage.md.Numbers
Behind/after are the same fixtures with the spill dialed in; resident control
included. Arms interleaved, fresh process per point, host otherwise idle.
pciecpuqwen3.5-9b.mq4(--spec off, 128 tok)qwen3.5-9b.mq4qwen3.8-27b.mq3-xt(64 layers)battery.jsonattachment, whose 9B is adifferent artifact): 9B md5
31a8d8dc7603226801b08d8319015602, promptbenchmarks/prompts/humaneval_3_below_zero.txtmd537c5aad9f9efe93b5c47f27256bdf149,daemon md5
6b2f85588aac1abdbd233e415044eb12,hipfiremd57075cd546b1ce61d39a36d59ce149cba.27B md5
80bb9198e6a565fc006b2ae1b7c89eca, probe prompt md55835c71e471849b4a72e1dc8e39695e7,daemon md5
e39c3adb85f608db1439b128471764e8,hipfiremd557a0cf5795a8ce376d274eba78bc733a.Host: 1× RX 9070 XT (
gfx1201, 16304 MB), ROCm/HIP 7.2, Ryzen 7 7800X3D, 28 GB RAM.vram_free_mbis identical across arms at the same budget, and the 27B's ownfootprint (~13070 MB) matches stock — the resident path costs no extra VRAM. An
offloaded=Nlog line is not evidence on its own; this control is.≈41 ms/token on
cpuversus ≈51 ms/token onpcie, against the link's measured27.1 GB/s. Remaining headroom is the kernel, not the design — the same row dot
streams a 195.8 MB host-mapped buffer at 44.5 GB/s over 16 rayon threads.
cpuarm as 3–6× slower; it was measured while aworkspace build was running on the same 16 threads. Recorded as superseded rather
than deleted, because the way it was produced is the reusable lesson.
battery.jsonattachment — the 9B fixtureattached there is a different artifact (md5
296092bf…) from the records' 9B (31a8d8dc…).docs/perf-checkpoints/2026-09-27-*(lifecycle
historical, exploratory, not product claims).Test plan
./scripts/no-gpu-ci.shpasses — red on this host, for a pre-existing reason:its Python phase fails 5 tests in
tests/test_mq4c_repack.py(
mq4c_repackhas no attributemain/HfqmError/parse_hfqm_index). Reproducesidentically in an unmodified checkout of this repository, and this change touches
no file under
tools/ortests/. The Rust phases of the same script are green(
cargo check --workspace --examples,hipfire-cpulib 26 passed,hipfire-config89 passed, env-docs drift clean).
cargo build --releaseclean —cargo build --release --locked -p hipfire-daemon -p hipfire-clifinished, warnings only
cargo test --lib --workspacepasses — every crate's lib target green(
hipfire-runtime691,rdna-compute251,hipfire-dispatch273,hipfire-arch-qwen35199,hipfire-quantize49, …), 0 failedserve_harness.py --mode battery(and--mode chain)myself on hardware, with a real spill on the
cpuarm: 5/5 turns each,runaway=0 empty=0 attractor=0 retrieval_miss=0,avg_decode=26.9 tok/s(chain: prefix-cachehits at 388/717/828/952 tokens), decoded and eyeballed — each turn's
assistant_contentis inthe attachment below. Fixture
qwen3.5-9b.mq4sha256ba83acf5…, md5296092bf…;target/release/daemonmd51c6e1c0a…(the harness reports the same digest it ran). The daemonserve log proves the spill rather than assuming it:
partial offload: 24 resident / 8 offloaded, i_gpu_start=8andcpu exec: 8/8 spilled layers fully covered; uncovered quants: none.Separately,
redline_daemon_harness.pyon thepciearm with a spill is recorded indocs/perf-checkpoints/2026-09-27-gfx1201-cpu-exec-offload.md(stable=True, launches=427,AQL/PM4 shadow
exact=True; on thecpuarm it fails the load with the intended refusal)../scripts/speed-gate.shwithin ±2% of locked baselines — not run; itneeds a 4B fixture, and the resident path this gate protects is covered by the byte-identity
(zero-diff) guard in the records above, while the spill's cost is documented rather than hidden.
scripts/leanup-thresholds.txt: RATCHET-RAISE …local serve_harness battery.json (load / serve / kernel changes)
Command
--speculation off --dflash off --thinking off: no τ in the picture, and a reasoningmodel cannot close a think block inside this budget (the daemon fails that turn closed
and releases no text).
--mode chainwas run with the same flags (cached=388/717/828/952,5/5 turns,
avg_decode=26.9).Fixture identity
qwen3.5-9b.mq4(5,313,750,016 B)ba83acf5bfd5d4e334b0afc26d779734e31623bb7f74e807c3581dfecb3128ad, md5296092bf1e6a45d78c1acf815eb93366target/release/daemon1c6e1c0ae86ba651d446fe1c833a729b(the harness prints this asdaemon_binary_md5; it matches the build under test)target/release/hipfiredc305c1c6de2119ed03a9376b42cfdcbgfx1201, 16304 MB), ROCm/HIP 7.2This is not the
qwen3.5-9b.mq4of the perf-checkpoint records (those pin md531a8d8dc7603226801b08d8319015602at 5,297,456,128 B), so these numbers are aseparate fixture and are not comparable with that record's.
The spill actually engaged (daemon serve log)
Per-turn rows
turns=5 runaway=0 empty=0 attractor=0 retrieval_miss=0 avg_prefill=165.1tok/s avg_decode=26.9tok/sharness
--outJSON (verbatim)[ { "request_id": "chatcmpl-1023366-1", "ctx": 44, "cached": 0, "gen": 343, "finish": "stop", "think_words": 0, "ans_words": 178, "prefill_ms": 969.5, "prefill_tok_s": 45.4, "decode_tok_s": 26.9, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 1.026, "wall_s": 13.713, "attractor": false, "empty": false, "runaway": false, "ans_preview": "```python\ndef merge_sorted(a, b):\n \"\"\"\n Merge two already-sorted lists into one sort", "assistant_content": "```python\ndef merge_sorted(a, b):\n \"\"\"\n Merge two already-sorted lists into one sorted list.\n\n This function assumes that both input lists `a` and `b` are sorted in ascending order.\n It returns a new list containing all elements from both inputs, also sorted in ascending order.\n The implementation uses a two-pointer approach with O(n + m) time complexity.\n\n Parameters:\n a (list): First sorted list of comparable items.\n b (list): Second sorted list of comparable items.\n\n Returns:\n list: A new merged sorted list containing all elements from `a` and `b`.\n\n Example:\n >>> merge_sorted([1, 3, 5], [2, 4, 6])\n [1, 2, 3, 4, 5, 6]\n \"\"\"\n i, j = 0, 0\n merged = []\n\n # Traverse both lists and append the smaller element to the result\n while i < len(a) and j < len(b):\n if a[i] <= b[j]:\n merged.append(a[i])\n i += 1\n else:\n merged.append(b[j])\n j += 1\n\n # Append any remaining elements from list a\n while i < len(a):\n merged.append(a[i])\n i += 1\n\n # Append any remaining elements from list b\n while j < len(b):\n merged.append(b[j])\n j += 1\n\n return merged\n```", "content": "```python\ndef merge_sorted(a, b):\n \"\"\"\n Merge two already-sorted lists into one sorted list.\n\n This function assumes that both input lists `a` and `b` are sorted in ascending order.\n It returns a new list containing all elements from both inputs, also sorted in ascending order.\n The implementation uses a two-pointer approach with O(n + m) time complexity.\n\n Parameters:\n a (list): First sorted list of comparable items.\n b (list): Second sorted list of comparable items.\n\n Returns:\n list: A new merged sorted list containing all elements from `a` and `b`.\n\n Example:\n >>> merge_sorted([1, 3, 5], [2, 4, 6])\n [1, 2, 3, 4, 5, 6]\n \"\"\"\n i, j = 0, 0\n merged = []\n\n # Traverse both lists and append the smaller element to the result\n while i < len(a) and j < len(b):\n if a[i] <= b[j]:\n merged.append(a[i])\n i += 1\n else:\n merged.append(b[j])\n j += 1\n\n # Append any remaining elements from list a\n while i < len(a):\n merged.append(a[i])\n i += 1\n\n # Append any remaining elements from list b\n while j < len(b):\n merged.append(b[j])\n j += 1\n\n return merged\n```", "reasoning_content": "", "tool_calls": [], "request_md5": "b2d411eb58ec2648ebd087ac828a51bb", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f", "expected_substrings": [], "retrieval_missing": [] }, { "request_id": "chatcmpl-1023366-3", "ctx": 55, "cached": 0, "gen": 286, "finish": "stop", "think_words": 0, "ans_words": 147, "prefill_ms": 304.1, "prefill_tok_s": 180.9, "decode_tok_s": 26.9, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.348, "wall_s": 10.946, "attractor": false, "empty": false, "runaway": false, "ans_preview": "To find the total distance traveled, we need to calculate the distance for each segment of", "assistant_content": "To find the total distance traveled, we need to calculate the distance for each segment of the trip separately and then add them together.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first part of the trip.**\n* Speed ($v_1$) = $60$ mph\n* Time ($t_1$) = $2.5$ hours\n$$ d_1 = 60 \\times 2.5 = 150 \\text{ miles} $$\n\n**Step 2: Calculate the distance for the second part of the trip.**\n* Speed ($v_2$) = $40$ mph\n* Time ($t_2$) = $1.5$ hours\n$$ d_2 = 40 \\times 1.5 = 60 \\text{ miles} $$\n\n**Step 3: Add the distances together to find the total distance.**\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 + 60 = 210 \\text{ miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210** miles.", "content": "To find the total distance traveled, we need to calculate the distance for each segment of the trip separately and then add them together.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first part of the trip.**\n* Speed ($v_1$) = $60$ mph\n* Time ($t_1$) = $2.5$ hours\n$$ d_1 = 60 \\times 2.5 = 150 \\text{ miles} $$\n\n**Step 2: Calculate the distance for the second part of the trip.**\n* Speed ($v_2$) = $40$ mph\n* Time ($t_2$) = $1.5$ hours\n$$ d_2 = 40 \\times 1.5 = 60 \\text{ miles} $$\n\n**Step 3: Add the distances together to find the total distance.**\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 + 60 = 210 \\text{ miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210** miles.", "reasoning_content": "", "tool_calls": [], "request_md5": "85b33e9aedabf7a9fa0d1d81cc6f8356", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "640e0fd4f55996cb175a422f0a12cef5", "expected_substrings": [], "retrieval_missing": [] }, { "request_id": "chatcmpl-1023366-5", "ctx": 25, "cached": 0, "gen": 81, "finish": "stop", "think_words": 0, "ans_words": 70, "prefill_ms": 126.1, "prefill_tok_s": 198.2, "decode_tok_s": 27.1, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.163, "wall_s": 3.122, "attractor": false, "empty": false, "runaway": false, "ans_preview": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its or", "assistant_content": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its orbital plane around the Sun. As the planet orbits, this fixed tilt causes different hemispheres to receive varying amounts of direct sunlight and experience changes in day length throughout the year. Consequently, when a hemisphere leans toward the Sun, it experiences summer due to more intense solar radiation, while the opposite hemisphere experiences winter.", "content": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its orbital plane around the Sun. As the planet orbits, this fixed tilt causes different hemispheres to receive varying amounts of direct sunlight and experience changes in day length throughout the year. Consequently, when a hemisphere leans toward the Sun, it experiences summer due to more intense solar radiation, while the opposite hemisphere experiences winter.", "reasoning_content": "", "tool_calls": [], "request_md5": "d34383b66fd5f7af0cfced19d3777983", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8f66b4c97988825bd8e7840aaf44357e", "expected_substrings": [], "retrieval_missing": [] }, { "request_id": "chatcmpl-1023366-7", "ctx": 33, "cached": 0, "gen": 102, "finish": "stop", "think_words": 0, "ans_words": 81, "prefill_ms": 211.1, "prefill_tok_s": 156.3, "decode_tok_s": 26.9, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.249, "wall_s": 4.01, "attractor": false, "empty": false, "runaway": false, "ans_preview": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against t", "assistant_content": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against the jagged rocks below. One stormy night, a strange object tumbled onto the shore, glowing with an eerie, pulsating blue light that defied the darkness. He approached cautiously to investigate and discovered it was not debris, but a small, intricate clockwork device humming with impossible energy. Realizing this invention could change everything, Elias knew his life of solitude had just become the most significant chapter of his story.", "content": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against the jagged rocks below. One stormy night, a strange object tumbled onto the shore, glowing with an eerie, pulsating blue light that defied the darkness. He approached cautiously to investigate and discovered it was not debris, but a small, intricate clockwork device humming with impossible energy. Realizing this invention could change everything, Elias knew his life of solitude had just become the most significant chapter of his story.", "reasoning_content": "", "tool_calls": [], "request_md5": "643a78544be2dab1edcfe5b8e85c3718", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817", "expected_substrings": [], "retrieval_missing": [] }, { "request_id": "chatcmpl-1023366-9", "ctx": 31, "cached": 0, "gen": 82, "finish": "stop", "think_words": 0, "ans_words": 62, "prefill_ms": 126.7, "prefill_tok_s": 244.6, "decode_tok_s": 26.9, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.165, "wall_s": 3.182, "attractor": false, "empty": false, "runaway": false, "ans_preview": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep", "assistant_content": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep functions small, focused on a single responsibility, and easy to test.\n3. Write clear comments only for complex logic or non-obvious design decisions.\n4. Follow consistent coding conventions and formatting rules across the entire project.\n5. Refactor code regularly to remove technical debt and improve readability over time.", "content": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep functions small, focused on a single responsibility, and easy to test.\n3. Write clear comments only for complex logic or non-obvious design decisions.\n4. Follow consistent coding conventions and formatting rules across the entire project.\n5. Refactor code regularly to remove technical debt and improve readability over time.", "reasoning_content": "", "tool_calls": [], "request_md5": "9af4a6122e824b100c1d31cb525d7816", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4", "expected_substrings": [], "retrieval_missing": [] } ]Correctness on the changed surface, in more detail:
cargo test -p hipfire-cpu --lib— 26 passed (GPU-free; the crate is inscripts/no-gpu-ci.sh).cargo test -p hipfire-config— 89 passed.python3 scripts/check-env-docs.py— clean.hipfire_cpu::gemvparity on identical bytes, both arms ofthe rotation contract: worst 9.3e-7 relative against a 1e-4 tolerance, 15 of 23
formats bit-exact. qt 49 (every projection of the 27B MQ3 fixture, and a format
with no canonical
dequantize_to_f32arm) is oracled on a real 27B tensor:1.195e-7 / 3.530e-7.
dequant_groupreproduces the canonical decoder bit-for-bit over 1695 realtensors. Zero-diff guard: stock ≡ this change ≡ spill+
pcie, byte-identical greedytext (2B 1343 chars, 9B 2734 chars).
9B, 8 spilled, greedy — both arms agree for 723 characters (~190 tokens) and then
differ by one whitespace token; the path is deterministic across processes and
across
RAYON_NUM_THREADS=1, so it is not CPU-side scheduling. On the 27B MQ3fixture both completions are 610 B and byte-identical.
--ignoredand fixture-dependent, so it is notin the GPU-free CI set; two formats (qt 11/12) report
SKIP — production launcher has no kernel on this archon gfx12.rustfmt --check --config skip_children=trueis clean on every file this branch touches, and
clippyis an advisory job.scripts/verify-bind-thread.sh(a pre-commit hook, not part of the required CIgatesjob) reports 12 of 94pub fninimpl Gpumissing thebind_threadcontract against 8 of 86 before this change. The four added host-mapped entry
points were fixed in this branch — the two allocators bind the calling thread, the
two state queries carry the gate's reasoned skip marker. The remaining 8 are
byte-identical to the pre-existing set and are left alone.
Known limits
arch-agnostic (
hipfire-runtime/src/llama.rsroutes the llama-familydown-projection through it); llama and qwen2 declare no offload support, and
MoE/A3B is structurally out of reach —
load_moe_ffnnever sees the offload flag.never enters
execute_steps, so their spilled weights still cross PCIe.27B MQ3 fixture, 8 of 64 layers (≈1.2 GB pinned) loads; 16 (≈2.5 GB) stalled twice
with 0 GB free.
byte-identical between the
cpuandpciearms (cpu.txt/pcie.txtindocs/perf-checkpoints/data-2026-09-27-cpu-exec-mq3-avx2/). The earlier record's "no decodedtext was read on this fixture" is a gap in that earlier run, closed by the later one — but a
second prompt or genre is still absent, and there is no ≥3-process arm-crossing byte diff on
this fixture — the cross-process evidence above is the 9B
cpuarm's determinism, which is adifferent property.
mq4g256_row_dot_scalaris nowunreachable).
Hardware validation request (optional)
Omitted deliberately: the CPU-executed path needs
HIPFIRE_GPU_LAYER_BUDGETandHIPFIRE_OFFLOAD_EXECset at load, which a registry-tag route cannot express, and aroute without the spill proves nothing about this change.
How this merges (direct review)
Merge authority is direct maintainer review plus the required no-GPU CI checks.
hw-gate automation is optional evidence delivery, not a prerequisite or substitute.
build (workspace, no GPU),unit tests (lib, no GPU),gates (ratchets, layering, registers)from.github/workflows/ci.yml.docs/VALIDATION.md. Prefer attaching local harness output (serve_harness.py,redline_daemon_harness.py,test_kernels, etc.). A lonehipfire runtranscript is not evidence.scripts/coherence-gate*.shandtools/change_gate/ agentic-review are historical only — never acceptance evidence.Architecture-trait change?
No —
crates/hipfire-runtime/src/arch.rs(theArchitecturetrait) is unchanged.