Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
37 commits
Select commit Hold shift + click to select a range
eb62684
feat(offload): CPU-exec GPU-idle accounting, and the pass-back headro…
Oct 1, 2026
c696b30
fix(docs)
Oct 1, 2026
965c9a4
feat(offload): memory.offload_exec=passback — split a spilled step ac…
Oct 1, 2026
f907988
docs(offload): pass-back mode — design § 6.2.2, config/env tables, TU…
Oct 1, 2026
80bb649
fix(offload): time the pass-back probe with a real sync, and read the…
Oct 1, 2026
02a324a
docs(perf): pass-back split checkpoint — gfx1201, 9B MQ4, budget 24 (…
Oct 1, 2026
619116a
feat(offload): put the scheduler's no-wait floor on the split trace line
Oct 1, 2026
42c5430
test(cli): cover offload-bench's refusal predicate and share recommen…
Oct 1, 2026
e728392
docs(perf): offload-amount sweep chart, 27B comparison point, in-plac…
Oct 1, 2026
ac01393
docs(perf): chart labels that cannot collide — values in the legend, …
Oct 1, 2026
62e172a
docs(perf): fix the amended record's stale ranges, its dangling secti…
Oct 1, 2026
534aff7
docs(perf): drop the long-horizon arc from the record, keep what the …
Oct 1, 2026
8ee62f5
fix(dispatch): censored, probe-anchored estimator for the passback share
Oct 2, 2026
308f8a9
perf(dispatch): re-estimate the passback share every 8 steps, not 64
Oct 2, 2026
99d8ca5
docs: describe pass-back as scheduled co-inference of the spilled layers
Oct 2, 2026
7262136
revert(dispatch): the 8-step estimator cadence is a regression
Oct 2, 2026
fc5c072
perf(dispatch): co-infer the dense FFN down-projection in passback
Oct 2, 2026
8f656c4
revert(dispatch): the censored estimator is a regression; keep the sh…
Oct 2, 2026
01f0457
docs: correct the FFN-down gain to the shipped-estimator measurement
Oct 2, 2026
9c72282
perf(dispatch): let the passback latch re-open when the balance point…
Oct 2, 2026
f97f836
docs: add the passback perf-gap and scheduling-algorithm investigations
Oct 2, 2026
ce8f3e5
perf(dispatch): tune the passback latch's re-open policy
Oct 2, 2026
00bb0fe
docs(perf): the passback coverage term is 4-6% of wall, not the deficit
Oct 2, 2026
0cc2cb5
docs(perf): record the passback latch band A/B as the null it is
Oct 2, 2026
d85faa9
docs(perf): amendments to the two 2026-10-02 passback checkpoints
Oct 2, 2026
6a082f6
docs(perf): correct the oracle bound direction (amendments 1 -> 2)
Oct 2, 2026
7443401
docs(perf): amendment 3 — fix the amendment-2 self-contradiction, clo…
Oct 2, 2026
a7e3e1c
docs(perf): current-HEAD passback mode A/B; make the changelog claim …
Oct 2, 2026
289b38a
docs(perf): pass-back cross-model spread (2B/9B/27B); pass-back loses…
Oct 2, 2026
b0b4141
docs(perf): quant spread, pass-back nondeterminism, mq6 blocker, mode…
Oct 2, 2026
2b1d60d
docs(perf): correct determinism framing, 27B point, mq6 root cause
Oct 2, 2026
c32e41c
docs(perf): quant comparison figure for the pass-back quant spread
Oct 2, 2026
d76f6be
docs(perf): scatter figure for the pass-back spreads (PR evidence)
Oct 2, 2026
059f879
fix(offload): route the FFN down-projection through the step seam; on…
Oct 2, 2026
196feda
fix(rdna-compute): serialize JIT compiles of one cache key
Oct 2, 2026
a39c7d4
chore(crate-maps): refresh the generated blocks this branch drifted
Oct 2, 2026
9599bd7
fix(hipfire-xdna): allow clippy::mut_from_ref on the documented BO view
Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -755,6 +755,7 @@ Caveats that are part of the fixture, not trivia:
| "Numbers don't match the README" | Forgot `HIPFIRE_NORMALIZE_PROMPT=1` (pre-2026-04-26) | Now default ON. Pull latest. If you opted out via `prompt_normalize=false`, that overrides the default — flip back. |
| "27B DFlash regressed 30-40% suddenly" | PR #32 (cleanup-dead-wmma-kernels) on master removed `gemm_hfq4g256_residual_wmma{,2,_k4}.hip` thinking dead. Dispatch fell back to slower variants. | Verify against canonical 199 tok/s @ max=120 with default flags. If kernel files missing in `kernels/src/`, `git checkout` from a known-good commit (see commit 9a2c667 for the full recovery context). |
| `HIPFIRE_GRAPH=1` reports plausible tok/s but output is garbage | Dangling stack-pointer kernargs from raw `self.hip.launch_kernel(...)` calls in `forward_scratch_layers` (kv_cache_write_*, attention_flash_*, fused_qkv_hfq4g256, rmsnorm_batched, rope_partial_interleaved_f32, gated_delta_net_q8, etc.) — captured pointers dangle past `end_graph_capture` | Bench tok/s alone never proves graph correctness. Always eyeball under `HIPFIRE_GRAPH=1` and run the claim-scoped VALIDATION serve route — never retired coherence-gate scripts as acceptance. Fix: migrate every raw-launch helper used in forward_scratch_layers to `launch_maybe_blob` (model after `conv1d_silu_split_f32_n`). |
| `memory.offload_exec` modes disagree on greedy output, or `passback` differs between two identical runs | Expected: the modes run different GEMV engines, and `passback`'s row split is scheduled per process (probe + online arm timings). Measured (2B): `cpu`×2 and `pcie`×2 byte-identical; `cpu` vs `pcie` differ; `passback`×2 differ; `HIPFIRE_OFFLOAD_PASSBACK_SHARE=0` byte-identical to `cpu` | NEVER assert a greedy/golden reference across a mode pair (`cpu` vs `pcie` is not bit-identical either), and NEVER assert `passback` output with the share on `auto` — pin `memory.offload_passback_share`, or use `0` to reach the exact `cpu` twin. See [`docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201-amendment-1.md`](docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201-amendment-1.md) § 5. |


---
Expand All @@ -778,7 +779,8 @@ Caveats that are part of the fixture, not trivia:
| `HIPFIRE_HOST_TIMING` | Per-cycle host timing probe | OFF |
| `HIPFIRE_VERIFY_GRAPH` | Verify-forward graph capture (0 = off) | ON |
| `HIPFIRE_GPU_LAYER_BUDGET` | Resident-layer budget for partial GPU offload: `N` keeps the last `N` layers on the GPU and spills the prefix to host RAM. Counts layers **on** the GPU, not offloaded — `3` on a 64-layer model spills 61. Unset/`auto`/`-1`/unparseable = fully resident (never forces offload). Sibling of `HIPFIRE_OFFLOAD_EXEC`. See [docs/plans/partial-gpu-offload-design.md](docs/plans/partial-gpu-offload-design.md). | unset (fully resident) |
| `HIPFIRE_OFFLOAD_EXEC` | Which engine multiplies a spilled layer's host-mapped weights: `pcie` (default) or `cpu`. Sibling of `memory.gpu_layer_budget`; `cpu` executes the weight-reading GEMVs on the CPU instead of reading them over PCIe. See [docs/plans/partial-gpu-offload-design.md](docs/plans/partial-gpu-offload-design.md) § 6.2.1. | `pcie` |
| `HIPFIRE_OFFLOAD_EXEC` | Which engine multiplies a spilled layer's host-mapped weights: `pcie` (default), `cpu`, or `passback` (scheduled co-inference — both engines run the same spilled step concurrently, divided by output rows). Sibling of `memory.gpu_layer_budget`; `cpu` executes the weight-reading GEMVs on the CPU instead of reading them over PCIe, and `passback` runs the `pcie` path on the step's first output rows and the `cpu` path on the rest, concurrently. See [docs/plans/partial-gpu-offload-design.md](docs/plans/partial-gpu-offload-design.md) § 6.2.1 and § 6.2.2. | `pcie` |
| `HIPFIRE_OFFLOAD_PASSBACK_SHARE` | Schedule of a spilled step's output rows between the two engines under co-inference (`memory.offload_exec=passback` only): the GPU takes the first rows (the `pcie` arm), the CPU the rest (the `cpu` arm). `auto` (default) schedules it from this host's measured engine rates and refines it online; a number in `(0, 0.5]` pins it; `0` collapses it onto the CPU engine alone, so the mode is byte-identical to `cpu`. Inert (one notice at load) in any other mode. Measure or override with `hipfire offload-bench`. See [docs/plans/partial-gpu-offload-design.md](docs/plans/partial-gpu-offload-design.md) § 6.2.2. | `auto` |
| `HIPFIRE_DDTREE_*` | Various DDTree diagnostics | various |

| `hipfire bench` flag | Purpose |
Expand Down
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,20 @@
# Changelog

## Unreleased
- **Offload pass-back: `memory.offload_exec=passback` is scheduled co-inference of the spilled layers — both engines run each spilled step concurrently.** It is a mixture of the two existing offload paths — the GPU reading the host-mapped weights over the link (`pcie`) and the CPU reading host RAM (`cpu`) — divided by output rows, not a third engine and not a fallback. A CPU-executed spill step is a host sync point (blocking D2H, host GEMV, blocking H2D), so the GPU is idle for most of its wall (77.7 % at 8 of 32 layers spilled on the 9B, 94.6 % at 24 of 32, 66.0 % on a 27B at 8 of 64 — [`docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md`](docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md)). A spilled step's output rows are independent, so the mode divides the step **by output rows**: rows `[0, g)` go to the GPU (the `pcie` path — the production kernel, reading the same host-mapped weight over the link) and rows `[g, m)` stay on the CPU (the `cpu` path), with the two arms running concurrently.
- **Both arms are the existing paths over a row range.** The GPU arm is `pcie`'s launch restricted to rows `[0, g)` (`launch_op_rows`, a byte view of the weight plus a pointer-offset output: every covered kernel indexes `A + row*row_stride` and writes `y[row]`), the CPU arm is `cpu`'s step restricted to `[g, m)` (`cpu_arm_prepare` / `cpu_arm_finish`). No third engine and no new numerics.
- **Ordering, not streams.** The blocking D2H is issued *before* the GPU arm is enqueued (so it drains only the step's producer), the GPU arm runs async, the CPU multiplies while it executes, and the blocking H2D of the CPU's rows is the join — stream-ordered after the GPU arm on the same (default) stream. No second stream, no events: the copies are `k*4` bytes down and `(m-g)*4` up (~30 KB at 9B shapes) against tens of MB of weight bytes.
- **`memory.offload_passback_share`** (`HIPFIRE_OFFLOAD_PASSBACK_SHARE`) selects the schedule: `auto` (default) schedules the split, a number in `(0, 0.5]` pins it, `0` collapses it onto the CPU engine alone so the mode is byte-identical to `cpu`. Read only in `passback` mode; elsewhere it is inert and prints one informational line at load.
- **The share is scheduled, not hardcoded** — the optimum is a property of the host, so nothing upstream depends on this box's constant. The first split-eligible step of each `(dtype, k)` seeds itself by timing both engines on that step's own weight buffer, and every split step refines the share online from its own arm timings. The join is read relative to the shape's own measured no-wait floor (the smallest join seen) plus a hysteresis margin: only an excess over that can be a wait for the GPU arm, and then the arm's duration is recoverable as `cpu_ns + excess` and the share moves toward the balance point. At the floor the CPU is the straggler and the share rises. Shares are clamped to `[0.05, 0.50]`, adjusted every 4 steps, and a shape latches once it has applied 64 adjustments with the last four under 0.005 — but the latch is **not permanent**: it keeps measuring, and re-opens and re-tracks (then re-latches) when the fresh proposal leaves `REOPEN_EPS` (0.01) in the same direction for `REOPEN_CONFIRM` (3) consecutive adjustments, so a long run whose host load shifts the optimum is followed rather than stuck at its first plateau. The same-direction run is what rejects the single-proposal excursion a host-load transient or one badly sampled join produces; `0.01` is twice the freeze band (the largest proposal a just-latched shape can hold), and it halves the original `0.02`'s balance-point dead zone (`REOPEN_EPS / ALPHA`: 0.067 → 0.033), which is what let a slow host-load drift slip past the latch. The band is not field-identifiable: the re-open path is live (the trace's `reopens` counter shows unlatches on serving runs), but a four-arm A/B left every arm inside its own run-to-run spread (±6 % at 16 of 32 spilled), so the value rests on that freeze-band margin and a deterministic host-load model in `offload_split`'s tests, not on a field win.
- **Output determinism: modes differ by engine, and `passback` varies per run.** `cpu` and `pcie` are not interchangeable byte-wise (different GEMV engines), so they need not agree — expected, not a bug (2B: `cpu` 2046 B vs `pcie` 3554 B). Within a mode at the same layer split the output is deterministic (`cpu`×2 and `pcie`×2 byte-identical). `passback` is the one mode **not** reproducible run-to-run: its row split is scheduled per process from the probe + online arm timings, so identical greedy invocations mix the engines differently (3670 vs 3078 B). `HIPFIRE_OFFLOAD_PASSBACK_SHARE=0` is byte-identical to `cpu`. Pin `memory.offload_passback_share` in any gate that diffs pass-back output against a reference; the 2026-10-01 record's 9B byte-identity held on its own prompts, where the argmax was insensitive.
- **Covered shapes:** `Gemv{Raw}`, `Gemv{Prerotated}`, and either residual form when the dtype has a fused residual kernel (the GPU arm of a residual step is `dispatch_residual`, whose dtype set is exactly `for_gemv_residual`'s) — qwen3.5 dense only; MoE/routed paths never reach it. Everything else runs on the CPU engine alone — the schedule's degenerate point, never `pcie` — so the mode is never less robust than `cpu`: a padded `row_stride`, a weight whose byte length is not exactly `m * row_bytes(q, k)`, no vector row dot, a non-F32 or short output, a weight below 2 MiB, or a residual step with no GPU residual kernel.
- **The dense FFN down-projection is now co-inferenced too.** `weight_gemv_swiglu_residual` fuses the SiLU into the down GEMV's launch, so the seam cannot see it as a step and it ran wholly on the CPU. Under `passback` it now hands that residual GEMV over as a `Step::GemvResidual` (falling back to the whole-CPU path when the seam declines). Measured on the 9B: passback **+3.8 %** at 8 of 32 spilled and **+2.6 %** at 16 of 32, interleaved fresh-process pairs against the build without it.
- **Pass-back vs engine, model size, and quant** (gfx1201, 2B/9B/27B, 128 tokens, greedy, `noslots`/`stateless`, interleaved fresh-process rounds, arm order rotated): pass-back beats `cpu` by **+19.6–21.6 %** on the 9B (25–75 % of 32 layers spilled), **+9.0–17.7 %** on the 27B (6.25–25 % of 64 — its deeper points are RAM-bound, not mode-bound, on a 28 GB host), **+15.2–19.4 %** on 9B mq3, and **+12–13 %** on the 2B; every paired round positive, so *over `cpu`* the gain is roughly flat across size and quant. What changes is **which route wins**: `pcie` is the loser on the 9B/27B (−15 to −28 % vs `cpu`) but the **winner on the 2B** (+42–45 % over `cpu`, ~20 % ahead of pass-back), and on mq3 `pcie` moves to ≈`cpu` (+3.0 % / −1.8 %) where on mq4 it was −15 to −22 %. So the recommended `memory.offload_exec` is size- and quant-dependent — `pcie` for a small model, `passback` for a mid/large one. Why pass-back loses on the 2B is not established: its scheduler settles at share 0.27–0.47 (measured from the split trace), so the `SHARE_MAX` cap is *not* binding. Output identity is verified for none of these arms: the modes are not byte-interchangeable (different GEMV engines) and `passback` is not reproducible per run — see the determinism bullet; the 2026-10-01 record's 9B mq4 byte-identity held on its own prompts only. — [`checkpoint`](docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201.md) + [`amendment`](docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201-amendment-1.md) + [figure](docs/perf-checkpoints/data-2026-10-02-offload-passback-model-spread/passback-gain-vs-spill.png); [`quant checkpoint`](docs/perf-checkpoints/2026-10-02-offload-passback-quant-spread-gfx1201.md).
- **`qwen3.5:9b-mq6` does not decode in this build** — a **plan-vs-fused routing mismatch**, not a pass-back result and not missing kernel coverage. Its tensors are the legacy `MQ6G256`; the table's plan for that dtype is the *unfused* `[SiluMulRotate, GemvResidual]` (`KernelKey::gemv_steps`) and `dispatch_gemv_residual` already wires `MQ6G256` to `gemv_hfq6g256_residual` (`families/gemv.rs:622`), but `weight_gemv_swiglu_residual` (`crates/hipfire-runtime/src/llama.rs:1543`) takes the *fused* launcher, and `dispatch_swiglu_residual` (`families/gemv.rs:637`) has no `MQ6G256` arm — so it hits the catch-all and every mode (passback / cpu / fully resident) fails with `unsupported gemv.swiglu_residual for /`. Exposed alongside: the key is declared/mapped/registered with no fused arm (`coverage_tests` checks key↔table, not arms); the `_ =>` error carries `arch/quant: ""` so it names no dtype (why 30 runs and a code read were needed); and the registry advertises this tag while the shipped build cannot decode it.
- **New `hipfire offload-bench`** measures the host directly: one row per dense format (`cpu`/`gpu` GB/s, the share, and the two-engine speedup bound), with `--json` and `--write` to persist the recommendation. It is a raw device benchmark and refuses while a daemon pid file names a live process.
- **Accounting stays honest.** A split step is not charged to the CPU-idle numerator (its wall contains GPU work) and gets its own `split: …` line under `HIPFIRE_CPU_EXEC_TRACE=1`, which also prints the controller's own state (`gpu samples`, `floor`, the last `waited`, the last `target`, `applied`, `reopens`, `frozen`) so a share nobody can explain is diagnosable from one run. The `cpu exec: idle` window line appends the split steps it excluded.
- **Design:** [`docs/plans/partial-gpu-offload-design.md`](docs/plans/partial-gpu-offload-design.md) § 6.2.2. Config keys in [`docs/CONFIG.md`](docs/CONFIG.md) / [`docs/env-vars.md`](docs/env-vars.md); TUI "Pass-back GPU share" row.

## v0.4.0 — 2026-09-30
- **Qwen3.8-Flash-Next (Qwen4, arch 16): `auto` KV is fp8 QSA K/V on exact gfx1201; the GDN recurrent state is Q8 by default; `max_seq` goes up to the native 262,144.** The QSA K/V arenas store E4M3 codes with one f16 scale per head and token, and the indexer's raw and pooled keys are stored as BF16, which holds the same values as the F32 arenas. `auto` keeps the exact `bf16` state (F32 arenas) on gfx1100, gfx1151 (Halo) and every other arch. An explicit `fp8` is refused off gfx1201, and every other `kv_cache` value is refused. The GatedDeltaNet state uses Qwen3.5's Q8 DeltaNet format on every arch; the `state_quant: "fp32"` load parameter opts out. Past 15,360 pooled blocks, the QSA selector keeps its score rows in global memory.
- **Validation (gfx1201, host-mapped experts N=12, tp=1).** There is no Flash-Next KLD reference yet (the MI300X requant is 0.4.1), so fp8 was validated by parity and greedy agreement, **not by KLD**. At 2K, 8K, 16K, 32K and 64K, every sampled QSA row matches the CPU F32-FMA reference in all three arms (fp8+Q8, bf16+Q8, bf16+fp32). The final-row argmax equals bf16 at every context, KL(bf16‖fp8) ≤ 5.1e-3, and top-5 overlap is 4 or 5 out of 5. Greedy text on the three flash prompts is coherent but not byte-identical to bf16 (first divergence after 40–442 characters). The serve battery loads with `auto` and serves 5 of 5 turns with 0 runaway and 0 empty turns.
Expand Down
1 change: 1 addition & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

4 changes: 2 additions & 2 deletions crates/hipfire-arch-qwen35/map.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside
| [`src/qwen35/config.rs`](src/qwen35/config.rs) | 1,947 | 43 | 25 |
| [`src/qwen35/ep_batch.rs`](src/qwen35/ep_batch.rs) | 5,641 | 23 | 8 |
| [`src/qwen35/forward.rs`](src/qwen35/forward.rs) | 7,549 | 32 | 12 |
| [`src/qwen35/load.rs`](src/qwen35/load.rs) | 7,632 | 16 | 14 |
| [`src/qwen35/load.rs`](src/qwen35/load.rs) | 7,663 | 16 | 14 |
| [`src/qwen35/oracle.rs`](src/qwen35/oracle.rs) | 1,243 | 20 | 0 |
| [`src/qwen35/prefill.rs`](src/qwen35/prefill.rs) | 17,032 | 18 | 70 |
| [`src/qwen35/weights.rs`](src/qwen35/weights.rs) | 2,991 | 43 | 11 |
Expand Down Expand Up @@ -100,6 +100,6 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside

### Totals

- 30 modules · 85,859 lines · 544 public items · 279 tests · 18 examples
- 30 modules · 85,890 lines · 544 public items · 280 tests · 18 examples

<!-- crate-map:generated:end -->
Loading
Loading