Repository navigation
Conversation
added 30 commits
October 1, 2026 10:00
…om record `HIPFIRE_CPU_EXEC_TRACE=1` now also prints, on the global step count's doubling schedule, `cpu exec: idle N% — window ending at step S ...`: the share of a decode window's wall spent inside CPU-executed steps. A CPU step is a host sync point (the blocking D2H, the multiply, the H2D) and prefill never enters the seam, so that share is a lower bound on the GPU's idle fraction — the headroom a scheme that hands part of a spilled step back to an idle GPU would be spending. Windowed rather than cumulative, because `hipfire bench --runs N` decodes N times inside one process and the gaps between runs are neither. Measured with it (gfx1201, RX 9070 XT, PCIe 5.0 x16, HIP 7.2, daemon md5 7f115e384f4fe1521d02dc57fbd02745): the idle share rises with the CPU's share of the split — 9B 77.3/77.7/77.8% at 8 of 32 spilled, 89.8% at 16, 92.4% at 20, 94.6% at 24; 27B-mq3-xt 66.0% (repeat 66.3%) at 8 of 64. A four-process contention probe (CPU gemv stream over a 205 MB host-mapped buffer, PCIe H2D of the same bytes, then both) shows the two streams are close to additive: the CPU retains 86-94% of its solo rate while the GPU runs at 102-103% of its, combined 75-78 GB/s — ~1.4x CPU-alone and ~90% of this host's DDR5 peak. So the idle window is real, and a row-split pass-back would be bounded by host DRAM rather than by either engine. Record and raw captures: docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md and its data directory (bench JSON, per-window idle lines, the probe source). Rates there are n=1 per process and labelled exploratory; the idle fractions and the probe are the evidence. The record also documents the 27B load hang (layer 45/64, 99% CPU, 0 GB free) as the capacity precondition the 2026-09-27 amendment already flagged, and that the 2026-09-27 9B fixture is not in any registry entry or HF revision, so both arms are measured fresh on the registry artifact (qwen3.5:9b, sha256 ba83acf5...). Diagnostic only: no behavior change to the resident or offload paths, no new env var. Docs updated in env-vars.md, partial-gpu-offload-design.md 6.2.1 and the benchmark handoff.
LLM incorrectly determined that GPU was a PCIe 5.0 x16 link. This is correct in what the GPU is capable of, but incorrect in what it actually has available. Both the GPU and CPU on this system support a PCIe 5.0 x16 link - the B650 motherboard does not. As a result the link is capped at PCIe 4.0.
…ross both engines
A CPU-executed spill step is a host sync point (blocking D2H, host GEMV,
blocking H2D), so the GPU is idle for most of its wall: 77.7 % at 8 of 32 layers
spilled on the 9B, 94.6 % at 24 of 32, 66.0 % on a 27B at 8 of 64
(docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md). The same
record measured the headroom: a CPU GEMV stream and a GPU PCIe read of the *same*
host-mapped bytes are near-additive to 75–78 GB/s combined against ~46 GB/s for
the CPU while the GPU reads.
A spilled step's output rows are independent, so this mode splits the step by
output rows: rows [0, g) go back to the GPU (the production `pcie` kernel on the
same host-mapped weight) and rows [g, m) stay on the CPU, concurrently.
* Both arms are the *existing* paths over a row range, so there is no third
engine and no new numerics — and `memory.offload_passback_share = 0` is
byte-identical to `memory.offload_exec=cpu`.
- `launch_op_rows(gpu, ctx, step, Some(0..g))` restricts a `pcie` launch: a byte
view of the weight's rows plus `sub_offset` views of the output and residual.
Every covered kernel indexes `A + row*row_stride` from the passed base and
writes `y[row]`, so a row-shifted base *is* a row-shifted launch. `None` is the
pre-existing path, unchanged and still the only one `pcie` uses.
- `cpu_exec::cpu_arm_prepare` / `cpu_arm_finish` are `cpu` mode's step in its two
observable phases; `run_host_mapped_gemv{,_residual}` and `run_step` keep their
signatures and run the two back to back, so `cpu` mode's device-operation
order is unchanged.
* Ordering, not streams. The blocking D2H is issued before the GPU arm is
enqueued (so it drains only the step's producer), the GPU arm runs async, the
CPU multiplies while it executes, and the blocking H2D of the CPU's rows is the
join — stream-ordered after the GPU arm on the same stream. The copies are
`k*4` bytes down and `(m-g)*4` up against tens of MB of weight bytes, so a
second stream plus events would buy the overlap of a ~2 µs copy.
* `memory.offload_passback_share` (`HIPFIRE_OFFLOAD_PASSBACK_SHARE`): `auto`
(default) schedules the share, a number in (0, 0.5] pins it, `0` disables the
pass-back. Inert outside `passback` (one informational line at load). Requires
the same config plumbing as the mode: `OffloadExec::Passback`,
`ValueRule::PassbackShare`, the TUI row/knob and the schema arms.
* The share is scheduled, not hardcoded — the optimum is a property of the host,
so nothing upstream depends on this box's constant. The first split-eligible
step of each (dtype, k) seeds itself by timing both engines on that step's own
weight buffer (GPU arm through a scratch-output probe, so a residual probe
cannot be double-counted), and every split step refines the share from its own
arm timings: `r_cpu` from the contended CPU arm every step, and `r_gpu` only
when the join outlasts the CPU's own multiply — the one regime where the
blocking H2D demonstrably waited for the GPU arm. Below it the join *is* the
CPU's duration, so that sample would read ~20× low and pin the split to its own
floor with both engines idle in turn; there the share rises instead. Clamped to
[0.05, 0.50], adjusted every 4 steps, frozen once a shape has applied 64
adjustments with the last four under 0.005.
* A failed seeding probe degrades to DEFAULT_GPU_SHARE rather than failing a step
that plain `cpu` mode would have run, and every refusal falls back to the
whole-CPU step — never to `pcie`. Refused: a padded `row_stride`; a weight whose
byte length is not exactly `m * row_bytes(q, k)` (the invariant that makes
`&host_bytes[g*row_bytes..]` the CPU arm's row `g`); no vector row dot (with the
scalar decoder the best share is all-GPU, which `pcie` does better); a non-F32 or
short output; a weight below 2 MiB; and a residual step with no GPU residual
kernel (`dispatch_residual`'s dtype set, which is `for_gemv_residual`'s).
* Accounting stays honest: a split step is not charged to the CPU-idle numerator
(its wall contains GPU work), gets its own `split:` trace line under
`HIPFIRE_CPU_EXEC_TRACE=1` (`join` = the blocking H2D, `cpu=`/`gpu≥` the sampled
and bounded rates), and the `cpu exec: idle` window names the split steps it
excluded.
* `hipfire offload-bench` measures the host directly (one row per dense format:
cpu/gpu GB/s, the share, the two-engine speedup bound; `--json`, `--write`),
building its own buffers via `offload_split::probe_synthetic` and refusing while
a daemon pid file names a live process. It lives in `hipfire-runtime` so the CLI
grows no dependency.
Evidence: `cargo test -p hipfire-config` 101, `-p hipfire-dispatch` 310,
`-p hipfire-tui` 165 pass. New `#[ignore]`d row-offset equivalence arms
(`tests/gpu_gemv_parity.rs`) are bit-exact on gfx1201 over 15 cases — real 2B/9B
MQ4/MQ3-Lloyd/MQ6/HFQ6 tensors, their odd-row-count prefixes, and synthetic
MQ4G256V2/MQ4CG256/MQ3G256V2/HFQ4G256/Q8_0 — for both arms of both step forms,
including that rows past `m` are never written and that a residual shape with no
GPU arm is refused. `memory.offload_passback_share=0` output is byte-identical to
`memory.offload_exec=cpu`, and `pcie` output is byte-identical with and without a
stray share set.
…I row, AGENTS, changelog * `docs/plans/partial-gpu-offload-design.md` § 6.2.2: the row split, why ordering rather than a second stream buys the overlap, the four covered shapes, the refusal list, the scheduler (probe + the join-vs-gemv online signal), the share key, and the measured ceiling from the 2026-10-01 headroom record. * `docs/CONFIG.md` / `docs/env-vars.md`: the generated lifecycle/inventory rows (`scripts/check-lifecycle.py --write`) plus the hand-written prose row for `HIPFIRE_OFFLOAD_EXEC` (`passback`) and a new `HIPFIRE_OFFLOAD_PASSBACK_SHARE` row. * The pass-back share is described for operators as a scheduled value with a pinned override, and as the byte-identity A/B twin at 0. * `AGENTS.md` § 7 flag table gains `HIPFIRE_OFFLOAD_PASSBACK_SHARE` and names the third `HIPFIRE_OFFLOAD_EXEC` value. * `CHANGELOG.md`: an `## Unreleased` entry for the mode. Gates: `scripts/check-lifecycle.py` and `scripts/check-env-docs.py` both exit 0.
… join against its own floor Two measurement bugs found while verifying `memory.offload_exec=passback`, both of which produced plausible-looking wrong numbers rather than errors. 1. **`sync_with_deadline` cannot time anything.** It polls completion with `std::thread::sleep(SYNC_POLL_INTERVAL)`, `SYNC_POLL_INTERVAL = 2 ms` (`crates/rdna-compute/src/dispatch.rs:1264`), so a probe that times `launch → sync_with_deadline` measures the *sleep*. The seeding probe and `hipfire offload-bench` therefore reported 4.1 GB/s for an 8 MiB host-mapped weight whose true rate is 24.4 GB/s, and the artifact looked exactly like a size-dependent fixed cost (~1.7 ms/launch) — which is how it survived a first read: `--buffer-mb` 8/16/24/48/96/192 gave 4.1/8.1/12.2/12.2/16.3/24.4 GB/s, a textbook `a + b·bytes` fit, and the fit's slope (~32 GB/s) was right while its intercept was the poll interval. `offload_split`'s probe now uses `gpu.hip.device_synchronize()`; deadline-bearing paths keep the polling sync. The seeded share moved from ~0.23 to ~0.36, and the bench's GPU column became size-independent (24.4–26.6 GB/s over 8 MiB–192 MiB). 2. **The controller read the join as a GPU duration in the regime where it is the CPU's.** The join is `copy + max(0, gpu_ns − cpu_ns)`, sampled as `rows_gpu·row_bytes / (cpu_ns + join_ns)`. Whenever the CPU arm is the straggler (the common case: a 0.1–0.6 ms CPU multiply against a 30–50 µs copy) that denominator is the *CPU's* duration, so the sample read ~13× low — and since a seeded `r_gpu` then wins `next_share`'s both-rates arm forever (the EWMA only feeds on waits), it pinned every shape to 0.071–0.081, just above `SHARE_MIN`, with both engines idle in turn. Any absolute microsecond threshold has the same failure on this host, because the copy alone costs more than the 20 µs the threshold assumed. The join is now measured against the shape's **own** no-wait floor (the smallest join seen) plus a hysteresis margin (10 µs, or a quarter of the floor): only the excess over that can be a wait, `gpu_ns = cpu_ns + excess`, and at the floor the share rises instead — the direction the design always intended. Measured shares: 0.279–0.360 on the 9B's field shapes, against 0.28–0.45 by format and 0.36–0.40 for the same sizes from the standalone bench. Also in this commit: * The `split:` trace line now prints the controller's own state (`gpu samples`, the last `waited`, the last `target`, `applied`, `frozen`). A share nobody can explain is diagnosable from one run; the two bugs above took one run each to localise because of it. * `offload-bench`'s recommendation is the **median** share over the largest measurements. Every format is measured over the same `buffer_mb`, so "the largest measured format" is normally the whole matrix and `max_by_key(bytes)` degenerated to whichever format happened to be last. * The residual gate refuses *any* `GemvResidual` whose dtype has no GPU residual kernel, not just the `Raw` form: the `Prerotated` form would have failed to launch where `cpu` mode handles it. (Found by the parity test's Q8_0 case.) * The row-offset parity arms report the **two arms' mutual delta** and require at least one case where the engines' numbers differ, so "the split mixes both engines" is asserted rather than assumed: 12 of 15 cases differ (1.2e-7–1.5e-5). Verification (gfx1201, RX 9070 XT, 9B MQ4 registry artifact, budget 24 = 8 of 32 spilled, greedy, `--spec off`, pinned prompt md5 5835c71e…): * 15 parity cases bit-exact per arm, plain and residual forms, including no writes at or past `m`. * `share = 0` output byte-identical to `cpu` (555-byte and 642-byte completions, `cmp`), and no `split:` line at all. * `pcie` with a stray share byte-identical to `pcie` alone, one notice line. * `hipfire bench`, three interleaved fresh-process rounds: `cpu` 27.40, `passback` 30.60 (+11.7 %), `pcie` 23.20 tok/s; `vram_free_mb` identical across arms. * `offload-bench` exit 0 with sane rates (GPU 23.0–28.2 GB/s vs the record's 27.3; CPU 34.7–68.4 vs 52), and exit 1 naming the live pid while a daemon runs. * `-t 0 --spec off -n 256/512` text: `cpu`, `passback` and `pcie` all byte-identical on this fixture (the 2026-09-27 cpu-vs-pcie divergence was measured on a different, local-only 9B).
…8 of 32 spilled) Dated, fixture-bound `historical` record for `memory.offload_exec=passback`, the companion of the headroom record that motivated the mode. One host (RX 9070 XT, PCIe 4.0 x16, Ryzen 7 7800X3D, HIP 7.2), one model (registry 9B MQ4, sha256 ba83acf5…), one prompt (md5 5835c71e…), one budget (24 of 32 layers resident), greedy `--spec off`. Establishes: +11.7 % decode over `cpu` on the median of three interleaved fresh-process rounds (`cpu` 27.40 / `passback` 30.60 / `pcie` 23.20 tok/s, VRAM identical across arms); the scheduler converges unaided to 0.279–0.360 on the field shapes against the standalone bench's 0.28–0.45 by format; the two arms are bit-exact per arm over 15 parity cases while 12 of them mix numerically different engines; `share = 0` is byte-identical to `cpu`; and the probe's pre-fix table, which measured the 2 ms `sync_with_deadline` poll interval instead of the kernel, is kept because the wrong number is the kind that gets quoted. Also records what is **not** measured: the serve/slots path, prefill, any other host/link/arch/model, long-context decode, and the two-stream contingency.
`ShapeSnapshot` gains `min_join_ns`, and the `split:` line prints it as `floor=0.04ms`. The floor is the quantity the controller actually compares every join against (the smallest blocking-H2D time seen for the shape, i.e. what the copy costs when the GPU arm finished first), so a share that looks wrong is now diagnosable from one run without reading the state: `waited=false` with a floor at the observed join means the share is being pushed up because no step has waited yet, and `gpu samples 0` says the GPU rate the share rests on has never been observed at all.
…dation Both paths are user-facing and had no coverage — the refusal is the only thing between a user and two engines sharing the device, and the recommendation is what `--write` persists. * `live_daemon_pids` is pure filesystem + `kill(pid, 0)`, so it is tested without a GPU or a real daemon: a live pid file is listed, a pid above the kernel's `pid_max` and a non-numeric file are not, `serve.pid`/`daemon.pid.bak` are not pid files, and a missing root is empty rather than a panic. * The recommendation is extracted from `offload_bench_command` into `recommended_share`, which pins the tie behavior: one measured format is named outright, and equal-size measurements (the normal case — every format is measured over the same `buffer_mb`) take the **median**, not whichever format `max_by_key` happened to return last (q8_0, whose rate profile is not the model's). `--help` for the new subcommand is rendered and checked; `--write` reuses the `config set` path (`load_global` → `set_cli` → `write_global_toml`), which the config tests already cover.
…e record amendment
Adds to the dated pass-back checkpoint:
* § 8 — the offload-amount sweep the record owner asked for. 9B mq4, 4/8/12/16 of
32 layers spilled, 2 interleaved rounds per point: pass-back beats `cpu` by
**+8.4 % / +8.5 % / +13.2 % / +13.9 %** and `pcie` by +20 % / +30 % / +39 % /
+44 %. The advantage grows with the spill, and grows faster against `pcie`
(which pays the whole link read per step while `cpu` and pass-back share host
DRAM). Plus one 27B point (`qwen3.8-27b.mq3-xt`, its pinned md5, 8 of 64 spilled):
`cpu` 14.40 / `passback` **17.10 (+18.8 %)** / `pcie` 14.40 tok/s — a bigger model
gains more at the same *number* of spilled layers, because each spilled
projection is a larger vector and the CPU arm's serialized host time per step is
longer. A second 27B round was abandoned when the `pcie` arm wedged under memory
pressure (`free` at 0 GB), the stall the headroom record already documents.
* The chart: `data-2026-10-01-offload-passback-split/mode-sweep-vs-spilled-layers.{png,svg}`
— (a) tok/s per engine against layers spilled with a titled legend, (b) the 27B
point with the `cpu` baseline marked, (c) the gains on their own axis so no
percentage label sits under a data line.
* § 9 — the `offload-bench` CLI paths, unrecorded until now: the live-daemon
refusal (exit 1, empty stdout, the pid named), `--json`, `--write` →
`config get` → `config reset` leaving the config byte-identical.
* § 10 — the 32k-token long-range arc's fixture and intent (its numbers land in a
follow-up commit once the three arms finish).
Amended **in place** on the record owner's explicit instruction; the file's header
now says so and marks every changed passage `[amended]`, because
`docs/perf-checkpoints/README.md` otherwise routes corrections to a new dated file.
Corrected there: the summary bullet's VRAM range (said "11308–11308 MB"; the data
are 11308–11336 across the nine arms), and the identity table's host-load
provenance (the load is `orcaslicer_main` at ~55 % of a core, not transient
YouTube, so the host was never near the plan's `uptime ≈ 0.5` and the absolute
rates are depressed — the ratios are what carry).
`benchmarks/prompts/long_range_count.txt` is the committed arc prompt (md5
`abab011aadb5cb1d3ef86a0e646cc1a3`): a self-verifying task, so a reviewer can check
the 32k-token outputs mechanically rather than reading prose.
…gains on their own axis
The record owner's complaint was that the data eclipsed the percentage labels. Two
changes fix it by construction rather than by tuning label offsets:
* panel (a) carries no in-plot text at all. Every series' four values live in the
legend, one column per spill amount (cpu 43.8 / 28.1 / 20.5 / 16.2, and likewise
for passback and pcie), so no number can land on a line, on another number, or
outside the frame. The earlier in-plot version collided at two of the four x
positions and pushed the rightmost label past the axis.
* the pass-back's gain is its own panel (c), so no percentage ever sits under a data
line.
Panel (b) gains the cpu-baseline dashed line and a "+18.8 % over cpu" caption, and
loses a redundant legend (its x tick labels already name the engines) and a caption
that rendered red-on-red across the cpu bar's own top edge.
The chart tool that produced the figure from the raw bench JSON is committed beside
the artifact, so it is reproducible rather than a hand-drawn one-off.
No further plot work: the figure at
`docs/perf-checkpoints/data-2026-10-01-offload-passback-split/mode-sweep-vs-spilled-layers.{png,svg}`
is final.
…on, and sharpen the 27B claim Three corrections to the amendment that landed in e728392, plus one analytic sharpening: * Bullet 6 quoted "19.6–28.5 GB/s" for the GPU arm and "29–73" for the CPU — the 19.6 and the 29 are from the **pre-sync-fix** probe table that § 3 keeps only as a discarded comparison, so the bullet contradicted § 2's own shipped table. It now quotes § 2 (22.8–28.2 and 34.7–68.4 GB/s) and says what moved. * The § 4 reading said the probe "seeds ~0.23–0.36"; 0.233 is the pre-fix seed the dead `sync_with_deadline` produced. The shipped seed is ~0.36, and the text now says so. * § 10 pointed at a § 10.1 that does not exist yet (the arc outlives this commit). It now points at the briefing that does exist, with § 10.1 promised on completion. * The 27B claim now compares at equal **spill fraction**: 8 of 64 spilled is 12.5 % of the model, the same fraction as the 9B's 4-of-32 point (+8.4 %), so +18.8 % isolates model size from spill amount. Comparing at equal spill *count* (8 of 64 vs 8 of 32) confounds the two, and the 27B wins there too despite spilling half the fraction.
…attempt taught The record owner asked for a long-horizon *text* run and then dropped that line of work, so the arc (its prompt, its outputs and its scratch directory) is gone rather than half-reported. § 10 now says exactly that, and replaces the promised numbers with the two pitfalls the attempt established — both are facts a future long-budget run needs: * `hipfire run` has no reasoning flag. With `reasoning.mode = on` (the config default; the 9B is a reasoning model) two arms spent ~20 minutes of real decode each and wrote **1 byte**: a bare `println!()` around an empty visible answer, because the whole budget went into the invisible think block. `reasoning.mode = off` streamed text immediately (27 tok/s, matching the sweep's `cpu` point). AGENTS section 7 documents the `bench` side of this; `run` exposes only the config key. * `-n` does not set an arc's length — the model's own stop does. Told to "keep going as long as you are allowed to", the 9B EOSed at 1000 (997 integers, 3887 bytes, ~3.9k tokens, 143 s), so a long run has to get its length from the prompt. Also removes the now-unused `benchmarks/prompts/long_range_count.txt` (its md5 was recorded only as that arc's fixture), and restores `reasoning.mode` to `on` in the global config, which the attempt had flipped. Closing gate: `cargo build --release` clean; `hipfire-config` 101, `hipfire-dispatch` 311 (+1 ignored), `hipfire-tui` 165, `hipfire-cli` 306 tests pass; `scripts/check-lifecycle.py` and `scripts/check-env-docs.py` exit 0.
The passback share controller estimated the GPU arm's rate from an EWMA of only the steps where the GPU overran the CPU multiply — a sample conditioned on the GPU's slow tail — against a running-minimum join floor, with the blocking D2H folded into the CPU arm's rate, and then froze permanently once it had applied 64 adjustments. Replace the estimator, keeping the setpoint: the DLT two-processor balance point r_gpu/(r_gpu+r_cpu) is the correct optimum and is not the defect (Cheng & Robertazzi's optimality principle, via arXiv:1902.01898 §II-A). * the CPU arm's time-per-byte is sampled from the multiply alone; * the GPU arm's is a censored (Tobit) EM estimate over both exact (overrun) and censored (finished-first) observations, anchored by the one-time probe's directly measured rate — which is also the estimator's identifiability anchor when every online observation is censored; * the permanent freeze and the blind ratchet are removed, so the share keeps adapting for the whole decode rather than ceasing to move at 64 adjustments; * if the automatic probe fails the share escalates upward while unanchored instead of sticking at the default, and the trace marks `anchor=none`; * the split trace now prints the controller's own est_cpu/est_gpu (from the same state the share was computed from) rather than a per-shape mean. Known limits, documented at the estimator and measured on the reference fixture: the anchor is a *solo* measurement while the optimum is contended, so the fixed point sits near 0.46 against a measured optimum of ~0.42; and the GPU level does not re-base over a long horizon — fading the anchor lets floor-noise overruns run the estimate away downward (est_gpu 2-9 GB/s), so the floor must be made robust before the anchor can decay. Sources are cited above the scheduler.
The estimator's own lag is the investigation doc's §1.2 weakness 1 ("lag, not
offset"): the GPU arm's estimate re-ran on a 64-step batch, so the share trailed
the drifting balance point by the batch length. Keep the same statistics over a
64-observation rolling window, re-estimated every 8 steps — the doc's §5 Phase 1
lever ("reduce the target EWMA's lag") applied to the estimator.
Also rewrites the module doc to the co-inference framing: pass-back is the two
offload paths (the GPU reading host-mapped weights over the link, the CPU reading
host RAM) run concurrently on one spilled step, not a fallback, and the schedule
is per step, not per layer.
Pass-back is a mixture of the two offload paths — `pcie` (the GPU reading the host-mapped weights over the link) and `cpu` (the CPU reading host RAM) — run concurrently on the same spilled step and divided by output rows; the share is that schedule. Corrects every authored description: the design doc §6.2.2 (and its heading anchor), the config schema (which feeds the env-vars table), AGENTS.md, the CHANGELOG entry, and the benchmark-handoff note. A step placed wholly on one engine is the schedule's degenerate point (`share -> 0`), not a validation failure and not a fallback to a worse path.
Measured on the 9B fixture (qwen3.5:9b, budget 24 = 8/32 spilled, the pinned gpu_offload_probe prompt, five interleaved fresh-process pairs, daemon binaries built from 8ee62f5 and from 308f8a9): the 8-step rolling window gave 31.60 tok/s median against 32.50 for the 64-step batch, losing all five pairs, and pushed the settled share to 0.495 against 0.464 (further from the optimum). The estimator lag was low-pass-filtering the anchor bias, not costing throughput, so the investigation doc's §1.2 "reduce the target EWMA's lag" lever does not hold on this fixture. Restores ESTIMATE_EVERY = ESTIMATE_WINDOW, at which the rolling buffer degenerates to the disjoint batch 8ee62f5 shipped. Keeps the VecDeque mechanism and the window constant, with the measurement recorded at the constant so the lever is not re-attempted blind.
`weight_gemv_swiglu_residual` fuses the SiLU into the down GEMV's launch, so the passback seam cannot see it as a `Step`: it self-split at the op level (SiLU on the GPU, the whole GEMV+residual on the CPU) and, because its guard is the mode-agnostic `host_mapped_cpu_capable`, that happened under passback too — so the dense FFN down-projection, the largest non-split shape in the trace, ran wholly on one engine (~24% of the spilled bytes on the CPU engine alone). Under passback it now builds a `Step::GemvResidual` over the real AWQ sidecar and hands it to `offload_split::run_with`, falling back to the whole-CPU GEMV when the seam declines. cpu mode, pcie and resident loads are untouched: the branch is gated on `offload_split::enabled()`, and the certified GPU fused arm is not reached. Measured on the 9B (qwen3.5:9b, budget 24/16 = 8/16 of 32 spilled, the pinned gpu_offload_probe prompt, interleaved fresh-process pairs, passback mode): **+4.0 % at 8 spilled and +5.0 % at 16 spilled** against the build without it, so the passback-vs-cpu gain at 16 spilled widens from +12.9 % to +17.6 %. Validation: hipfire-dispatch 313 lib tests, hipfire-runtime 956; serve_harness battery + chain on the passback arm 5/5 turns with runaway/empty/attractor 0.
…ipped one Interleaved passback A/B, fresh process per run, daemons built from 534aff7 and 8ee62f5 (both predating the FFN-down change, so this isolates the estimator): the shipped estimator gave 32.40 vs 31.80 tok/s at 8 of 32 spilled (-1.9%) and 19.00 vs 18.80 at 16 of 32 (-1.1%), losing 3/3 pairs at both points. The premise of 8ee62f5 — that the shipped estimator's tail-conditioned bias drives the share below the balance point — is smaller on this fixture than the solo-probe anchor bias the replacement introduced (placement ~0.46 vs the shipped 0.40-0.44), so it traded a theoretical bias for a measured one. Restores offload_split.rs and cpu_exec.rs to 534aff7's estimator and trace and re-applies only the co-inference framing to the module doc (that change is orthogonal). The FFN-down co-inference from fc5c072 is untouched and still measures +3.8% (8 of 32 spilled) / +2.6% (16 of 32) on the shipped estimator, so the stack is now net-positive with no measured regression left in it. Validation: hipfire-dispatch 311 lib tests; serve_harness battery + chain on the passback arm 5/5 turns, runaway/empty/attractor/retrieval_miss 0.
The +4.0%/+5.0% figures were measured on the censored estimator, since reverted in 8f656c4. On the shipped estimator the FFN-down co-inference measures +3.8% (8 of 32 spilled) and +2.6% (16 of 32), interleaved fresh-process pairs, passback mode. Corrects CHANGELOG.md and the design doc § 6.2.1.
… moves The freeze gate latched a converged shape permanently — the investigation doc's §1.2 weakness 3: it never reopens, so a long run whose context or host load moves the optimum stays stuck at its first plateau. Removing the latch outright was tried first and measured as a large regression (31.60 vs 32.50 tok/s at 8 of 32 spilled, -7.0%, and -11.8% at 16 of 32, 4/4 pairs): unfrozen, the shipped estimator's proposal wanders. So the latch stays and gains a re-open. A latched shape keeps *measuring* — the rate EWMAs never stop — and every APPLY_EVERY samples `observe` recomputes the proposal; if it differs from the latched share by more than `REOPEN_EPS` (0.02) the shape unlatches, re-tracks and can re-latch. The band sits above the largest proposal a just-latched shape can be holding (`FREEZE_EPS / ALPHA ≈ 0.0167`), so the bands cannot overlap and the latch cannot thrash. Measured: latch-with-reopen vs plain latch is a tie on the static fixture (+0.6% / +1.5% at 8/16 of 32 spilled, -0.3% at 512 tokens, all inside ±1.5%), which is the intent — it behaves exactly as the latch until the balance point actually moves. A `REOPEN_EPS` spread (0.005/0.01/0.02/0.05/no-reopen) confirmed the structural floor: 0.005 — below `FREEZE_EPS / ALPHA` — was the worst (-1.8% / -5.2%), while 0.01/0.02/0.05 sat inside the bench's noise. Calibrating the band from the measured proposal scatter at runtime is the next step; a static bench cannot resolve it. Validation: hipfire-dispatch 7 scheduler tests, including a latch-then-reopen unit test; serve_harness battery + chain on the passback arm 5/5 turns, counters 0.
The research record behind the passback scheduling work: the perf gap (pcie 23.6 / cpu 28.1 / passback 31.1 tok/s at 8 of 32 spilled), the algorithm survey, the constraints (no deliberate excitation, O(1)/step, cross-machine), and the Phase 0-3 plan. Untracked until now.
The freeze latch re-opened on a single out-of-band proposal compared against a fixed 0.02 band, both measured on the ALPHA-filtered proposal. Two things break there. A slow host-load drift moves the balance point a little per adjustment, so no single proposal ever leaves a band that size — on the balance-point axis `0.02 / ALPHA ~= 0.067` — and a latch that re-latches after one step change then sits up to 6.7 points off the drifting optimum. And one proposal past the band can come from a host transient or a single badly sampled join, so unlatching on it hands the controller the noise the latch exists to reject. * `REOPEN_EPS` 0.02 -> 0.01: halves the balance-point dead zone (0.067 -> 0.033) while keeping a 2x margin over the freeze band (`FREEZE_EPS` = 0.005, the largest proposal a just-latched shape can be holding). A band at the freeze band re-opens on the latch's own convergence noise: the model re-opens a *clean* static plant nine times in 1024 steps at 0.002. * `REOPEN_CONFIRM` = 3: unlatch only after three consecutive same-direction out-of-band proposals. A genuine move drives the balance point one way for many adjustments; noise flips sign, so the run separates them without a wider band. It is insurance with a cost — ~1-2 us of extra reopen latency per genuine move in the model, no measured throughput change — and it removes the reopens a noisy static host produced without it. * `ShapeState` / `ShapeSnapshot` gain a `reopens` counter, printed on the `split:` trace line. `frozen` alone cannot show whether the path was exercised. No runtime calibration of the band: the only per-shape scatter the loop holds is the converged proposal magnitude, which the freeze guarantees is under `FREEZE_EPS`, so a scatter-derived band would land with no margin. Validation. The band is not field-identifiable: a four-arm A/B (0.005 without the streak, 0.005/0.01/0.02 with it, 3 interleaved fresh-process rounds at 8 and 16 of 32 spilled on qwen3.5-9b.mq4, prompt md5 5835c71e471849b4a72e1dc8e39695e7, model md5 296092bf1e6a45d78c1acf815eb93366) left every arm inside its own run-to-run spread (+-6% at 16 of 32) and flipped the ordering between budgets; the 0.005 penalty an earlier spread measured did not reproduce on this host. So the value rests on the freeze-band margin and a deterministic host-load model in `offload_split`'s tests (step / ramp / noisy-static / transient scenarios), not on a field win. On live serving the trace's `reopens` shows shapes unlatching 0-5 times, so the path is real. * hipfire-dispatch 316 lib tests (11 scheduler, 4 new); hipfire-runtime offload and hipfire-cli `recommended_share` / `live_daemon` pass. * serve_harness battery + chain on the passback arm: 5/5 turns, runaway/empty/ retrieval 0. The sampled and greedy runs flag a token attractor on the prose turn; it reproduces on the `cpu`-mode control under identical settings (worse there: 3 vs 2 under greedy), so it is this model's greedy think-channel repetition, not the scheduler. Binaries: daemon b35c25f032f80882d757512d0f8ee98a, hipfire ee24d2671dc3480ac8949a43df71b156. Not touched: the estimator's open defects (monotone `min_join_ns` floor; `r_gpu` sampled only on a waited join) — documented in docs/plans/partial-gpu-offload-design.md 6.2.2 "Known estimator limitations".
Phase 0 of the offload pass-back investigation: the measurement docs/investigations/2026-10-01-offload-passback-perf-gap.md 7 names, and the gate its 5 Phase 0 put on further scheduler tuning. The perf-gap record's open question was whether the refused (never-split) steps carry a disproportionate share of CPU-executed wall — the "32% of steps never split" reading, tagged [INFERENCE] and explicitly "an assumption, not a measurement". Bucketing the HIPFIRE_CPU_EXEC_TRACE lines by wall time on the pinned 9B fixture answers it: spilled decode tok/s coverage (refused wall / cpu-exec wall) 8 / 32 33.0 3.73% 16 / 32 19.5 3.72% 24 / 32 13.9 6.30% The refused set is a single shape at every point: m=32 k=4096 (the DeltaNet beta/alpha rows, ~68 KiB, under MIN_SPLIT_BYTES), 8 calls per token. At 8 spilled that is ~2.7% of token wall. So the coverage term collapses toward zero, not toward "most of the shortfall": the 9B's remaining gap is per-step cost on the steps that do split (blocking copies, same-stream ordering), not coverage and not the scheduler. The scheduler tuning itself is ce8f3e5. Fixture: qwen3.5-9b.mq4 md5 296092bf1e6a45d78c1acf815eb93366, prompt md5 5835c71e471849b4a72e1dc8e39695e7, gfx1201, passback share auto, spec off, greedy, 128 tokens, one fresh process per point. Daemon b35c25f032f80882d757512d0f8ee98a, hipfire ee24d2671dc3480ac8949a43df71b156. One run per point, host load uncontrolled; the numbers are within-run wall fractions, which are load-robust, but a second session would be needed to quote them across sessions. New: docs/perf-checkpoints/2026-10-02-offload-passback-coverage-phase0.md. Design doc 6.2.2 "Known estimator limitations" gains the pointer and the note that the SHARE_MAX ratchet is latent rather than a live bug (it self-heals when a wait reappears).
The four-arm A/B behind ce8f3e5 is a tie: every arm sits inside its own run-to-run spread (+-6% at 16 of 32 spilled) and the ordering flips between budgets, so the static fixture does not resolve REOPEN_EPS. The earlier 0.005 penalty (9c72282) did not reproduce. Recorded so the null is not re-run as if open, with the arm-order caveat stated. Arms b005c1/b005c3/b01c3/b02c3, 3 interleaved fresh-process rounds at 8 and 16 of 32 spilled, qwen3.5-9b.mq4 md5 296092bf1e6a45d78c1acf815eb93366, prompt md5 5835c71e471849b4a72e1dc8e39695e7, hipfire ee24d2671dc3480ac8949a43df71b156 held constant across arms.
Ledger rule: the originals are immutable, corrections are new dated files. coverage-phase0-amendment-1: - the trace prints a shape's line only at a power-of-two call count, so the 1024/2048/4096 calls are snapshots and the '8 per token' parenthetical (and the per-token wall figure derived from it) is withdrawn; the 3.7/3.7/6.3% headline is a ratio of two same-quantized walls and stands. - classify the one non-split shape: m=32 k=4096 = 68 KiB < MIN_SPLIT_BYTES, so it is size-refused on every step, not capacity-refused; no deficit conflation here. - [DERIVED] loop-closer: for m=12288 k=4096 the run's own rates give a balance floor of ~0.49 ms against a measured step wall of 0.52 ms and a balance point ~0.39 against a settled share 0.428, so the largest step is already at its balanced floor and the residual is the sync points. latch-band-ab-amendment-1: - 'Host load 3-6' was not measured in the 2026-10-02 session; it is carried from the 2026-10-01 split record. Replace with 'uncontrolled and not measured'. The tie is unaffected.
Amendment-1 3 computed the two-engine floor with the trace's gpu>=22.2 GB/s, which is a LOWER bound on the GPU rate, so m*rb/(r_cpu+r_gpu) is an UPPER bound on the achievable floor: the true floor is smaller and the slack larger than the '~5%, already at the floor' reading. State the direction and recover the rate from the converged balance (r_gpu = share*r_cpu/(1-share) ~ 25.6 GB/s): r_gpu 22.2 (lower bound) -> floor 0.494 ms -> slack 5.0% (lower bound on slack) r_gpu 25.6 (share-implied) -> floor 0.467 ms -> slack 10.2% (bound, not measured) So ~5-10% of the m=12288 step is scheduling slack, and d2h+join (0.11 ms, ~21%) is pure copy/sync. Conclusion unchanged: the residual is sync, not the share. The share-implied rate is circular for confirming placement and is labelled as a bound.
…se the chain Amendment-2's 'What survives' said 'at most a few percent is scheduling slack', contradicting its own table (5.0-10.2%). Corrected sentence states the residual is the d2h+join sync (0.11 ms ~ 21%) plus ~5-10% scheduling slack. Consolidated reading, with the bound directions re-checked: - coverage 3.7/3.7/6.3%, one size-refused shape -> not the deficit; - the floor built from gpu>= is an UPPER bound (<= ~0.494 ms), so the slack is >= ~5%; ~10% only under the converged-share estimate (r_gpu ~ 25.6 GB/s), which is an estimate, not a bound; - d2h+join = 0.11 ms ~ 21% is pure copy/sync. Conclusion unchanged: the residual is per-step sync cost, not share placement. Final word on this checkpoint.
…traceable The passback headline was split across two builds: 2026-10-01 split record +11.7%/+13.9% (built at f907988, pre-FFN-down and pre-latch) and fc5c072's +17.6% at 16 spilled (commit message only). Neither described HEAD. New checkpoint: cpu/passback/pcie x {8,16 of 32 spilled} x 4 interleaved fresh-process rounds, arm order rotated by round, identity + load in a header row. spilled cpu passback pcie passback/cpu pcie/cpu 8/32 28.65 34.10 23.90 +19.0% -16.6% 16/32 16.65 19.60 12.95 +17.7% -22.2% Consistent with the FFN-down co-inference landing on top of the old figure; the latch work is neutral per its own band A/B. Host load 3.78 -> 8.42 (1-min), uncontrolled and rising, so the absolutes are contended-host numbers and only the interleaved relative deltas are readable. HEAD 7443401, daemon b35c25f032f80882d757512d0f8ee98a, hipfire ee24d2671dc3480ac8949a43df71b156, model 296092bf1e6a45d78c1acf815eb93366, prompt 5835c71e471849b4a72e1dc8e39695e7. Also: changelog gains the current-HEAD A/B pointer and the `reopens` trace field.
… to pcie on a small model Multi-model end-to-end memory.offload_exec A/B on a spill-FRACTION grid so the models share an axis (2B 24 layers, 9B 32, 27B 64; 2B/9B at 25/50/75% spilled, 27B at 12.5/25%), 5 rounds per point for 2B/9B and 3 for 27B, arm order rotated by round, one fresh process per run. Median decode tok/s: model spilled % cpu passback pcie pb/cpu pb/pcie 2B 6/24 25.0 69.90 79.30 101.00 +13.4% -21.5% 2B 12/24 50.0 42.10 47.10 60.20 +11.9% -21.8% 2B 18/24 75.0 30.00 34.00 42.60 +13.3% -20.2% 9B 8/32 25.0 28.10 33.60 23.80 +19.6% +41.2% 9B 16/32 50.0 16.50 19.90 13.20 +20.6% +50.8% 9B 24/32 75.0 11.60 14.10 9.10 +21.6% +54.9% 27B 8/64 12.5 15.00 16.90 11.70 +12.7% +44.4% 27B 16/64 25.0 9.60 11.30 6.90 +17.7% +63.8% Every passback-cpu paired round positive (5/5 at each 2B/9B point, 3/3 at each 27B point). Findings: - 9B is the target regime: +19.6..21.6% over cpu, rising with the spill; pcie the loser. - On the 2B pass-back is the WRONG mode: pcie beats it at every spill fraction (+42..45% over cpu vs pass-back's +12..13%). The 2B's whole-model rates put the balance point ~0.59, above SHARE_MAX=0.50, so the scheduler cannot collapse to the GPU-only route. On the 9B/27B the balance point is 0.42..0.46, under the cap, and does not bind. - A bigger model does NOT give a bigger gain (25% spilled: 9B +19.6% vs 27B +17.7%); the gain is roughly a rate ratio, flat in model size. What model size changes is which route wins. Fixture: HEAD a7e3e1c, daemon b35c25f032f80882d757512d0f8ee98a, hipfire ee24d2671dc3480ac8949a43df71b156, prompt 5835c71e471849b4a72e1dc8e39695e7, models 9ed6628f/296092bf/d1292b4d. Load uncontrolled (2B 7.3->5.2, 9B 2.4->7.3, 27B 4.45->6.58), so absolutes are contended and only the interleaved relative deltas are readable. Also: figure + raw jsonl + chart script committed under data-2026-10-02-offload-passback-model-spread/ (the 2026-10-01 figure's source data lived in /tmp and is not reproducible); amends the 2026-10-02 mode-A/B record (withdraws an unsupported outlier cause, adds the paired per-round statistic); CHANGELOG bullet updated to the cross-model result.
…l-spread amendment New records and corrections from the 2B/9B/27B + mq3/mq4/mq6 pass-back sweeps. quant-spread-gfx1201.md (new): same-size, different-quant (9B, 8 and 16 of 32 spilled). pass-back vs cpu: mq3 +15.2%/+19.4%, 5/5 paired positive. Quant changes which route wins: on mq3 `pcie` moves to +3.0%/-1.8% over cpu where on mq4 it was -15.3%/-20.0%. mq6 is UNMEASURABLE on this build and it is a real coverage bug, not a missing kernel: the model's tensors are legacy MQ6G256; gemv_steps(MQ6G256, WithSwiGLUResidual) plans the UNFUSED [SiluMulRotate, GemvResidual] and dispatch_gemv_residual already wires MQ6G256 -> gemv_hfq6g256_residual (gemv.rs:622), but weight_gemv_swiglu_residual (hipfire-runtime/src/llama.rs:1543) takes the FUSED launcher, and dispatch_swiglu_residual (gemv.rs:637) has no MQ6G256 arm -> catch-all -> `unsupported gemv.swiglu_residual for /`. Plan-vs- fused routing mismatch. Exposed alongside: the key is declared/mapped/registered with no fused arm (coverage_tests only checks key<->table, so it cannot catch it); the `_ =>` error carries arch/quant "" so it names no dtype (why 30 runs and a code read were needed); and the registry advertises qwen3.5:9b-mq6 while the shipped build cannot decode it. model-spread amendment-1 (new): corrects four things in the committed record. (1) the SHARE_MAX claim is REFUTED by measurement — the 2B settles at share 0.268-0.471, under the cap, so the cap is not binding; why pass-back trails pcie on the 2B is open, not the cap. (2) the 27B's usable spill ceiling on this host is RAM-bound (24/64 spilled thrashes 28 GB), adding the 6.25% point: cpu 20.6, passback 22.5, pcie 17.8 (n=6 across two durable blocks), +9.0%. (3) the 9.5 outlier did not reproduce (re-run 22.6/22.6/23.5) and is kept VISIBLE in the committed data rather than excluded. (4) baseline-dependent framing: over cpu the gain is ~flat across size and quant; against the best single-engine route it grows with size because that route flips identity. Nondeterminism (design doc 6.2.2, CHANGELOG, AGENTS.md pitfalls): pass-back is NOT bit-reproducible. Measured on the 2B: cpu x2 byte-identical, pcie x2 byte-identical, passback x2 differ, HIPFIRE_OFFLOAD_PASSBACK_SHARE=0 byte- identical to cpu. Structural: the share is scheduled per process, so the row split and the mixed-engine numerics vary. cpu/pcie deterministic, passback not; pin the share in any gate that diffs output. The 2026-10-01 9B byte-identity held on its prompts, not as a general property. Figure data made durable and self-contained: sp_27b_extra/sp_27b_60 committed, spread-chart.py reads its own directory, both non-reproducing/outlier rows kept. Data: qwen3.5-9b.mq3 (sha256 c379dbbc..., 4.57 GB) and .mq6 (69b0e3b2..., 7.30 GB) pulled and sha256-verified.
added 5 commits
October 2, 2026 16:07
Follow-up to b0b4141 (kept as published rather than rewritten). - Output determinism stated three-part: the modes are not byte-interchangeable (cpu vs pcie differ -- different GEMV engines), within a mode at the same layer split the output is deterministic (cpu x2, pcie x2 byte-identical), and passback alone varies per run because its scheduled row split does. - 27B 6.25% point cites the durable blocks (cpu 20.6, passback 22.5, pcie 17.8, n=6) with the non-reproducing 9.5 kept visible; +9.0%. - mq6 reframed as a plan-vs-fused routing mismatch (gemv_steps prescribes the unfused route; the runtime takes the fused launcher; dispatch_swiglu_residual has no MQ6G256 arm), plus phantom coverage and the nameless-dtype error. - AGENTS.md pitfalls row, CHANGELOG bullet, design 6.2.2 updated to match.
Adds data-2026-10-02-offload-passback-quant-spread/quant-chart.py and
quant-vs-engine.{png,svg}, rendered from the committed jsonl (9b_mq3.jsonl +
sp_9b.jsonl). Panel (a) decode by engine for mq3 vs mq4 at 8/32 and 16/32
spilled; panel (b) pass-back's gain over cpu and over pcie, which shows the two
findings directly: on mq3 pcie is ~cpu (+12%/+22% gain over it) while on mq4 pcie
is far behind (+41%/+51%); pass-back's gain over cpu is lower on mq3
(+15%/+19%) than mq4 (+20%/+21%). Linked from amendment-1; no numbers change.
passback-scatter.{png,svg} + scatter-chart.py, rendered from the committed jsonl.
Four panels: (a) paired 1:1 passback-vs-cpu per run (every dot above the line
except the one visible 27B outlier); (b) gain vs spilled BYTES -- the curves do
not collapse, so bytes is not the sole driver; (c) gain vs spill fraction --
9B-arch peaks, 2B flat, 27B rises; (d) gain vs model depth at 25% spilled -- no
monotone size trend. Data only; no record numbers change.
…e exec-mode snapshot Two structural fixes for the partial-offload direction the maintainer set out in the review of the base PR: arch/dispatch code stays central, and the generic step seam owns the execution decision rather than each call site. - hipfire-runtime/src/llama.rs (weight_gemv_swiglu_residual): the CPU-executed arm hand-built a Step::GemvResidual and called offload_split::run_with directly, bypassing the seam. Hand the step to pipeline::execute_steps instead, so the seam decides cpu vs pass-back. host_mapped_cpu_capable gates on the CPU-exec predicate, so pcie never reaches this arm; the seam's plan_step/run_step path reproduces the old cpu route. - hipfire-dispatch/src/cpu_exec.rs: collapse the two cached LazyLock<bool> mode predicates (cpu_exec_enabled/passback_enabled) into a single cached HostExecMode snapshot; the predicates become thin views. Verification (gfx1201, qwen3.5-9b.mq4, 8/32 layers spilled, greedy, thinking off): the cpu arm's decoded output is identical between the pre-change (HEAD^) and current builds on a 160-token generation (sha256 57a42f72...), i.e. the seam swap preserved the cpu route. The pass-back arm reproduces the same text as cpu on the smoke prompt, and HIPFIRE_CPU_EXEC_TRACE shows the down-projection (gemv m=4096 k=12288 rotated residual) split by output rows under passback and run whole-CPU under cpu.
A cold kernel cache with parallel callers launched one hipcc per caller for the same module. Those redundant concurrent compiles are where the ROCm 7.2 clang-22 frontend crashed (parser stack fault, clang::Parser::ParsePragmaLoopHint -> SIGSEGV in Lexer::LexTokenInternal, tensor_ops.hip), failing 5 rdna-compute GPU tests in no-gpu-ci. One compile per cache key now, with the waiters re-checking the published blob and reusing it; distinct keys still compile in parallel, which is the compile_batch_for_symbols contract. Not fixed: the clang-22 ICE itself. It is intermittent and was not reproduced by hand; this removes the redundant compiles that were present in every failing run. Verified: cargo test -p rdna-compute --lib with HIPFIRE_KERNEL_CACHE pointed at a fresh dir, three runs, 0 failed (before the fix: 5 failed with the shared cache, 1 failed cold). ./scripts/no-gpu-ci.sh end to end on a cold cache: exit 0.
added 2 commits
October 2, 2026 18:45
scripts/check-crate-maps.py --check is the required "Crate maps match the tree" gate and it failed on PR warpfront#809: hipfire-arch-qwen35, hipfire-cli, hipfire-config, hipfire-dispatch, hipfire-runtime, hipfire-tui and rdna-compute had drifted -- two new modules (dispatch src/offload_split.rs, runtime src/offload_calibrate.rs), the new parking_lot dependency edge in rdna-compute, and the branch changed line/API/test counts. Regenerated with scripts/check-crate-maps.py <crate>; only the generated marker blocks changed, the hand-written prose is preserved byte for byte. --check now reports 45 maps matching the tree.
cargo clippy --workspace --all-targets fails on hipfire-xdna with "mutable borrow from immutable input(s)" at src/lib.rs:598 (clippy::mut_from_ref is deny-by-default). Pre-existing and unrelated to this branch -- the crate is untouched here and the code is identical on master -- but it is the only hard error in the advisory clippy job, so that job stays red until it is allowed. The &self receiver is deliberate: the BO mapping is interior-mutable by contract, the Arc keeps it alive, and the caller aliasing rule is documented on as_slice/as_mut_slice. Verified: cargo clippy -p hipfire-xdna --lib now exits 0 (4 warnings, no error).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Author's Note
This is another relatively large PR and deserves an explanation of its own. This PR directly depends on PR#793 - the partial CPU offload of qwen3.5 dense models.
When I was setting up the offload branch, I noticed that using the CPU to offload compute the layers that were within system RAM was noticeably faster (and it is what llama.cpp does) than just having the GPU access the layers over the PCIe bus. That remains true. But I also noticed that my GPU (a 9070 XT) did not hit 100% utilization while the CPU was computing the layers that were assigned to it. I don't think that this realization is going to surprise most people.
As a result I had a thought and decided to pursue it. The result is this PR, and a feature I am calling 'passback offload' mode.
Passback offload is scheduled co-inference of the layers in system RAM by both the GPU and CPU on a system at the same time. This works and is a speedup if a CPU is unable to saturate the bandwidth between it and the system RAM it is attached to. If there is bandwidth left over, there is sufficient leftover bandwidth for the GPU to read the layers in system RAM over the PCIe bus, and the GPU is fast enough to have some idle time while waiting for the CPU, then this results in an inference speedup of the offloaded layers. In essence, it is a hybrid mode between the existing 'pcie' and 'cpu' offload target modes.
On my machine, which meets all of those requirements, this results in up to a 21.6% speed increase over the 'CPU' offload target pushed in PR#793 depending on model size, and offloaded layer amount. Compared to the 'CPU' offload target, the speedup increases with layers offloaded while the overall speed decreases due to the number of layers offloaded. Better specific benchmarks are available in the PR.
There are a lot of requirements there, so it's worth talking about my personal PC for a moment - because my machine is relatively high end, but it fits a fairly common stereotypical mold.
Specifically, my PC is equipped with:
a) a 7800X3D,
b) 32GB of DDR5 running at 6000MT/s running @ CL30,
c) a B650 Motherboard, which limits my PCIe devices to gen 4.0, and
d) a 9070 XT 16GB.
Those of you who have been paying attention to the computer hardware scene, should note that this is a fairly typical gaming PC setup. It follows the trend of:
a) Fast CPU with less cores than the max available, but also a CPU that has either a high clock speed or in the case of the X3D chips, 3D V-Cache
b) Fairly large but not crazy large amount of decently fast system RAM.
c) A motherboard that may or may not support the latest PCIe generation
d) A GPU that is actually intended only for gaming.
I don't know for certain, but suspect that any PC in a similar "gaming PC shape" to mine will see a similar speedup.
If your PC is running a threadripper or xeon CPU with 96 cores, you probably want to use CPU inference. If you have a dual core CPU running a PCIe 5.0 x16 GPU, you probably want to use 'pcie' offload.
This does add complexity sadly, at the cost of having to have a basic 'scheduler' inside hipfire to allocate the compute of the offloaded layers between the existing 'cpu' and 'pcie' modes in the configuration dynamically. I need to be clear here that this scheduler keeps the layers offloaded to system RAM in system RAM, its only a question of whether the CPU computes inference, or the GPU reads the layer over the PCIe bus.
As before this targets only qwen35 dense models. As before the CPU engine is not byte identical. Further, because the scheduler is dynamic and tries to change the targeted offload amount all the time while running, behavior is not byte deterministic even between runs. By this I mean even though bit parity is not provided between the CPU and GPU runtimes, because passback dynamically allocates layers between the 'cpu' and 'pcie' offload modes output can change from specific run to run - whereas if you offloaded the same number of layers to the cpu runtime you should always get the same output as long as you kept the same number of layers allocated to the cpu runtime.
Summary
memory.offload_exec(envHIPFIRE_OFFLOAD_EXEC) gains a third value,passback, which runs a spilled layer's step on both offload engines concurrently — the GPU arm ispcie's launch restricted to output rows[0, g), the CPU arm iscpu's step restricted to[g, m). The newmemory.offload_passback_share(envHIPFIRE_OFFLOAD_PASSBACK_SHARE,autoby default) picksgfrom the host's measured engine rates and refines it online. The default stayspcie; a fully resident load changes nothing. A step that cannot be split runs wholly on the CPU (the schedule's degenerate point), so the mode is never less robust thancpu.Behavior, in one line: with
memory.offload_exec=passbackand a nonzero spill, the GPU stops idling through the host-executed steps and the two engines share each spilled step's output rows.Which surface(s) does this touch?
crates/hipfire-dispatch(newoffload_splitmodule; row-restricted launch views inpipeline/steps.rs;families/gemv.rsrow-stride helper),crates/rdna-compute(DTypegainsPartialOrd/Ordfor the per-(dtype, k)schedule map; and, from the bundled fix below, the JIT compile path incompiler.rs— a per-cache-key compile lock)crates/hipfire-config(OffloadExec::Passback,PassbackShare+ schema rule),hipfire-arch-qwen35load.rs(CPU-exec coverage line reports splittable layers + the stray-share notice),crates/hipfire-runtime/src/config.rs(retained-Redline exclusion coverspassback)crates/hipfire-runtime(llama.rs: the dense FFN down-projection is handed to the generic step seam; newoffload_calibrate.rs)hipfire-arch-qwen35(load-path coverage report + newtests/gpu_gemv_parity.rs)crates/hipfire-quantize/ quant formats — not touchedhipfire-cli(newhipfire offload-bench,--json/--write),hipfire-tui("Pass-back GPU share" row, exec-mode labels)Architecture-trait change? No.
crates/hipfire-runtime/src/arch.rsis untouched.What it is, precisely
pcie's launch restricted to rows[0, g)(launch_op_rows— a byte view of the weight plus a pointer-offset output; every covered kernel indexesA + row*row_strideand writesy[row]). The CPU arm iscpu's step restricted to[g, m)(cpu_arm_prepare/cpu_arm_finish). No third engine and no new numerics: the GPU arm is bit-identical row-for-row to a fullpcielaunch, andmemory.offload_passback_share=0is byte-identical tomemory.offload_exec=cpu(measured — it is the A/B twin).k*4down and(m-g)*4up (~30 KB at 9B shapes) against tens of MB of weight bytes.(dtype, k)seeds itself by timing both engines on that step's own weight buffer; every split step refines the share from its own arm timings. The join is read relative to the shape's own measured no-wait floor plus a hysteresis margin, so only the excess over that can be a wait for the GPU arm. Shares are clamped to[0.05, 0.50], adjusted every 4 steps, and a shape latches once converged — but the latch re-opens when the balance point moves (a livereopenscounter is on the trace line).Gemv{Raw},Gemv{Prerotated}, and either residual form when the dtype has a fused residual kernel. Dense qwen3.5 only; MoE/routed paths never reach it. Everything else runs on the CPU engine alone.pcie): not host-mapped; paddedrow_stride; weight byte length not exactlym * row_bytes(q, k); no vector row dot for the format; non-F32/short output; weight below 2 MiB; residual form with no fused GPU residual kernel.split: …line underHIPFIRE_CPU_EXEC_TRACE=1, which also prints the controller state (gpu samples,floor,waited,target,applied,reopens,frozen).hipfire offload-benchmeasures the host directly (one row per dense format:cpuGB/s,gpuGB/s, the share, the two-engine speedup bound), with--jsonand--writeto persist the recommendation. It is a raw device benchmark and refuses while a daemon pid file names a live process. Its measured limits — the shape matrix and the--writepin — are in "Exposed by this work" below.docs/plans/partial-gpu-offload-design.md§ 6.2.2. Config keys indocs/CONFIG.md/docs/env-vars.md(regenerated byscripts/check-lifecycle.py --write); changelog under Unreleased.Measured (gfx1201, RX 9070 XT, PCIe 4.0 ×16, Ryzen 7 7800X3D)
Read this as fixture-bound measurement, not a product default. Absolutes are
contended-host; only the interleaved relative deltas transfer. Protocol and
caveats:
docs/methodology/perf-benchmarking.md;each number lives in an immutable
historicalcheckpoint record (list below).cpupassbackpcieData:
docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201.md(+ amendment-1),…-quant-spread-gfx1201.md(+ amendment-1),…-mode-ab-gfx1201.md,…-split-gfx1201.md,…-headroom-idle.md.What the numbers say:
cpuon every model and quant measured (+9 % to +22 %), and every paired round was positive — so overcputhe gain is roughly flat across size and quant. What changes is which single-engine route it beats:pcieis the loser on the 9B/27B but the winner on the 2B, so the recommendedmemory.offload_execis size- and quant-dependent.d2h+joinis ≈ 21 % of the step wall. Refused (never-split) steps are only 3.7–6.3 % of CPU-executed wall — coverage is not the deficit (…-coverage-phase0.md+ amendments 1–3).cpustep's wall at 8/32 spilled, 89.8 % at 16/32, 94.6 % at 24/32, 66.0 % on a 27B at 8/64; a CPU GEMV stream and a GPU PCIe read of the same host-mapped bytes are near-additive to ~75–78 GB/s.Step::GemvResidual: measured +3.8 % at 8/32 and +2.6 % at 16/32 (9B, interleaved fresh-process pairs).…-latch-band-ab.md+ amendment-1). The mode's measured headline does not depend on it.Figures (committed):
passback-gain-vs-spill.pngpassback-scatter.pngquant-vs-engine.pngmode-sweep-vs-spilled-layers.pngDeterminism (please read before diffing output)
The modes are not byte-interchangeable, and
passbackis the only one that isnot run-to-run reproducible:
cpuandpcierun a layer's GEMV on different engines, so they need not agree — expected, not a bug (measured on the 2B:cpu2046 B vspcie3554 B). Within a mode at the same layer split the output is deterministic (cpu×2 andpcie×2 byte-identical).passbackschedules its row split per process (seeding probe + online arm timings), so two identical greedy invocations mix the engines differently. Pinmemory.offload_passback_sharein any gate that diffs pass-back output against a reference, or use0to reach the exactcputwin.AGENTS.md's pitfall table: never assert a greedy/golden reference across a mode pair.Exposed by this work, not fixed here
qwen3.5:9b-mq6does not decode in a spill configuration, for a reason unrelated to pass-back: its tensors are legacyMQ6G256, whose plan is the unfused[SiluMulRotate, GemvResidual], butweight_gemv_swiglu_residualtakes the fused launcher anddispatch_swiglu_residualhas noMQ6G256arm — the catch-all fires and every mode (passback/cpu/fully resident) fails withunsupported gemv.swiglu_residual for /. Flagging it for a separate fix; the pass-back PR does not touch it.min_join_nsis a monotone running minimum that never rises (a contention-lifted floor biases the balance point down);r_gpuis refreshed only on a waited join, so a shape whose joins sit at the floor carries a stale GPU-rate anchor. Neither is where this fixture's gap is: the 2026-10-01 investigation measured the share sitting where its probe said it should in that regime (docs/investigations/2026-10-01-offload-passback-perf-gap.md§ 5), and the skipped-step wall term is 3.7–6.3 %. (That regime's probe agreement does not carry to the defaultoffload-benchmatrix on today's host — see the next bullet.)hipfire offload-bench's default shape matrix does not represent a real projection, and--writeis a hard pin. The probe times each engine alone on a syntheticStep::Gemv{Raw}shape —k=5120,mderived from--buffer-mb(192 MiB → m≈37k–97k rows), never a real projection and neverStep::GemvResidual— so its recommendation is a solo-regime bandwidth ratio. Measured on gfx1201, 2026-10-02: it recommends 0.331 formq4g256(mq4g256v20.334,mq4cg2560.331,mq6g2560.395,mq5g256v20.450,q8_00.313), while the scheduler's own per-shape values on the same host are 0.489–0.492 for the non-residual shapes and 0.353 for the FFN down-projection;--k 4096 --buffer-mb 4reproduces the live number (0.494). Two further notes on that run: the recommendation is taken over the five formats that win a ~1.6 KB tie among equal 192 MiB measurements (median over all nine: 0.345), and the per-format number moves ~±0.05 run to run. Because--writepins the value for every shape (offload_split.rs:Share(f) => fskips theAutobranch where the seeding probe and the controller run; the pin is only recorded for the trace) — unlikeseed(), which installs a fallback the controller still refines — writing that default here would hold the non-residual steps ~0.16 low.autois the right setting on this host;--writeis not. A realistic default shape, a residual probe mode, and a--writewarning are separate follow-up work, not part of this PR.Test plan
./scripts/no-gpu-ci.shpasses — after the bundled JIT fix below; the post-fix run was end-to-end withHIPFIRE_KERNEL_CACHEpointed at a fresh directory (exit 0, 285 s). Before the fix, a cold-cache run failed 5rdna-computeGPU tests on aclang-22frontend crash (details and fix in "Bundled CI fix"); everything else in the script passed on the first run (the script exits at the first failing step, so those steps were run individually and are green:hipfire-cpu26,hipfire-arch-qwen35 moe_prefill20, config/registry/client/cli/tui 165, pytest 388+6 subtests, redline unittest 179,test_install_revision.py,test_uninstall.py,check-env-docs.py,check-lifecycle.py268 keys / 1356 env vars).cargo check --workspace --examplesclean (first step ofno-gpu-ci.sh);cargo build --releaserun to completion (47 s incremental, warnings only, no errors).cargo test --lib --workspaceas a whole was not run —no-gpu-ci.shruns the no-GPU subset (rdna-compute,hipfire-cpu,hipfire-arch-qwen35 moe_prefill, config/registry/client/cli/tui) and that subset is green. Flagging the gap rather than claiming the wider command.scripts/serve_harness.pyon hardware myself for the exact model and settings under test, and attached the per-turn JSON below (passback battery + chain, pluscpuand fully-resident controls on the same fixture).kv_backend: "vmm",kv_backend_legacy: false,kv_backend_reason: null. No legacy token was used.serve_harnessis ticked as user-facing serve semantics only, perdocs/VALIDATION.md's claim→route map — it is explicitly not numerical/state parity evidence, and no cross-mode golden-output assertion is possible on this path:cpuandpcierun different engines, andpassbackschedules its row split per process, so two identical invocations mix the engines differently. What the offload path's parity rests on instead: (a)memory.offload_passback_share=0is byte-identical tocpu(measured, and the A/B twin), (b) the GPU arm is bit-identical row-for-row to a fullpcielaunch by construction — the same kernels over a row range, no new numerics (the design doc's claim; unlike (a) it is not a separate measured byte-diff), and (c)crates/hipfire-arch-qwen35/tests/gpu_gemv_parity.rs— the GPU↔CPU GEMV parity matrix for the host-mapped path, new in this branch,#[ignore]d because it needs a GPU../scripts/speed-gate.sh— not run. Its locked baselines (tests/speed-baselines/gfx1201.txt) were captured on a 4× R9700 / Threadripper host; this host is a single RX 9070 XT, so a comparison would report the hardware, not the change. The perf claim here is fixture-bound and lives in the checkpoints above; the resident control in the harness evidence is the regression smoke for the shared step seam.scripts/leanup-thresholds.txtis raised by this branch (the file is untouched), so noratchet-raiselabel is owed.local serve_harness harness output (load / serve / kernel changes)
Fixture:
~/.hipfire/models/qwen3.5-9b.mq4(md5296092bf1e6a45d78c1acf815eb93366), greedy,--speculation off --thinking off, spill viaHIPFIRE_GPU_LAYER_BUDGET=24(24 resident of 32 → 8 spilled); the
passback/cpuarms addHIPFIRE_OFFLOAD_EXEC, the resident control runs the same model with the budgetunset. Daemon md5
d8e21f57834240b7607dba80546be863;hipfiremd531c057d13c637446aafab0685e8035e6— the build before the bundled JIT fixbelow (post-fix build: daemon
80e52b2c3655a55e051f783563491729,hipfire7bc03ee5898c28abfd13c43dde42f551). The fix is behavior-neutral for thisevidence: it only serializes duplicate compiles of one cache key and touches no
forward-path code.
--mode battery,HIPFIRE_OFFLOAD_EXEC=passback:[ { "request_id": "chatcmpl-846787-1", "finish": "stop", "gen": 343, "decode_tok_s": 33.2, "attractor": false, "empty": false, "runaway": false, "retrieval_missing": [], "expected_substrings": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_reason": null, "prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f", "request_md5": "b2d411eb58ec2648ebd087ac828a51bb", "ans_preview": "```python\ndef merge_sorted(a, b):\n \"\"\"\n Merge two already-sorted lists into one sort" }, { "request_id": "chatcmpl-846787-3", "finish": "stop", "gen": 286, "decode_tok_s": 33.6, "attractor": false, "empty": false, "runaway": false, "retrieval_missing": [], "expected_substrings": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_reason": null, "prompt_md5": "640e0fd4f55996cb175a422f0a12cef5", "request_md5": "85b33e9aedabf7a9fa0d1d81cc6f8356", "ans_preview": "To find the total distance traveled, we need to calculate the distance for each segment of" }, { "request_id": "chatcmpl-846787-5", "finish": "stop", "gen": 81, "decode_tok_s": 33.5, "attractor": false, "empty": false, "runaway": false, "retrieval_missing": [], "expected_substrings": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_reason": null, "prompt_md5": "8f66b4c97988825bd8e7840aaf44357e", "request_md5": "d34383b66fd5f7af0cfced19d3777983", "ans_preview": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its or" }, { "request_id": "chatcmpl-846787-7", "finish": "stop", "gen": 102, "decode_tok_s": 33.6, "attractor": false, "empty": false, "runaway": false, "retrieval_missing": [], "expected_substrings": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_reason": null, "prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817", "request_md5": "643a78544be2dab1edcfe5b8e85c3718", "ans_preview": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against t" }, { "request_id": "chatcmpl-846787-9", "finish": "stop", "gen": 82, "decode_tok_s": 33.5, "attractor": false, "empty": false, "runaway": false, "retrieval_missing": [], "expected_substrings": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_reason": null, "prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4", "request_md5": "9af4a6122e824b100c1d31cb525d7816", "ans_preview": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep" } ]The
cpuand fully-resident controls run the same battery prompts as thepassback battery above, turn for turn — their
request_md5values match itexactly, which is what makes the three battery arms comparable.
chainis therelated-turn mode, so from turn 2 onward it carries prior turns into the request
and its
request_md5values legitimately differ (efde613f…at turn 2 vs thebattery's
85b33e9a…).chaincpubatteryDecoded text read: coherent on all four arms. On this fixture
passbackreproduced the
cpuand fully-resident text exactly — all five batteryturns byte-equal (
assistant_contentdiffed turn-by-turn, 5/5 against each arm),and the same token counts (343 / 286 / 81 / 102 / 82). That is a property of this
fixture's argmax, not a guarantee:
passback's row split is scheduled perprocess, which is why the determinism section above forbids building a
golden-output gate on it.
passback33.5 vscpu28.0 tok/s here is +19.6 % —the same delta as the 8/32 row of the checkpoint table. Full per-turn
JSON (every field) for all four arms was kept at
/tmp/passback-pr/{pbn_battery, pbn_chain,cpuN_battery,residentN_battery}.jsonon the author's host.Bundled CI fix (independent of pass-back) — the reviewer may prefer this split out
rdna-compute's JIT path launched onehipccper caller for the samemodule whenever several threads needed it on a cold kernel cache (the
rdna-computelib test binary runs its GPU tests on a thread pool, and everytest builds its own
Gpu). Every failing run had exactly that condition, andthe ROCm 7.2
clang-22frontend crashed in it — a parser stack fault,clang::Parser::ParsePragmaLoopHint→ SIGSEGV inclang::Lexer::LexTokenInternal, current parser token'xa'attensor_ops.hip:1332insidehyper_norm_gate_f32. The source file on disk isintact and byte-identical between the failing and the passing cache
(
md5 37119ab9…, 160694 B), so this is not a torn read of the shared{stem}.hip; it is the compiler crashing, intermittently, while identicalcompiles ran.
compiler.rsnow takes a process-global lock per cache key across thecompile (
compile_lock), and a waiter re-checks the published blob and reusesit instead of compiling the same source again. Distinct keys still compile in
parallel — that is
compile_batch_for_symbols's contract — so this is not aglobal compile lock. It removes the redundant concurrent compiles (5 identical
hipccruns become 1) that were present in every failing run.The compiler ICE itself is not claimed fixed — it is intermittent, was not
reproduced by hand, and the backtrace below is worth reporting to LLVM/ROCm.
Evidence —
cargo test -p rdna-compute --lib(400 tests,HIPFIRE_KERNEL_CACHEpointed at a fresh directory so every module JIT-compiles):
scripts/no-gpu-ci.sh, the reported failure)scripts/no-gpu-ci.shend to end, cold cachecargo test -p rdna-compute --lib compiler::compile_lock_is_per_cache_key_and_mutually_exclusive)Three clean runs do not prove an intermittent compiler ICE is gone — they are
the strongest evidence available on one host. What is proven is the
deterministic half: one compile per cache key, and the new unit test asserts the
lock is per-key and mutually exclusive.
clang-22 crash backtrace (first frames)
Hardware validation request (optional)
Deliberately omitted. The claim needs
HIPFIRE_OFFLOAD_EXEC=passbackplus amemory.gpu_layer_budgetspill, and the<!-- hw-gate-request -->block'sroutes[]schema cannot express env for a route — a hw-gate run on the defaultfixtures would exercise
pcie, not this mode, and report a non-reproduction thatis not a result. Local
serve_harnessoutput is attached instead. If amaintainer wants it on automation, the route is:
qwen3.5:9bbattery + chainwith
HIPFIRE_OFFLOAD_EXEC=passback HIPFIRE_GPU_LAYER_BUDGET=24(fixture:~/.hipfire/models/qwen3.5-9b.mq4, md5296092bf1e6a45d78c1acf815eb93366).How this merges (direct review)
Merge authority is direct maintainer review plus the required no-GPU CI
checks. hw-gate automation is optional evidence delivery, not a prerequisite or
a substitute.
build (workspace, no GPU),unit tests (lib, no GPU),gates (ratchets, layering, registers)from.github/workflows/ci.yml.docs/VALIDATION.md; the localserve_harnessoutput and the perf checkpoints are attached/linked above.master.scripts/coherence-gate*.shandtools/change_gate/ agentic-review are historical only — never acceptance evidence.Architecture-trait change?
No.
crates/hipfire-runtime/src/arch.rsis untouched; noArchitecturetraitsurface changes, so nothing ripples to the arch crates beyond
hipfire-arch-qwen35's own load-path coverage line.