Skip to content

Give a carried state's input edge no transfer, so recurrences with an initial state lower - #152

Merged
asyms merged 1 commit into
mainfrom
fix/carried-state-inputs
Oct 6, 2026
Merged

asyms merged 1 commit into
mainfrom
fix/carried-state-inputs

Conversation

@asyms

@asyms asyms commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

A recurrence whose carried state comes in as a workload input (the Mamba state-update block's h_prev, the linear-attention state) failed to lower with StopIteration in update_tensor_ssis. The carried state already gets no transfer, since it stays resident on its node's cores, but its input edge was kept in the steady-state graph with nothing to feed, and the iteration-space pass expects every input edge to have a successor.

build_transfer_graph now leaves out an input edge that only delivers carried state. The state keeps the iteration space of the node that holds it, as before.

Test: the Mamba and linear-attention catalog blocks solve end to end on the TPU-like quad core, with no transfer and no input edge for their state. Both fail without the change.

Stacked on #151.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 5 cell(s) flagged (total_latency > 0.1% tol): hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 149684492 +0.00% 52% 26%
eyeriss_like_quad_core 101581557 → 101581557 +0.00% 52% 19%
eyeriss_like_single_core 303038604 → 303038604 +0.00% 52% 25%
fusemax 233766951 → 233766951 +0.00% 9.3% 0.26%
meta_prototype 101711933 → 101711933 +0.00% 100% 6.2%
simba 125344507.60 → 125344507.60 +0.00% 65% 1.0%
simba_small 106627594 → 106627594 +0.00% 65% 3.0%
tpu_like_quad_core 101580877 → 101580877 +0.00% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 5 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 74761 ⚠ ↑+0.17% 51% 23%
eyeriss_like_quad_core 39917 → 40209 ⚠ ↑+0.73% 51% 22%
eyeriss_like_single_core 115003 → 115003 +0.00% 68% 31%
fusemax 187612 → 187676 ↑+0.03% 1.4% 0.05% low array utilization
meta_prototype 15341 → 15487 ⚠ ↑+0.95% 51% 19%
simba 4610.67 → 4610.67 +0.00% 50% 4.9%
simba_small 8706 → 8998 ⚠ ↑+3.35% 50% 16%
tpu_like_quad_core 12300 → 12592 ⚠ ↑+2.37% 42% 11%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

@asyms
asyms changed the base branch from fix/tile-search-group-artifacts to main October 6, 2026 07:44
@asyms
asyms force-pushed the fix/carried-state-inputs branch from f5870b9 to 651ea69 Compare October 6, 2026 07:44
@asyms
asyms merged commit b79e6b8 into main Oct 6, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant