Skip to content

Give readers that index a tensor along unrelated dims their own transfer, so self-attention keeps its query and key axes apart - #154

Open
asyms wants to merge 1 commit into
mainfrom
fix/multicast-reader-axes
Open

asyms wants to merge 1 commit into
mainfrom
fix/multicast-reader-axes

Conversation

@asyms

@asyms asyms commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

The steady-state lowering moves each tensor to all its readers with one multicast transfer, and a transfer copies along identity maps, so every reader's index dims become the transfer's. When two readers index the tensor along unrelated unique dims, that merges them. In a self-attention the query projection and the key/value projections both read the input rows, so the query and key axes of the scores collapsed into one: splitting the queries over four cores also split the keys, and each core computed only its diagonal block of the scores (four 32x32 tiles of a 128x128 matrix).

_reader_walks now groups a tensor's readers so that, at every index of the tensor, the readers in a group share a unique dim, and each group gets its own transfer, named after its first reader (Transfer(x for proj_q)). Readers that slide different windows along the same axis, like a 3x3 and a 5x5 conv, still share one copy. A tensor with one group lowers exactly as before.

DecisionSpace.copied_bits sums what every transfer reading a tensor delivers, so off-chip traffic counts both copies.

Test: in the attention catalog block, the scores' query and key dims stay distinct in the steady state and the input gets one transfer per walk. It fails without the change. The rest of the fast suite is unchanged.

…fer, so self-attention keeps its query and key axes apart
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 5 cell(s) flagged (total_latency > 0.1% tol): hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 149684492 +0.00% 52% 26%
eyeriss_like_quad_core 101581557 → 101581557 +0.00% 52% 19%
eyeriss_like_single_core 303038604 → 303038604 +0.00% 52% 25%
fusemax 233766951 → 233766951 +0.00% 9.3% 0.26%
meta_prototype 101711933 → 101711933 +0.00% 100% 6.2%
simba 125344507.60 → 125344507.60 +0.00% 65% 1.0%
simba_small 106627594 → 106627594 +0.00% 65% 3.0%
tpu_like_quad_core 101580877 → 101580877 +0.00% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 5 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 74761 ⚠ ↑+0.17% 51% 23%
eyeriss_like_quad_core 39917 → 40209 ⚠ ↑+0.73% 51% 22%
eyeriss_like_single_core 115003 → 115003 +0.00% 68% 31%
fusemax 187612 → 187676 ↑+0.03% 1.4% 0.05% low array utilization
meta_prototype 15341 → 15487 ⚠ ↑+0.95% 51% 19%
simba 4610.67 → 4610.67 +0.00% 50% 4.9%
simba_small 8706 → 8998 ⚠ ↑+3.35% 50% 16%
tpu_like_quad_core 12300 → 12592 ⚠ ↑+2.37% 42% 11%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant