Skip to content

Place and size each copy of a multicast with the node it reaches - #156

Open
asyms wants to merge 1 commit into
fix/single-core-footprintfrom
fix/multicast-copy-placement
Open

asyms wants to merge 1 commit into
fix/single-core-footprintfrom
fix/multicast-copy-placement

Conversation

@asyms

@asyms asyms commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

A transfer that multicasts a tensor to readers on different cores gave every copy the transfer's whole set of destination cores, and sized every copy with the transfer's own inter-core tiling. So the copy for a reader left on one core (a residual Add on a vector core, a softmax exp) was also reserved on the cores of the other readers, and was sized as one slice of the other readers' split instead of the whole tile it reads.

  • DecisionSpace narrows each copy's placements to the cores of the node it reaches; a copy for another transfer or an edge keeps the transfer's.
  • A chosen route only requires a copy on the targets where that copy can live.
  • Workload.get_tensor_single_core sizes a copy that feeds a computation node directly with that node's split; a copy staged for another transfer keeps the transfer's tiling. sliding_halo returns no halo for a copy its reader reads without a window.

Tests: in the attention block on the TPU-like quad core, each copy of a multicast can only be placed with its reader, and the copy for the unsplit exp holds and receives the whole tile. Both fail without the change.

Stacked on #155.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 5 cell(s) flagged (total_latency > 0.1% tol): hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 149684492 +0.00% 52% 26%
eyeriss_like_quad_core 101581557 → 101581557 +0.00% 52% 19%
eyeriss_like_single_core 303038604 → 303038604 +0.00% 52% 25%
fusemax 233766951 → 233766951 +0.00% 9.3% 0.26%
meta_prototype 101711933 → 101711933 +0.00% 100% 6.2%
simba 125344507.60 → 125344507.60 +0.00% 65% 1.0%
simba_small 106627594 → 106627594 +0.00% 65% 3.0%
tpu_like_quad_core 101580877 → 101580877 +0.00% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 5 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 74761 ⚠ ↑+0.17% 51% 23%
eyeriss_like_quad_core 39917 → 40209 ⚠ ↑+0.73% 51% 22%
eyeriss_like_single_core 115003 → 115003 +0.00% 68% 31%
fusemax 187612 → 187676 ↑+0.03% 1.4% 0.05% low array utilization
meta_prototype 15341 → 15487 ⚠ ↑+0.95% 51% 19%
simba 4610.67 → 4610.67 +0.00% 50% 4.9%
simba_small 8706 → 8998 ⚠ ↑+3.35% 50% 16%
tpu_like_quad_core 12300 → 12592 ⚠ ↑+2.37% 42% 11%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant