Skip to content

Place each operator on the cores that execute it - #160

Open
asyms wants to merge 5 commits into
feat/conv-block-mappingfrom
fix/operator-placement
Open

asyms wants to merge 5 commits into
feat/conv-block-mappingfrom
fix/operator-placement

Conversation

@asyms

@asyms asyms commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #159.

FuseMax's 256x1 vector core declared no operator_types, so core selection treated it as a second generic core and split every conv evenly between it and the 256x256 array. Each conv then ran at the pace of the vector core.

  • fusemax_vec.yaml lists the elementwise, reduction and pooling ops it serves, the same set as the TPU7x VPU. Convs and matmuls run on the array. The accelerator is renamed from quad_core to fusemax.
  • With unequal generic cores, the generator keeps the largest-array cores whose even split finishes first: n cores of at least u units deliver n * u MACs per cycle.
  • A fixed mapping that places an operator on a core whose operator_types exclude it is rejected when it is read.
  • Silu, Sigmoid and Exp now get their 4 cycles per op on the ZigZag path. The match compared the node type against lowercase names.
  • The utilization report divides a node's ideal cycles by the cycles of one call rather than by its per-iteration latency, so a node idle on some iterations no longer reports an efficiency above 1. A node counts as a fallback when its estimator priced it at ideal cycles, not whenever it has no ZigZag evaluation, so kernel-library costs are no longer flagged.

The FuseMax conv-window scenarios that split convs over the vector core are replaced by a conv on the array feeding a max pool on the vector core through the shared memory. The shared-memory test is re-derived for this placement: apart, each core fits its tiles in 133 KB, while sharing one memory they need 136 KB, so 134 KB is infeasible.

Fast suite: 907 passed. The slow FuseMax, generic and ResNet tests fail in the same two cases as on #159.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 7 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[fusemax], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware (⚠ 1 flagged)

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 149684492 +0.00% 52% 26%
eyeriss_like_quad_core 101581557 → 101581557 +0.00% 52% 19%
eyeriss_like_single_core 303038604 → 303038604 +0.00% 52% 25%
fusemax 233766951 → 510918663 ⚠ ↑+118.56% 6.2% 0.75%
meta_prototype 101711933 → 101711933 +0.00% 100% 6.2%
simba 125344507.60 → 125344507.60 +0.00% 65% 1.0%
simba_small 106627594 → 106627594 +0.00% 65% 3.0%
tpu_like_quad_core 101580877 → 101580877 +0.00% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 6 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 74761 ⚠ ↑+0.17% 51% 23%
eyeriss_like_quad_core 39917 → 40209 ⚠ ↑+0.73% 51% 22%
eyeriss_like_single_core 115003 → 115003 +0.00% 68% 31%
fusemax 187612 → 22546 ⚠ ↓-87.98% 1.6% 0.48% low array utilization
meta_prototype 15341 → 15439 ⚠ ↑+0.64% 51% 19%
simba 4610.67 → 4610.67 +0.00% 50% 4.9%
simba_small 8706 → 8998 ⚠ ↑+3.35% 50% 16%
tpu_like_quad_core 12300 → 12592 ⚠ ↑+2.37% 42% 11%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant