Skip to content

Judge an AIE matmul against its MAC unit's peak - #161

Merged
asyms merged 1 commit into
mainfrom
fix/aie-matmul-peak
Oct 9, 2026
Merged

asyms merged 1 commit into
mainfrom
fix/aie-matmul-peak

Conversation

@asyms

@asyms asyms commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Stacked on #160.

The AIE estimator took a matmul's ideal cycles from a fixed rate of 32 bf16 ops per cycle per tile. A measured call of the library's matmul runs at about 150 MACs per cycle, so the reported efficiency came out above 1.

A matmul kernel's ideal now comes from the kernel library's MAC tile, one of which the matmul unit computes per cycle. For the AIE2P library that is 8x8x8: the bfp16 mmul<8,8,8> the bf16 matmul issues, 512 MACs per cycle, the same rate as int8 on that tile. On AIE2 the native bf16 tile is 4x8x4, 128 MACs per cycle. Vector kernels, and nodes without a library, keep the per-datatype rate. A measured call is still what prices the node.

Test: tests/unit/test_aie_cost_estimator.py. Fast suite: 911 passed.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 7 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[fusemax], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware (⚠ 1 flagged)

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 149684492 +0.00% 52% 26%
eyeriss_like_quad_core 101581557 → 101581557 +0.00% 52% 19%
eyeriss_like_single_core 303038604 → 303038604 +0.00% 52% 25%
fusemax 233766951 → 510918663 ⚠ ↑+118.56% 6.2% 0.75%
meta_prototype 101711933 → 101711933 +0.00% 100% 6.2%
simba 125344507.60 → 125344507.60 +0.00% 65% 1.0%
simba_small 106627594 → 106627594 +0.00% 65% 3.0%
tpu_like_quad_core 101580877 → 101580877 +0.00% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 6 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 74761 ⚠ ↑+0.17% 51% 23%
eyeriss_like_quad_core 39917 → 40209 ⚠ ↑+0.73% 51% 22%
eyeriss_like_single_core 115003 → 115003 +0.00% 68% 31%
fusemax 187612 → 22546 ⚠ ↓-87.98% 1.6% 0.48% low array utilization
meta_prototype 15341 → 15439 ⚠ ↑+0.64% 51% 19%
simba 4610.67 → 4610.67 +0.00% 50% 4.9%
simba_small 8706 → 8998 ⚠ ↑+3.35% 50% 16%
tpu_like_quad_core 12300 → 12592 ⚠ ↑+2.37% 42% 11%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

@asyms
asyms force-pushed the fix/aie-matmul-peak branch from d5e7d8b to ab2a786 Compare October 9, 2026 10:59
@asyms
asyms changed the base branch from fix/operator-placement to main October 9, 2026 10:59
@asyms
asyms merged commit 56f44cd into main Oct 9, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant