Skip to content

Land the allocator and AIE codegen changes IRON's stream-dse operators need - #137

Merged
asyms merged 67 commits into
mainfrom
iron-integration
Oct 4, 2026
Merged

asyms merged 67 commits into
mainfrom
iron-integration

Conversation

@asyms

@asyms asyms commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

What IRON's stream-dse operators (amd/IRON#227) need from stream, on top of main.

  1. Allocator plumbing: model expressions registered as named quantities, SolverModel.value, and optional constraint families selected per solve through SolveOptions.families, validated when the options are resolved.
  2. A registry of the top-level memory ports of each core (AIE tile DMA channels, ZigZag memory ports, measured per-port bandwidth) and the memory_ports family. With interval and burst off it only reports, which is how IRON runs it: per-port, shared-bandwidth and link activity in the performance view.
  3. Transfer targets credited with the tile their consumer core receives.
  4. AIE model and codegen fixes found by checking stream's estimates against NPU2:
    • kernel dimensions addressed by operand axis, called once per leading batch index, and a matmul batch axis of one indexed by the node's batch dimension
    • a loop around a re-reading shim transfer chained into one task where its buffer descriptors fit
    • a window an outer loop moves on may be held in one buffer, and each run is charged the fill of the off-chip tensors it holds that way
    • fifos deepened as far as their tile has slack, a broadcast's readers together
    • a transferred tensor placed only where its chosen path delivers it, and a loop of one kept out of a tile's local axes
    • each call linked to the object built for the tile it takes
    • tile DMA channels at 64 bits a cycle, the reconfiguration per column calibrated to IRON's dispatch, and a design whose steady state is one unsplit iteration
    • matmul_PV passed the key length mlir-aie 1.4.4.dev73's kernel reads
  5. Dead code removed: the ComputeAllocator module, unused constraint profiles, visualizations, allocator helpers and accessors.

The fast suite passes. With IRON's branch on mlir-aie 1.4.4.dev73, every MHA and SwiGLU stream-dse operator test passes on an NPU2 (85 tests, extensive included).

asyms added 30 commits October 3, 2026 22:48
asyms added 24 commits October 3, 2026 22:49
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 15 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[eyeriss_like_dual_core], hardware_swiglu[eyeriss_like_quad_core], hardware_swiglu[eyeriss_like_single_core], hardware_swiglu[fusemax], hardware_swiglu[meta_prototype], hardware_swiglu[simba], hardware_swiglu[simba_small], hardware_swiglu[tpu_like_quad_core], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 1847bde900d7f4688ec06e9712148aaed915db03 | date 2026-08-18 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware (⚠ 8 flagged)

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 170918067 → 149684492 ⚠ ↓-12.42% 52% 26%
eyeriss_like_quad_core 127074661 → 101581557 ⚠ ↓-20.06% 52% 19%
eyeriss_like_single_core 343933021 → 303038604 ⚠ ↓-11.89% 52% 25%
fusemax 76808227 → 233766951 ⚠ ↑+204.35% 9.3% 0.26%
meta_prototype 110100501 → 101711933 ⚠ ↓-7.62% 100% 6.2%
simba 38666519 → 125344507.60 ⚠ ↑+224.17% 65% 1.0%
simba_small 123142418 → 106627594 ⚠ ↓-13.41% 65% 3.0%
tpu_like_quad_core 103546916 → 101580877 ⚠ ↓-1.90% 97% 3.1%
hardware_two_conv — 8 hardware (⚠ 7 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 76675 → 74631 ⚠ ↓-2.67% 51% 24%
eyeriss_like_quad_core 41961 → 39917 ⚠ ↓-4.87% 51% 22%
eyeriss_like_single_core 114999 → 115003 ↑+0.00% 68% 31%
fusemax 186148 → 187612 ⚠ ↑+0.79% 1.4% 0.05% low array utilization
meta_prototype 17385 → 15341 ⚠ ↓-11.76% 51% 19%
simba 4316 → 4610.67 ⚠ ↑+6.83% 50% 4.9%
simba_small 10750 → 8706 ⚠ ↓-19.01% 50% 17%
tpu_like_quad_core 14344 → 12300 ⚠ ↓-14.25% 42% 12%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

@asyms
asyms merged commit 32b8eaa into main Oct 4, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant