Skip to content

Report the latency the solve models, held to every core, link and memory port - #163

Open
asyms wants to merge 1 commit into
feat/ir-tensor-dimsfrom
fix/throughput-bound-transfers
Open

asyms wants to merge 1 commit into
feat/ir-tensor-dimsfrom
fix/throughput-bound-transfers

Conversation

@asyms

@asyms asyms commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #162.

The cycles a group reported were not the latency the solve modelled:

  • Lone nodes reported a throughput bound computed after the solve that counted only how long each core computes. A node that streams its operands was reported at its compute time. On the TPU-like quad core, a 256 x 8192 x 2048 matmul reported 2.15 M cycles while its DRAM port alone needs 12.7 M.
  • Fused groups reported the latency objective level, which summed the modelled latency with two charges:
    • the off-chip traffic over the off-chip link bandwidth, which the overlap already times through its off-chip contention bound, so it was counted twice;
    • the DMA channel peaks, a count added to cycles.
  • Memory port rates bounded nothing unless a caller selected memory_ports, so a solve could run a DRAM port above its rate.

Now:

  • memory_ports is a default family, with its interval and burst bounds on, so every iteration is held to its busiest memory port inside the MILP. It can still be set to report only through default_families(options=...).
  • The latency objective level is the modelled latency alone.
  • The DMA peaks get their own level after the off-chip traffic.
  • The off-chip charge applies only when no family times the off-chip links: the overlap registers offchip_timed when its off-chip contention is on.
  • Every group, lone or fused, and the tile search report the solved total_latency. The throughput bound computed after the solve is removed.

The matmul above now solves to 12.72 M cycles. That is its objective, its reported cycles and its DRAM port busy time, with the port at 100 % of the interval.

Tests: tests/unit/test_reported_latency.py, plus the objective-level, family-order and memory-port tests updated to the new levels and defaults. Fast suite: 914 passed. The slow suite fails in the same 10 tests as its base.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Stream AIE Metrics Regression Guard

⚠ 14 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[eyeriss_like_dual_core], hardware_swiglu[eyeriss_like_quad_core], hardware_swiglu[fusemax], hardware_swiglu[meta_prototype], hardware_swiglu[simba], hardware_swiglu[simba_small], hardware_swiglu[tpu_like_quad_core], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core]

16 of 16 cells captured

Provenance: baseline 32b8eaa94931143f86c3a135c774febd1d8b96ae | date 2026-10-04 | Python 3.12.3 | backend ortools_gscip
Note: mip_gap: null — OR-Tools GSCIP
Columns: array fill = per-layer PE-array spatial fill (dataflow quality); MAC eff (e2e) = useful MACs / (chip peak MACs/cycle × total latency), the true fraction of the chip's compute used (incl. idle cores, temporal stalls & transfers).

hardware_swiglu — 8 hardware (⚠ 7 flagged)

seq_len=256, embedding_dim=2048, hidden_dim=8192, bf16; layer-fused tiles seq=16/embedding=128/hidden=32

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 149684492 → 201589115 ⚠ ↑+34.68% 52% 19%
eyeriss_like_quad_core 101581557 → 201589494 ⚠ ↑+98.45% 52% 9.5%
eyeriss_like_single_core 303038604 → 303038597 ↓-0.00% 52% 25%
fusemax 233766951 → 26214403 ⚠ ↓-88.79% 6.2% 0.75%
meta_prototype 101711933 → 201588793 ⚠ ↑+98.20% 100% 3.1%
simba 125344507.60 → 201589015 ⚠ ↑+60.83% 65% 0.17%
simba_small 106627594 → 201589340 ⚠ ↑+89.06% 65% 1.6%
tpu_like_quad_core 101580877 → 201588814 ⚠ ↑+98.45% 97% 1.6%
hardware_two_conv — 8 hardware (⚠ 7 flagged)

batch=1, in_ch=8, H=32, W=32, out_ch1=16, out_ch2=32, kernel=3x3, bf16 (generic auto-tiling, not layer-fused)

Hardware total_latency (base → cur) Δ% array fill MAC eff (e2e) note
eyeriss_like_dual_core 74631 → 80083 ⚠ ↑+7.31% 51% 22%
eyeriss_like_quad_core 39917 → 45657 ⚠ ↑+14.38% 51% 19%
eyeriss_like_single_core 115003 → 114999 ↓-0.00% 68% 31%
fusemax 187612 → 18598 ⚠ ↓-90.09% 1.6% 0.48% low array utilization
meta_prototype 15341 → 20761 ⚠ ↑+35.33% 51% 14%
simba 4610.67 → 12052 ⚠ ↑+161.39% 50% 1.3%
simba_small 8706 → 14881 ⚠ ↑+70.93% 50% 9.7%
tpu_like_quad_core 12300 → 18040 ⚠ ↑+46.67% 42% 8.0%

To regenerate baseline: python scripts/analysis/render_metrics_comment.py --update-baseline

@asyms asyms changed the title Hold a group's cycles to its busiest core, link and memory port Report the latency the solve models, held to every core, link and memory port Oct 6, 2026
@asyms
asyms force-pushed the fix/throughput-bound-transfers branch from 360a3af to ce5d4b2 Compare October 6, 2026 23:27

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant