Repository navigation
Conversation
Stream AIE Metrics Regression Guard⚠ 14 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[eyeriss_like_dual_core], hardware_swiglu[eyeriss_like_quad_core], hardware_swiglu[fusemax], hardware_swiglu[meta_prototype], hardware_swiglu[simba], hardware_swiglu[simba_small], hardware_swiglu[tpu_like_quad_core], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core] 16 of 16 cells captured Provenance: baseline hardware_swiglu — 8 hardware (⚠ 7 flagged)
hardware_two_conv — 8 hardware (⚠ 7 flagged)
To regenerate baseline: |
…d the charges in their own levels
360a3af to
ce5d4b2
Compare
Stacked on #162.
The cycles a group reported were not the latency the solve modelled:
memory_ports, so a solve could run a DRAM port above its rate.Now:
memory_portsis a default family, with its interval and burst bounds on, so every iteration is held to its busiest memory port inside the MILP. It can still be set to report only throughdefault_families(options=...).offchip_timedwhen its off-chip contention is on.total_latency. The throughput bound computed after the solve is removed.The matmul above now solves to 12.72 M cycles. That is its objective, its reported cycles and its DRAM port busy time, with the port at 100 % of the interval.
Tests:
tests/unit/test_reported_latency.py, plus the objective-level, family-order and memory-port tests updated to the new levels and defaults. Fast suite: 914 passed. The slow suite fails in the same 10 tests as its base.