Repository navigation
Judge an AIE matmul against its MAC unit's peak - #161
Conversation
Stream AIE Metrics Regression Guard⚠ 7 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[fusemax], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core] 16 of 16 cells captured Provenance: baseline hardware_swiglu — 8 hardware (⚠ 1 flagged)
hardware_two_conv — 8 hardware (⚠ 6 flagged)
To regenerate baseline: |
b2682b9 to
db2da5a
Compare
d5e7d8b to
ab2a786
Compare
Stacked on #160.
The AIE estimator took a matmul's ideal cycles from a fixed rate of 32 bf16 ops per cycle per tile. A measured call of the library's matmul runs at about 150 MACs per cycle, so the reported efficiency came out above 1.
A matmul kernel's ideal now comes from the kernel library's MAC tile, one of which the matmul unit computes per cycle. For the AIE2P library that is 8x8x8: the bfp16
mmul<8,8,8>the bf16 matmul issues, 512 MACs per cycle, the same rate as int8 on that tile. On AIE2 the native bf16 tile is 4x8x4, 128 MACs per cycle. Vector kernels, and nodes without a library, keep the per-datatype rate. A measured call is still what prices the node.Test:
tests/unit/test_aie_cost_estimator.py. Fast suite: 911 passed.