Repository navigation
Conversation
Stream AIE Metrics Regression Guard⚠ 7 cell(s) flagged (total_latency > 0.1% tol): hardware_swiglu[fusemax], hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[fusemax], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core] 16 of 16 cells captured Provenance: baseline hardware_swiglu — 8 hardware (⚠ 1 flagged)
hardware_two_conv — 8 hardware (⚠ 6 flagged)
To regenerate baseline: |
28ba70b to
a1d8a1b
Compare
…or that priced it
Stacked on #159.
FuseMax's 256x1 vector core declared no
operator_types, so core selection treated it as a second generic core and split every conv evenly between it and the 256x256 array. Each conv then ran at the pace of the vector core.fusemax_vec.yamllists the elementwise, reduction and pooling ops it serves, the same set as the TPU7x VPU. Convs and matmuls run on the array. The accelerator is renamed fromquad_coretofusemax.ncores of at leastuunits delivern * uMACs per cycle.operator_typesexclude it is rejected when it is read.The FuseMax conv-window scenarios that split convs over the vector core are replaced by a conv on the array feeding a max pool on the vector core through the shared memory. The shared-memory test is re-derived for this placement: apart, each core fits its tiles in 133 KB, while sharing one memory they need 136 KB, so 134 KB is infeasible.
Fast suite: 907 passed. The slow FuseMax, generic and ResNet tests fail in the same two cases as on #159.