Repository navigation
Conversation
…fer, so self-attention keeps its query and key axes apart
Stream AIE Metrics Regression Guard⚠ 5 cell(s) flagged (total_latency > 0.1% tol): hardware_two_conv[eyeriss_like_dual_core], hardware_two_conv[eyeriss_like_quad_core], hardware_two_conv[meta_prototype], hardware_two_conv[simba_small], hardware_two_conv[tpu_like_quad_core] 16 of 16 cells captured Provenance: baseline hardware_swiglu — 8 hardware
hardware_two_conv — 8 hardware (⚠ 5 flagged)
To regenerate baseline: |
The steady-state lowering moves each tensor to all its readers with one multicast transfer, and a transfer copies along identity maps, so every reader's index dims become the transfer's. When two readers index the tensor along unrelated unique dims, that merges them. In a self-attention the query projection and the key/value projections both read the input rows, so the query and key axes of the scores collapsed into one: splitting the queries over four cores also split the keys, and each core computed only its diagonal block of the scores (four 32x32 tiles of a 128x128 matrix).
_reader_walksnow groups a tensor's readers so that, at every index of the tensor, the readers in a group share a unique dim, and each group gets its own transfer, named after its first reader (Transfer(x for proj_q)). Readers that slide different windows along the same axis, like a 3x3 and a 5x5 conv, still share one copy. A tensor with one group lowers exactly as before.DecisionSpace.copied_bitssums what every transfer reading a tensor delivers, so off-chip traffic counts both copies.Test: in the attention catalog block, the scores' query and key dims stay distinct in the steady state and the input gets one transfer per walk. It fails without the change. The rest of the fast suite is unchanged.