Repository navigation
flm/gemm: program B's memtile pool from the runtime sequence - #219
Merged
Merged
Conversation
Merged
5 tasks done
Contributor
CI performance trends
Krackan - OperatorsOperators dropped: Phoenix - OperatorsOperators dropped: |
hunhoffe
force-pushed
the
flm-gemm-programmed-b
branch
from
October 1, 2026 23:25
cddd4e8 to
665ddb3
Compare
…nels mlir-aie 18ca6c1 contains #3818, which adds the flm_gemma4 kernel factories that the ported FastFlowLM operators import. llvm-aie moves to the version that utils/peano-requirements.txt pins at that commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mlir-aie #3807 makes a Worker call the setup kernel of each kernel's contract once before its loop. Most eltwise, activation, norm and rope factories name set_rounding_conv_even as their setup kernel. The design therefore links set_rounding_conv_even_<digest>.ll. IRON did not build that file, and aiecc fails with "cannot read merge-mode link artifact". KernelObjectArtifact.from_extern adds the setup kernel as a dependency of the kernel's artifact. The build compiles it into the kernel directory, where aiecc finds it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
B no longer goes through an ObjectFifo. Its memtile BDs would be device configuration, so a resident fifo would put K and M into the one xclbin every shape shares. Instead each column's memtile holds a fixed pool of k-block slots with per-slot locks, fed and drained over explicit Flows, and the per-shape runtime sequence programs both memtile channels with dma_configure_task chains walked by repeat_count. Where a column-block fits, B is resident and DDR reads it once instead of once per row-block. M is cut into slabs so each arming fits a memtile channel's task queue, the shim outer dimension, and the lock range for residency. Co-Authored-By: Claude <noreply@anthropic.com>
Needs the mlir-aie dynamic-memtile-bd features (target-model limits, tile_dma_chain, Lock.set, repeat splitting, BD reclaim, slice-task decomposition, >4-D patterns, route-endpoint channels), which are not in a wheel yet. - Every hardware limit comes from the target model. B's memtile chains are tile_dma_chains restarted with Task.start(repeat_count=), the lock re-arm is Lock.set, and B's fills go through its Flow. No BD id is pinned. - The compiler owns BD liveness and queue depth. Fills and drains are unmanaged and only each column's last C of a slab carries a token. retire/pending/OVERLAP and the slab-end frees are gone. A memtile chain is pushed at the first block it covers, so slabs are cut only for B residency past the lock range. - The compiler owns splitting: one tap per leg per block. a_split_for, c_split_for, emit_split, the shim outer-dimension cut and op.py's _hw_stride_ok fallback are gone, and m_chunk's A is one 5-D tap. - B's Flows name no channel, and no tile is pinned; the compiler assigns B's channels around A and C, and the placer places every tile. On Strix Halo (pmode default): 180/180 tests pass. Over the 30 benchmark shapes, 8 interleaved rounds against the previous commit, errors are bit-identical and the median latency change is +0.32%, against +0.21% for a no-change control. E2B q and gateup at M=256 are 9-11% faster, from the compiler's queue depth. Co-Authored-By: Claude <noreply@anthropic.com>
#3791 changed during review. Memtile B chains are now tasks on the flow's endpoint (endpoint.task(...) then start()) instead of tile_dma_chain, Runtime.add_buffer is gone (a task declares its buffers), and BD reclaim is opt-in, so the instruction stream is built with aiecc --reclaim-runtime-bds. On Strix (npu2), mlir_aie 1.4.4.dev69: flm/gemm 180/180 pass. Co-Authored-By: Claude <noreply@anthropic.com>
- Test M=16384, K=512 on both devices: 64 row-block units, one past what a lock counts, so B is armed in two resident slabs. No shape reached that cut before. - Fix references the previous commits left stale: benchmark.py named the deleted test_gemm_split_leg_bounds, the E4B comment in test.py still described the design splitting its own legs, and the README called the NPU1 path unrun and dev69 the pin. - The README no longer claims every limit comes from the target model: B_MAX_SLOTS does not. - Import AIETileType and DMAChannelDir from aie.dialects.aie, make _Slab a dataclass, rename the host-side n_work() to work_blocks() so it no longer shares a name with the core's n_work, and spell out the set_up()/push_b_mt() branch and the runs=/repeat_count= offset. - python design.py ran into a NameError on MIN_M; drop unused imports. - Re-measure the README's numbers against #232 on dev73: median -23.9% (was -17.4% against the dev69 fifo version), errors bit-identical. The emitted MLIR is byte-identical to the previous commit for seven shapes across npu1 and npu2. On Strix Halo, flm/gemm passes 185/185. Co-Authored-By: Claude <noreply@anthropic.com>
hunhoffe
force-pushed
the
flm-gemm-programmed-b
branch
from
October 2, 2026 19:35
665ddb3 to
edcb2b5
Compare
Every flm.GEMM shape times out on NPU1 hardware with the programmed-B path. A run of those timeouts leaves the Phoenix runner returning EIO until it is reset, which fails every job after it. Skip both modules on npu1 until the hang is fixed; the npu1 parameters stay in place, so re-enabling is dropping the marker. Co-Authored-By: Claude <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unsupported NPU1 execution remains unguarded, and separate-dispatch sequence builds omit required descriptor reclamation.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Moves FLM GEMM’s B-operand memory programming into per-shape runtime sequences, enabling B reuse while preserving shared xclbins.
Changes:
- Adds resident/streamed B pools and slab scheduling, with compiler-managed DMA resources.
- Updates toolchain dependencies and builds kernel setup dependencies.
- Adds slab coverage, skips unsupported NPU1 tests, and updates documentation.
| File | Description |
|---|---|
| requirements.txt | Updates MLIR-AIE and LLVM-AIE versions. |
| iron/operators/flm/testing.py | Adds the NPU1 GEMM skip marker. |
| iron/operators/flm/gemm/test.py | Adds multi-slab coverage and updates split-path tests. |
| iron/operators/flm/gemm/README.md | Documents runtime B pools and compiler-managed transfers. |
| iron/operators/flm/gemm/op.py | Enables descriptor reclamation and simplifies chunk selection. |
| iron/operators/flm/gemm/design.py | Implements B pools, explicit flows, and per-shape slab scheduling. |
| iron/operators/flm/gemm/benchmark.py | Applies the NPU1 skip and updates explanatory comments. |
| iron/common/compilation/base.py | Builds contract setup kernels as dependencies. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| # on a push issued after it (see design.py's emit_slab). | ||
| extra_flags=["--reclaim-runtime-bds"], | ||
| ) | ||
| self.add_artifacts([self.xclbin_artifact, self.insts_artifact]) |
Comment on lines
+23
to
+25
| skip_flm_gemm_on_npu1 = pytest.mark.skipif( | ||
| _dev is not None and _dev.resolve().name == "npu1", | ||
| reason="flm.GEMM hangs on NPU1 hardware", |
hunhoffe
marked this pull request as ready for review
October 3, 2026 03:09
andrej
enabled auto-merge
October 3, 2026 19:28
# Conflicts: # iron/operators/flm/gemm/design.py
hunhoffe
added a commit
that referenced
this pull request
Oct 5, 2026
Brings in the FastFlowLM operator family and its Gemma 4 example (#226 prefill attention, #228 decode layer, #229 LM head, #230 flm common, #231 gemma4_flm, #233 the flm GEMM factory), #225's dynamic runtime sequence, #219's B memtile pool, #232's setup kernels and #236's TTFT and decode-rate benchmarks. Devel validated these designs on the NPU; this ports those decisions into iron-next's declared-Operator form, and iron-next's layout wins wherever the two differ: - PrefillAttention, PrefillSlidingAttention, DecodeLayer and LMHead are declared Operators with In/Out operands whose exported_design wraps devel's design functions; their dispatch parameters are DispatchTime values. They register in flm/__init__ beside GEMM and DequantBFP. - OperatorImage takes dispatch scalars as keyword arguments and carries no insts when a design's stream is generated per call (#225). - DecodeLayer's four layer types share the global type's xclbin; each type keeps its own stream through configuration(). - The flm designs take the dev88 taplib: fills and drains walk TensorAccessPattern views, and DecodeLayer's sequence runs on named Flow/PacketFlow unmanaged fills. - Prefill and layer validate L_end and context lengths before dispatch; a misaligned L_end would otherwise hang the NPU. - dispatch_params.py becomes a declared ChunkCopy with a DispatchTime n, checking that one image serves n = 3, 1, 16, 3. Its cpp/xclbin-rebuild and OperatorSequence/set_parameters cases go with those APIs, which iron-next does not have. - tests/compilation/factory_kernel_object.py goes with the IRON compilation rules it tested; mlir-aie's KernelObjectCache owns that path now. - gemma4_flm/build.py builds through OperatorImage and the compile cache (NPU_CACHE_HOME defaults to build/cache in the Makefile) and takes argparse flags; test.py reports TTFT and decode rate via record_property. - requirements.txt keeps iron-next's mlir_aie 1.4.4.dev88 pin. Co-Authored-By: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Program
flm/gemm's B operand from the per-shape runtime sequence, so B's memtile layout no longer has to be baked into the single xclbin that every shape shares.Previously B went through an ObjectFifo, and a resident fifo's memtile BDs are device configuration: K and M would end up in the xclbin. Now each column's memtile holds a fixed pool of k-block slots, and the runtime sequence programs it for the shape at hand. Where a column-block fits, B stays resident, so DDR reads it once instead of once per row-block.
Stacked on #232 (mlir_aie 1.4.4.dev73, which carries #3791's runtime memtile tasks,
Lock.setand--reclaim-runtime-bds).Added
endpoint.task(...), restarted withstart(repeat_count=)), and aLock.setre-arm per slab.Changed
B_MAX_SLOTS, since which BDs are static is the compiler's choice.aiecc --reclaim-runtime-bds.m_chunkA is a single 5-D tap.Removed
retire,pending,OVERLAPand the slab-end frees.a_split_for,c_split_for,emit_split, the shim outer-dimension cut, and op.py's_hw_stride_okfallback.Results
On Strix Halo (npu2), pmode default:
pytest iron/operators/flm/gemm/passes 185/185, including the new slab test.black --check, ruff andreuse lintare clean (no C++ changes).NPU1 (Phoenix) status: the npu1 path builds and places, but on hardware every flm/gemm shape currently times out. This PR's Phoenix CI run hit 46 consecutive timeouts, after which the runner returned EIO until it was reset. Until that is fixed, NPU1 is not a supported target for this operator.
PR Merge Checklist
develcommit and pointing todevel.