Skip to content

flm/gemm: program B's memtile pool from the runtime sequence - #219

Merged
andrej merged 9 commits into
develfrom
flm-gemm-programmed-b
Oct 3, 2026
Merged

andrej merged 9 commits into
develfrom
flm-gemm-programmed-b

Conversation

@hunhoffe

@hunhoffe hunhoffe commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Program flm/gemm's B operand from the per-shape runtime sequence, so B's memtile layout no longer has to be baked into the single xclbin that every shape shares.

Previously B went through an ObjectFifo, and a resident fifo's memtile BDs are device configuration: K and M would end up in the xclbin. Now each column's memtile holds a fixed pool of k-block slots, and the runtime sequence programs it for the shape at hand. Where a column-block fits, B stays resident, so DDR reads it once instead of once per row-block.

Stacked on #232 (mlir_aie 1.4.4.dev73, which carries #3791's runtime memtile tasks, Lock.set and --reclaim-runtime-bds).

Added

  • A per-column memtile pool of B slots, each with its own producer/consumer locks, fed from the shim and broadcast to the cores over explicit Flows.
  • Per-shape memtile tasks on each Flow's endpoint (endpoint.task(...), restarted with start(repeat_count=)), and a Lock.set re-arm per slab.
  • B residency: when a column-block's k-blocks fit the pool, B is read from DDR once per column-block.
  • A test at M=16384, K=512 on each device: 64 row-block units, past the 63 a lock can count, so B is armed in two resident slabs.

Changed

  • Hardware limits (task queue depth, repeat count, lock range, memtile and core memory sizes, compute rows) come from the target model. The exception is B_MAX_SLOTS, since which BDs are static is the compiler's choice.
  • The compiler owns BD lifetime and queue depth. Fills and drains are unmanaged, and the instruction stream is built with aiecc --reclaim-runtime-bds.
  • The compiler owns splitting: the design issues one tap per leg per block, and with m_chunk A is a single 5-D tap.
  • B's Flows name no channel and no tile is pinned. The compiler assigns channels and the placer places every tile.
  • README rewritten around the new B path.

Removed

  • B's ObjectFifo.
  • The hand DMA bookkeeping: retire, pending, OVERLAP and the slab-end frees.
  • The hand splitting: a_split_for, c_split_for, emit_split, the shim outer-dimension cut, and op.py's _hw_stride_ok fallback.

Results

On Strix Halo (npu2), pmode default:

  • Tests: pytest iron/operators/flm/gemm/ passes 185/185, including the new slab test.
  • 30-shape benchmark sweep (4 interleaved rounds against Bump mlir-aie to 1.4.4.dev73 and build the setup kernels of factory kernels #232): errors bit-identical on every shape; median latency -23.9% (range -37.7% to +1.8%), and -23.8% measured as speedup over FastFlowLM's frozen prebuilt overlay, the toolchain-neutral control.
    • Resident shapes at M ≥ 1024: -20.9% to -37.7%.
    • M = 256: +0.9% to -13.5%.
    • Streamed down projections at M ≥ 1024: -2.0% to +1.8%, i.e. noise.
  • Lint: black --check, ruff and reuse lint are clean (no C++ changes).

NPU1 (Phoenix) status: the npu1 path builds and places, but on hardware every flm/gemm shape currently times out. This PR's Phoenix CI run hit 46 consecutive timeouts, after which the runner returned EIO until it was reset. Until that is fixed, NPU1 is not a supported target for this operator.

PR Merge Checklist

  1. The PR is rebased on the latest devel commit and pointing to devel.
  2. Your PR has been reviewed and approved.
  3. All checks are passing.

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

CI performance trends

b5cf638 (2026-10-03T19:34:28Z)

Krackan - Operators

Operators dropped: swiglu_prefill_stream

Phoenix - Operators

Operators dropped: flm/gemm

⚠️ No results from: Krackan - Applications (missing). This comment covers the remaining suites only.

@hunhoffe
hunhoffe force-pushed the flm-gemm-programmed-b branch from cddd4e8 to 665ddb3 Compare October 1, 2026 23:25
andrej and others added 6 commits October 2, 2026 09:16
…nels

mlir-aie 18ca6c1 contains #3818, which adds the flm_gemma4 kernel
factories that the ported FastFlowLM operators import. llvm-aie moves to
the version that utils/peano-requirements.txt pins at that commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mlir-aie #3807 makes a Worker call the setup kernel of each kernel's
contract once before its loop. Most eltwise, activation, norm and rope
factories name set_rounding_conv_even as their setup kernel. The design
therefore links set_rounding_conv_even_<digest>.ll. IRON did not build
that file, and aiecc fails with "cannot read merge-mode link artifact".

KernelObjectArtifact.from_extern adds the setup kernel as a dependency of
the kernel's artifact. The build compiles it into the kernel directory,
where aiecc finds it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
B no longer goes through an ObjectFifo. Its memtile BDs would be device
configuration, so a resident fifo would put K and M into the one xclbin
every shape shares. Instead each column's memtile holds a fixed pool of
k-block slots with per-slot locks, fed and drained over explicit Flows,
and the per-shape runtime sequence programs both memtile channels with
dma_configure_task chains walked by repeat_count. Where a column-block
fits, B is resident and DDR reads it once instead of once per row-block.

M is cut into slabs so each arming fits a memtile channel's task queue,
the shim outer dimension, and the lock range for residency.

Co-Authored-By: Claude <noreply@anthropic.com>
Needs the mlir-aie dynamic-memtile-bd features (target-model limits,
tile_dma_chain, Lock.set, repeat splitting, BD reclaim, slice-task
decomposition, >4-D patterns, route-endpoint channels), which are not in
a wheel yet.

- Every hardware limit comes from the target model. B's memtile chains
  are tile_dma_chains restarted with Task.start(repeat_count=), the lock
  re-arm is Lock.set, and B's fills go through its Flow. No BD id is
  pinned.
- The compiler owns BD liveness and queue depth. Fills and drains are
  unmanaged and only each column's last C of a slab carries a token.
  retire/pending/OVERLAP and the slab-end frees are gone. A memtile
  chain is pushed at the first block it covers, so slabs are cut only
  for B residency past the lock range.
- The compiler owns splitting: one tap per leg per block. a_split_for,
  c_split_for, emit_split, the shim outer-dimension cut and op.py's
  _hw_stride_ok fallback are gone, and m_chunk's A is one 5-D tap.
- B's Flows name no channel, and no tile is pinned; the compiler assigns
  B's channels around A and C, and the placer places every tile.

On Strix Halo (pmode default): 180/180 tests pass. Over the 30 benchmark
shapes, 8 interleaved rounds against the previous commit, errors are
bit-identical and the median latency change is +0.32%, against +0.21%
for a no-change control. E2B q and gateup at M=256 are 9-11% faster,
from the compiler's queue depth.

Co-Authored-By: Claude <noreply@anthropic.com>
#3791 changed during review. Memtile B chains are now tasks on the flow's
endpoint (endpoint.task(...) then start()) instead of tile_dma_chain,
Runtime.add_buffer is gone (a task declares its buffers), and BD reclaim
is opt-in, so the instruction stream is built with
aiecc --reclaim-runtime-bds.

On Strix (npu2), mlir_aie 1.4.4.dev69: flm/gemm 180/180 pass.

Co-Authored-By: Claude <noreply@anthropic.com>
- Test M=16384, K=512 on both devices: 64 row-block units, one past what
  a lock counts, so B is armed in two resident slabs. No shape reached
  that cut before.
- Fix references the previous commits left stale: benchmark.py named the
  deleted test_gemm_split_leg_bounds, the E4B comment in test.py still
  described the design splitting its own legs, and the README called the
  NPU1 path unrun and dev69 the pin.
- The README no longer claims every limit comes from the target model:
  B_MAX_SLOTS does not.
- Import AIETileType and DMAChannelDir from aie.dialects.aie, make _Slab
  a dataclass, rename the host-side n_work() to work_blocks() so it no
  longer shares a name with the core's n_work, and spell out the
  set_up()/push_b_mt() branch and the runs=/repeat_count= offset.
- python design.py ran into a NameError on MIN_M; drop unused imports.
- Re-measure the README's numbers against #232 on dev73: median -23.9%
  (was -17.4% against the dev69 fifo version), errors bit-identical.

The emitted MLIR is byte-identical to the previous commit for seven
shapes across npu1 and npu2. On Strix Halo, flm/gemm passes 185/185.

Co-Authored-By: Claude <noreply@anthropic.com>
@hunhoffe
hunhoffe force-pushed the flm-gemm-programmed-b branch from 665ddb3 to edcb2b5 Compare October 2, 2026 19:35
Every flm.GEMM shape times out on NPU1 hardware with the programmed-B
path. A run of those timeouts leaves the Phoenix runner returning EIO
until it is reset, which fails every job after it. Skip both modules on
npu1 until the hang is fixed; the npu1 parameters stay in place, so
re-enabling is dropping the marker.

Co-Authored-By: Claude <noreply@anthropic.com>
@hunhoffe hunhoffe changed the title Flm gemm programmed b flm/gemm: program B's memtile pool from the runtime sequence Oct 2, 2026
@hunhoffe
hunhoffe requested a balanced review from Copilot October 2, 2026 19:37

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unsupported NPU1 execution remains unguarded, and separate-dispatch sequence builds omit required descriptor reclamation.

Review effort: Balanced
Findings: 2 High severity

Open (2)
What changed in this PR

Moves FLM GEMM’s B-operand memory programming into per-shape runtime sequences, enabling B reuse while preserving shared xclbins.

Changes:

  • Adds resident/streamed B pools and slab scheduling, with compiler-managed DMA resources.
  • Updates toolchain dependencies and builds kernel setup dependencies.
  • Adds slab coverage, skips unsupported NPU1 tests, and updates documentation.
File Description
requirements.txt Updates MLIR-AIE and LLVM-AIE versions.
iron/​operators/​flm/​testing.py Adds the NPU1 GEMM skip marker.
iron/​operators/​flm/​gemm/​test.py Adds multi-slab coverage and updates split-path tests.
iron/​operators/​flm/​gemm/​README.md Documents runtime B pools and compiler-managed transfers.
iron/​operators/​flm/​gemm/​op.py Enables descriptor reclamation and simplifies chunk selection.
iron/​operators/​flm/​gemm/​design.py Implements B pools, explicit flows, and per-shape slab scheduling.
iron/​operators/​flm/​gemm/​benchmark.py Applies the NPU1 skip and updates explanatory comments.
iron/​common/​compilation/​base.py Builds contract setup kernels as dependencies.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

# on a push issued after it (see design.py's emit_slab).
extra_flags=["--reclaim-runtime-bds"],
)
self.add_artifacts([self.xclbin_artifact, self.insts_artifact])
Comment on lines +23 to +25
skip_flm_gemm_on_npu1 = pytest.mark.skipif(
_dev is not None and _dev.resolve().name == "npu1",
reason="flm.GEMM hangs on NPU1 hardware",
@hunhoffe
hunhoffe marked this pull request as ready for review October 3, 2026 03:09
@andrej
andrej enabled auto-merge October 3, 2026 19:28
# Conflicts:
#	iron/operators/flm/gemm/design.py
@andrej
andrej added this pull request to the merge queue Oct 3, 2026
Merged via the queue into devel with commit b73e3f8 Oct 3, 2026
6 checks passed
@andrej
andrej deleted the flm-gemm-programmed-b branch October 3, 2026 23:11
hunhoffe added a commit that referenced this pull request Oct 5, 2026
Brings in the FastFlowLM operator family and its Gemma 4 example (#226
prefill attention, #228 decode layer, #229 LM head, #230 flm common, #231
gemma4_flm, #233 the flm GEMM factory), #225's dynamic runtime sequence,
#219's B memtile pool, #232's setup kernels and #236's TTFT and decode-rate
benchmarks. Devel validated these designs on the NPU; this ports those
decisions into iron-next's declared-Operator form, and iron-next's layout
wins wherever the two differ:

- PrefillAttention, PrefillSlidingAttention, DecodeLayer and LMHead are
  declared Operators with In/Out operands whose exported_design wraps
  devel's design functions; their dispatch parameters are DispatchTime
  values. They register in flm/__init__ beside GEMM and DequantBFP.
- OperatorImage takes dispatch scalars as keyword arguments and carries no
  insts when a design's stream is generated per call (#225).
- DecodeLayer's four layer types share the global type's xclbin; each type
  keeps its own stream through configuration().
- The flm designs take the dev88 taplib: fills and drains walk
  TensorAccessPattern views, and DecodeLayer's sequence runs on named
  Flow/PacketFlow unmanaged fills.
- Prefill and layer validate L_end and context lengths before dispatch; a
  misaligned L_end would otherwise hang the NPU.
- dispatch_params.py becomes a declared ChunkCopy with a DispatchTime n,
  checking that one image serves n = 3, 1, 16, 3. Its cpp/xclbin-rebuild
  and OperatorSequence/set_parameters cases go with those APIs, which
  iron-next does not have.
- tests/compilation/factory_kernel_object.py goes with the IRON compilation
  rules it tested; mlir-aie's KernelObjectCache owns that path now.
- gemma4_flm/build.py builds through OperatorImage and the compile cache
  (NPU_CACHE_HOME defaults to build/cache in the Makefile) and takes
  argparse flags; test.py reports TTFT and decode rate via record_property.
- requirements.txt keeps iron-next's mlir_aie 1.4.4.dev88 pin.

Co-Authored-By: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants