You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Runtime-sequence DMA on mem/core tiles, with compiler-owned BD bookkeeping - #3791
Runtime sequences could only program DMAs on shim tiles. This PR lets them program mem and core tiles too, and has the compiler handle the bookkeeping designs used to do by hand: picking DMA channels, splitting oversized transfers and repeat counts, and (optionally) recycling BD ids.
What's new
Runtime DMA on mem and core tiles. Runtime-valued buffer descriptors now work on every tile type, so a design with dispatch-time shapes can keep operands resident in a mem tile. Runtime values that would overflow a BD field are caught on the host instead of being silently truncated.
Compiler-assigned DMA channels. A DMA program or runtime task can name an aie.route_endpoint instead of a channel number, and --aie-objectfifo-allocate picks the channel. In IRON, leave the channels off a Flow and use flow.endpoint(tile).
Transfers and repeats of any size.
A transfer pattern too big for one BD, including constant patterns with more than 4 dimensions, is split into pieces (at most 1024).
A repeat count past the hardware maximum becomes several pushes of the same task.
Pieces of concurrent transfers are interleaved so a fill and the drain that depends on it can't deadlock.
Opt-in BD reclaim. With aiecc --reclaim-runtime-bds, running out of BD ids no longer fails compilation: the compiler polls for an earlier task to finish and reuses its BDs. It's off by default because the poll hangs if that task is waiting on something issued after it, which the compiler can't see (see Design rule below).
IRON additions.
endpoint.task(...) reprograms a mem or core tile's DMA from the runtime sequence, and Task.start(repeat_count=) restarts it.
Flow.fill/drain and PacketFlow.fill/drain handle host transfers.
Lock.set(value) sets a lock's value from the sequence.
Buffers and locks used by a TileDma no longer need add_buffer/add_lock.
Freeing a task twice, or starting it after it's freed, is an error, even across range_ loop iterations.
Placement. The sequential placer keeps cores that share a hub tile in one column, instead of placing them in creation order. The unpinned flm GEMM now compiles to the same instructions as its hand-pinned version.
Bugs fixed along the way:
BD-length verification per tile type.
Mem tile BD-id channel parity.
Stride checks with runtime sizes.
Runtime BD addresses on mem and core tiles.
Behavior changes
Breaking (Python):aie.dma_start's channel_index argument is now channel, and accepts an index or a route endpoint.
objectFIFOs can no longer be assigned a channel that a Flow or runtime task already uses explicitly.
New warnings:
a channel's BDs are overcommitted;
a repeat count needs more pushes than the queue holds.
A channel that's never resolved is reported as an error instead of crashing.
Known gap
PacketFlow.endpoint(tile) works, but a PacketFlow's channels still have to be given explicitly. Only plain Flows get compiler-assigned channels for now.
Design rule the compiler can't check
The compiler only sees one tile's DMA queue at a time. If a task can't finish until something issued later (on another core or tile) happens, a design must not need to reclaim that task's BDs. DMATasks.md covers this.
Testing
New hardware tests, run on Strix (npu2). Each one fails its negative control:
Test
Covers
Negative control
runtime_bd_reclaim
160 unfreed tasks through 16 BDs
Fails without the reclaim polls
decompose_split_tasks
36 interleaved fill/drain pieces over 16 BDs
Hangs if fills all go first
decompose_nd_task
A 6-D transfer with repeat
Swapped piece offsets → 4096 mismatches
flow_endpoints
Unpinned tiles, 5 endpoints on one mem tile, broadcast
—
Lit, non-hardware: no new failures. The one failure, checkpoint_resume_aiesim, also fails on main on this machine.
Lit, hardware (npu-xrt + unit_tests): 118 passed, no failures. 107 are npu1-only and skipped.
Downstream: the amd/IRON flm GEMM passes 180/180 on hardware. Across 30 shapes its results are bit-identical to the hand-managed design, and latency is within noise.
A runtime-valued buffer descriptor -- SSA len, sizes, strides, offset, or a
bd_id drawn from the runtime pool -- was rejected on anything but a shim NOC
tile, because the dynamic BD-word encoder only knew the shim register layout.
That is why a design with dispatch-time shapes cannot keep an operand resident
in a mem tile: the mem tile DMA can be programmed from a runtime sequence
today (test/npu-xrt/memtile_dmas/dma_configure_task_*), but only with values
fixed at compile time.
Split the encoder along the axis that actually varies. encodeBdCommon keeps
everything tile-independent -- the hardware size/stride encoding, the
linear-mode decision, buffer_length, the queue repeat count and the runtime
guards -- and three short packers write the register layouts, each meant to be
read line-for-line against its static counterpart in AIEDmaToNpu.cpp's
WriteBdToBlockWritePattern. The new dynamic-matches-static-words.mlir lowers
one descriptor twice per tile type, constant and runtime, and compares the
resulting words; that is what holds the two implementations together.
Field widths now come from the target model's existing
getDmaBdWrapBits/StepBits/IterBits/MaxLen accessors, deleting
ShimBdFieldWidths and the TODO asking for exactly that. The narrower fields
off shim make two guards load-bearing that were not there before: strides had
no runtime guard at all, and buffer_length needed none while it was 32 bits
wide. Both are added, so an oversized runtime value returns nullopt from the
generated builder instead of being silently masked; an oversized compile-time
length is now a diagnostic.
Also fixed along the way, all previously unreachable behind the shim-only
gate: setAddressForSingleBD ignored its runtime register address on the mem
and core tile branches, and folded a runtime offset to zero. A runtime bd_id
combined with offset_state_table_idx is now rejected rather than patching
whichever BD the compile-time literal named.
Co-Authored-By: Claude <noreply@anthropic.com>
Both date from mem tile dma_configure_task becoming usable
(test/npu-xrt/memtile_dmas/dma_configure_task_*) while the surrounding
machinery still assumed runtime sequences only ever drive shim tiles.
NpuWriteBdOp::verify bounds every BD field against the target model except
buffer_length, which needed no bound while it was the shim's full 32-bit word.
On a mem tile it is 17 bits and on a core tile 14, so an oversized constant
length was masked by the packer and silently transferred the wrong amount --
len = 200000 became 68928.
BD-id allocation passed channelIndex=0 unconditionally, with a comment
asserting runtime sequences configure BDs on shim and compute tiles only,
"which are channel-agnostic". A mem tile is not: it partitions its 48
descriptors by channel parity, so an unpinned BD on an odd channel was handed
an id from the low half that the channel cannot submit. Pass the configure's
own channel, and reject a pinned id in the wrong half rather than emitting a
descriptor that never runs.
Co-Authored-By: Claude <noreply@anthropic.com>
The dynamic counterpart of dma_configure_task_token: the same shim <-> mem
tile round trip, but the mem tile's buffer descriptors carry a runtime
transfer length and a runtime d1 wrap, so one xclbin serves every tile count.
The assertion that matters is the tail: the host checks the output past `len`
is untouched. Without it the test passes just as well when the runtime length
is ignored and a max-sized transfer runs every time.
--get-npu-cpp is required rather than convenient -- npu.blockwrite_values has
no encoding on the static binary target, so a runtime-valued BD can only reach
hardware through the generated C++ TXN builder.
Verified on NPU1 (Phoenix) at n = 1, 3 and 8 from a single xclbin.
Co-Authored-By: Claude <noreply@anthropic.com>
IRON could only drive shim DMAs at runtime: RuntimeEndpoint rejects non-shim
tiles and fill/drain always reach for shim_dma_single_bd_task, so the only way
to program a mem tile was TileDma -- which is structural, configured once when
the device loads, and cannot vary per dispatch. An operand held resident in a
mem tile therefore had to give up dynamic shapes.
tile_dma_task is the runtime-sequence peer of TileDma. It names a tile and
channel directly, moves a buffer on that tile, and takes dispatch-time sizes,
strides and length, so the descriptor is rebuilt on every call while the
resident buffer stays put. It reuses the existing vocabulary rather than
growing a parallel one: Acquire/Release for lock handoff, and the same Task
handle fill/drain return, so await_/free and range_ iter_args work unchanged.
A buffer reached only from the sequence body has neither a Worker's fn_args
nor a TileDma to carry it into the Program, and the body is emitted last, so
Runtime.add_buffer registers it the way add_lock and add_tile_dma already do.
The i64-sizes/i32-length split is the one hardware detail that leaked: a
dispatch-time scalar cannot feed both. shim_dma_single_bd_task already
narrowed inline for its repeat_count, so that expression becomes _as_bd_i32
next to _as_i32 in aie.py and both call sites adopt it. Generated MLIR is
byte-identical across every test in test/python.
Co-Authored-By: Claude <noreply@anthropic.com>
These four gated on ryzen_ai_npu2 and hardcoded aie.device(npu2), so they only
ever ran on Strix. Nothing about them is npu2-specific -- the gate came from
the board the author happened to have (#3624). Their untested sibling
matmul_whole_array_dynamic already carries no device gate at all.
Switch to the NPUDEVICE substitution the memtile_dmas tests use, and run every
tile count on both boards. FileCheck goes away with it: on a single-board
machine the absent board's runner expands to echo, which no CHECK line can
match, so the host's exit status becomes the assertion -- which is what
dynamic_conditional_passthrough already did.
All four verified on NPU1 (Phoenix): the designs place in one column and the
dynamic BD free-list pool path runs there unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
verifyConstBdRealizability skipped its positivity check whenever the SIZE was
runtime, even when the STRIDE was a compile-time zero. The encoder still
scales that stride to `stride * elemWidth / granularity - 1`, so zero becomes
-1 and the packer masks it into an all-ones step field -- a step large enough
to walk the BD out of its buffer.
Found by mutating test/npu-xrt/memtile_dmas/dma_configure_task_runtime_len to
zero its mem tile d1 stride: the design compiled clean, passed at n = 1 and 3,
and hung the channel at n = 8 (ERT status 8). A size the compiler cannot see
is the reason to check the stride, not a reason to skip it; only a statically
1 size makes the dimension safe, because that is the one case where hardware
never applies the stride.
Pre-existing, but the mem and core tile paths make it far easier to hit: their
step fields are 17 and 13 bits against shim's 20, and unlike shim they cannot
fold a contiguous scan into buffer_length to sidestep the stride entirely.
Co-Authored-By: Claude <noreply@anthropic.com>
All four already carried the full npu1 machinery -- the NPUDEVICE
substitution and a %run_on_npu1% line for every run -- and then gated
themselves to ryzen_ai_npu2, so the npu1 half never executed.
Verified on NPU1 (Phoenix): each substitutes its npu1 device, builds, and
runs on hardware.
Co-Authored-By: Claude <noreply@anthropic.com>
Nothing about it is npu2-specific: the AIE2 target model reports 63
out-of-order ids, so npu1 has the feature. The gate came from the board the
author had.
Verified on NPU1 (Phoenix). FileCheck goes away with the conversion because on
a single-board machine the absent board's runner expands to echo; the host's
exit status is the assertion.
Co-Authored-By: Claude <noreply@anthropic.com>
…ests
Two of the seven #2463 regressions already had an _npu1 sibling that reuses
the same IRON test.py under %run_on_npu1%; the other five only ever ran on
npu2. The designs pick their device through iron.get_current_device(), so
following the existing pattern is all that is needed.
Not verified locally: pyxrt is built for cpython-3.10/3.11 while lit's
%python is 3.12, so xrt_python_bindings is unavailable here and every
python-hardware test -- including the two pre-existing _npu1 variants -- is
Unsupported on this machine. CI is the check.
Co-Authored-By: Claude <noreply@anthropic.com>
A mem tile partitions its 48 buffer descriptors by channel parity -- an even
channel can only submit ids 0-23, an odd channel only 24-47 -- but the runtime
free-list was one flat pool per tile handing out id 0 first. An odd mem tile
channel drawing from it got an id it can never submit, so the transfer would
be configured, pushed, and silently never run.
This PR is what made that reachable: before it, a pooled mem tile task hit the
shim-only rejection in the dynamic BD path, so no mem tile ever drew from the
pool. The static allocator learned the same rule two commits ago; this is its
runtime counterpart.
dma_bd_pool_pop/push now carry the channel, and pools are keyed by the
accessible id RANGE rather than by channel, so channels sharing a partition
still share one pool and cannot be handed the same id twice. The range comes
from isBdChannelAccessible rather than a hardcoded split, so a tile type that
partitions differently needs no change here. A shim tile reports its whole
table for every channel and keeps exactly one pool -- splitting by parity
there would have halved the ids available to each channel.
Also fixes the target model reaching the converter: it runs over ops already
cloned out of the aie.device, where AIE::getTargetModel() falls back to a
default model, which would have sized every pool as if it were a shim tile.
That one is carried on DeviceResolved with the rest of the device state.
Co-Authored-By: Claude <noreply@anthropic.com>
Critical: Runtime local addresses can be truncated without range checks. Critical: Foldable runtime BD IDs can patch BD 0 instead of the configured ID. Moderate: Runtime local offsets lack alignment guards. Moderate: Runtime lengths lack granularity and minimum-size guards.
@copilot please address review feedback and CI errors
Addressed in 5098a9e: corrected runtime address/length bounds and alignment guards, negative field checks, constant SSA BD-id patching, pool FileCheck expectations, and dialect examples. Added lowering and generated-C++ regressions.
Standalone BD pool and transaction-encoding tests passed, along with formatting/whitespace and secret checks. MLIR build/lit tests remain blocked locally by missing MLIR development files/tools. GitHub showed no failed build/test jobs attached to the original PR head, so there were no CI failure logs to diagnose.
@copilot resolve the merge conflicts in this pull request
Resolved and pushed in merge commit a78179c. Both runtime buffer registration and main’s combined TileDma resolution are preserved. GitHub no longer reports merge conflicts.
Python syntax checks and both standalone host tests passed; focused review found no issues. Python regression/lint checks remain blocked by missing dependencies and tooling.
Fixed in 2ec6228: stale rolled-loop pool expectations, missing device substitutions in all five core-DMA fixtures, tile-DMA helper type annotations, and the missing API documentation entry.
Strict docs, targeted Pyright/Ruff/Black, pool/rolled-loop tests, and runtime guard execution passed. All five fixtures parse with the matching CI compiler. Full core compilation still needs CI’s Peano toolchain; CI rerun results are pending.
@copilot it looks like I'm seeing timed out tests. Can you help me understand why?
I looked at the failing run on 5b24796 (run 36652879651). I don't think this PR caused either failure. The test that times out was already close to the limit on main, and this runner was slow that night.
The timeout: python/npu/test_kernels_e2e.py, in both aie2p-8col jobs (XRT and pure HRX)
This is one lit test that runs 146 pytest kernel cases, and lit gives it 600 s (--timeout 600). In our run it was killed at 600 s with 65% of the cases done.
On main, the same test on the same runner (aie6, merge-queue run 36674736679, about 3 hours later) takes 445 s. Main's whole HRX lit step takes about 560 s. So a runner that is about 1.35x slower is enough to hit the limit.
In our run, everything on aie6 was slower, including tests this PR doesn't touch. Our time vs main's:
Test
Ours
Main
test_conv_compile
356 s
252 s
test_compile_link
176 s
95 s
test_prefix_symbols_in_object
142 s
80 s
aiecc/bank_placement_mismatch
138 s
41 s
aiecc/large_stack_chess (xchesscc)
87 s
43 s
The chess test only runs an external compiler, so the slowdown came from the machine, not from the MLIR passes. Overall lit time was 1242 s, against 857 s on main.
This PR doesn't change the kernel library or test_kernels_e2e. The only changes under python/utils/compile are docstrings.
The failure: npu-xrt/dma_s2mm_ooo, one variant only
The variant was --recv-tile core --channels 1 -n 4 --nonuniform. It failed with DRM_IOCTL_AMDXDNA_CREATE_HWCTX IOCTL failed (err=-22) while opening the hardware context, before the design ran.
The 10 variants before it passed. One of them (--recv-tile core -n 4 --packets 8) uses the same tile layout. --nonuniform only changes how many packets each sender sends.
Other PRs hit the same driver errors on the same night: allocator-exhaustive-fallback got CREATE_HWCTX -22 in reconfigure_loadpdi, and other PRs got CREATE_HWCTX or GET_INFO -110. That points to a transient driver problem.
run_on_npu.py only retries on GET_INFO ... (err=-22). A CREATE_HWCTX -22 failure isn't retried.
What to do: re-run the failed jobs. The Windows and wheel jobs weren't timeouts; they were still queued.
I haven't changed any code for this. Two separate fixes would make these failures less likely: splitting test_kernels_e2e (or giving it a per-test timeout above 600 s), and adding CREATE_HWCTX -22 to the retry list in run_on_npu.py. Both are CI changes outside this PR's scope. I can make either one if you'd like.
6c3b85e fixes the one open review finding, the adjacent-tile lock in program.py.
TileDma.all_tiles() now also returns the tiles of the Locks and pinned Buffers its BDs use. Program.resolve therefore resolves a neighbor tile's lock tile, as in memtiledma_channel4lock.mlir, before it calls lk.resolve(). This covers locks reached only through a TileDma BD.
black and ruff pass. I couldn't run the change end to end because the Python bindings aren't built here.
yield_ with a generator:yield_ now turns values into a list once and uses that list both to track tasks and for the scf.yield operands. Before, a generator was used up by the task tracking, so scf.yield got no operands.
Runtime tile BD length: when no length is given, a runtime-valued TileDma BD now multiplies only sizes[-3:]. The outermost of four sizes is the iteration count, as in shim_dma_bd and the C++ encodeBdCommon. Before, it was included, so sizes [2, 1, 1, 16] gave a length of 32 instead of 16.
black and ruff pass on both files. I couldn't run either change because the Python bindings aren't built here. The latest review summary counted 3 findings but showed only these two, so I couldn't find a third to fix.
Track only yields terminating the tracked loop body
python/iron/controlflow.py:95
yield_ also runs for nested scf.if regions inside a range_ body. Those yields currently update the enclosing loop's task metadata, so if the loop subsequently yields a raw SSA handle, its result can incorrectly inherit the nested branch's task type, endpoint, or freed state. Record a yielded task only when this is the scf.yield terminating the tracked scf.for body (and handle raw loop yields explicitly).
@copilot can you thoughtfully address review feedback?
The latest review had one new finding, which 2801bab fixes. I also rechecked the one it still lists as open.
New finding (nested scf.if yields changed the loop's Task state): this was a real bug. Before, any yield_ inside a range_ body updated the loop result's Task state. That included the yield_ that ends a nested scf.if branch. So a loop result could take the Task type, endpoint or freed state from inside an if branch. Now range_ stores its loop body block. yield_ records its values only when the current insertion block is that body block, which means it ends the loop itself. If the body's own yield_ passes a raw SSA handle where a Task was passed in, the result keeps the copy of the initial Task. I updated the range_ docstring to say so.
Still listed as open (adjacent-tile lock tiles): this is already fixed. TileDma.all_tiles() in tile_dma.py returns the tiles of the Buffers and Locks its BDs use, so Program.resolve resolves those tiles before lk.resolve(). The finding stays listed because its review thread was never marked resolved. Please resolve it if you agree.
The change hasn't been run against real MLIR, because the Python bindings aren't built here. I tested it with stubbed aie modules instead:
A Task yielded from an if branch, then a raw handle yielded from the loop body: the result keeps the initial Task's endpoint and live state.
A freed Task on another endpoint yielded from the loop body: the result takes that Task's endpoint and freed state.
black, ruff and pyright pass on controlflow.py, and CodeQL found no alerts.
Task.start() on a carried task:aie-dma-tasks-to-npu now follows a carried task back through scf control flow when lowering a start, as it already did for awaits.
If every value the task can carry comes from one configure, the push uses that configure's head BD, channel, repeat count and token. This is the case when the loop yields the task unchanged.
If it can come from more than one configure, for example a task reconfigured each iteration, there is no single BD id to push. The start now fails with a clear error saying to start the task where it is configured.
Before, both cases failed with only "failed to legalize". The pass also now stops once the task conversion fails, so the configures no longer produce a cascade of extra errors.
The Task.start() docstring states the restriction.
aie-split-long-repeats: a start's own repeat_count needs no configure, so it is split even when the task is carried. Without an override, the pass reads the count through control flow, using the same one-configure rule.
low_level_api.md: both mem-tile tasks were freed right after their push. They no longer are, and the text now says to free a task only once it is known to have finished. test_tile_dma_task_dispatch.py had the same pattern, with S2MM 0 and MM2S 0 sharing the even BD half, so I removed the .free() on its input task too.
New lit tests: dma-tasks-to-npu/start-loop-iter-arg.mlir covers the one-configure and two-configure cases, and repeat-split-carried.mlir covers splitting. I built aie-opt against the MLIR wheel and ran the .mlir lit tests under bd-chains-and-dma-tasks, dialect/AIE, dialect/AIEX, Targets and npu-xrt. There were no new failures; the only failures were tests needing aiecc or aie-translate, which I didn't build, and XFAIL tests. I also ran a runtime-bound loop through split, dynamic BD pool, assign-ids and lowering: the carried start becomes two pushes using the pool-popped BD id. I couldn't run the Python changes because the bindings aren't built here.
…IRON lifetimes
Compiler:
- aie-assign-runtime-sequence-bd-ids reclaim is now opt-in (reclaim-bds,
aiecc --reclaim-runtime-bds); without it, behavior matches main.
- aie-decompose-large-dma-bd caps a transfer at 1024 pieces.
- aie-split-long-repeats warns when a start needs more pushes than the
queue holds.
- Passes and targets that need a concrete DMA channel report a still-
symbolic one as an error instead of asserting.
- aie-objectfifo-allocate warns when a channel's BD demand exceeds what
its BD-parity half has left.
- aiex.npu.assert_bd_field gains min; a runtime transfer length must be
at least one address-gen granule.
- BD field widths are derived from the target model's register layout.
- DMAConfigureTaskLike op interface replaces the decompose pass's
per-op helpers; DmaQueueModel::keyOf keys channels everywhere.
IRON:
- Buffer.place, shared Task lifetime across copies and range_ results,
PacketFlow.endpoint, DmaChannel/Bd emit helpers as methods.
- aie.dma_start raises TypeError on a channel that is neither an index
nor a route endpoint.
Tests for the interface, loop-carried task lifetimes, PacketFlow
endpoints, the channel TypeError and unresolved channels; td
descriptions shortened in favor of DMATasks.md.
Co-Authored-By: Claude <noreply@anthropic.com>
… pattern
Factoring took the first divisor whose recursion succeeded at all, and
slicing now always succeeds, so a long strided row (Llama's vocabulary
split four ways in test_sample_e2e) sliced into ~1000 BDs per fill where
two factorings give one. Search factorings alone for a legal pattern first.
Co-Authored-By: Claude <noreply@anthropic.com>
@copilot resolve the merge conflicts in this pull request
4167781 merges main (c916318) and resolves the conflict.
The only conflict was in AIEDMAChannelAnalysis.h. This branch's streamedChannels map stays as it was. usedStreams now uses main's StreamClaim type (owner plus a packet flag), which replaces this branch's plain set. The rest of the merge applied without conflicts, and the merged AIEDMAChannelAnalysis.cpp uses both maps consistently.
aie-opt builds. I ran the .mlir lit tests that only need aie-opt under objectFifo-allocate, objectFifo-stateful-transform, Passes, lower-cascade-flows, bd-chains-and-dma-tasks and dialect, including main's new stream-port and cascade-order tests. All passed. I skipped the tests that need aiecc or aie-translate, since I didn't build those tools.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Runtime sequences could only program DMAs on shim tiles. This PR lets them program mem and core tiles too, and has the compiler handle the bookkeeping designs used to do by hand: picking DMA channels, splitting oversized transfers and repeat counts, and (optionally) recycling BD ids.
What's new
Runtime DMA on mem and core tiles. Runtime-valued buffer descriptors now work on every tile type, so a design with dispatch-time shapes can keep operands resident in a mem tile. Runtime values that would overflow a BD field are caught on the host instead of being silently truncated.
Compiler-assigned DMA channels. A DMA program or runtime task can name an
aie.route_endpointinstead of a channel number, and--aie-objectfifo-allocatepicks the channel. In IRON, leave the channels off aFlowand useflow.endpoint(tile).Transfers and repeats of any size.
Opt-in BD reclaim. With
aiecc --reclaim-runtime-bds, running out of BD ids no longer fails compilation: the compiler polls for an earlier task to finish and reuses its BDs. It's off by default because the poll hangs if that task is waiting on something issued after it, which the compiler can't see (see Design rule below).IRON additions.
endpoint.task(...)reprograms a mem or core tile's DMA from the runtime sequence, andTask.start(repeat_count=)restarts it.Flow.fill/drainandPacketFlow.fill/drainhandle host transfers.Lock.set(value)sets a lock's value from the sequence.TileDmano longer needadd_buffer/add_lock.range_loop iterations.Placement. The sequential placer keeps cores that share a hub tile in one column, instead of placing them in creation order. The unpinned flm GEMM now compiles to the same instructions as its hand-pinned version.
Bugs fixed along the way:
Behavior changes
aie.dma_start'schannel_indexargument is nowchannel, and accepts an index or a route endpoint.Known gap
PacketFlow.endpoint(tile)works, but a PacketFlow's channels still have to be given explicitly. Only plainFlows get compiler-assigned channels for now.Design rule the compiler can't check
The compiler only sees one tile's DMA queue at a time. If a task can't finish until something issued later (on another core or tile) happens, a design must not need to reclaim that task's BDs.
DMATasks.mdcovers this.Testing
New hardware tests, run on Strix (npu2). Each one fails its negative control:
runtime_bd_reclaimdecompose_split_tasksdecompose_nd_taskflow_endpointscheckpoint_resume_aiesim, also fails on main on this machine.npu-xrt+unit_tests): 118 passed, no failures. 107 are npu1-only and skipped.Companion IRON PR: amd/IRON#219
Checklist
clang-formatfor C++,blackfor Python — see CONTRIBUTING)ruff checkandpyrightpass locally for any touched Python file in their covered paths (see CONTRIBUTING)clang-tidypasses locally for any touched file on the enabled list (see CONTRIBUTING)🤖 Generated with Claude Code