Repository navigation
Report the Gemma 4 example's time to first token and decode rate as benchmarks - #236
Merged
Merged
Conversation
Contributor
CI performance trends
Krackan - OperatorsOperators dropped: Phoenix - OperatorsOperators dropped: |
The benchmark site plots only parametrizations with the bench mark, and only the metrics a test declares. The Gemma 4 test had neither, so CI recorded it as a pass with no measurements. The test now runs once per prompt, and each prompt carries the bench mark. It declares TTFT and TPS with the patterns of llama_3.2_1b and prints the IRON engine's prefill duration and decode rate from flm's /api/generate response. Each prompt's tokens must still equal the stock engine's. make engine and the stock engine's reference run once per module, not once per iteration: the stock engine is deterministic, so every iteration compares with one run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hunhoffe
added a commit
that referenced
this pull request
Oct 5, 2026
Brings in the FastFlowLM operator family and its Gemma 4 example (#226 prefill attention, #228 decode layer, #229 LM head, #230 flm common, #231 gemma4_flm, #233 the flm GEMM factory), #225's dynamic runtime sequence, #219's B memtile pool, #232's setup kernels and #236's TTFT and decode-rate benchmarks. Devel validated these designs on the NPU; this ports those decisions into iron-next's declared-Operator form, and iron-next's layout wins wherever the two differ: - PrefillAttention, PrefillSlidingAttention, DecodeLayer and LMHead are declared Operators with In/Out operands whose exported_design wraps devel's design functions; their dispatch parameters are DispatchTime values. They register in flm/__init__ beside GEMM and DequantBFP. - OperatorImage takes dispatch scalars as keyword arguments and carries no insts when a design's stream is generated per call (#225). - DecodeLayer's four layer types share the global type's xclbin; each type keeps its own stream through configuration(). - The flm designs take the dev88 taplib: fills and drains walk TensorAccessPattern views, and DecodeLayer's sequence runs on named Flow/PacketFlow unmanaged fills. - Prefill and layer validate L_end and context lengths before dispatch; a misaligned L_end would otherwise hang the NPU. - dispatch_params.py becomes a declared ChunkCopy with a DispatchTime n, checking that one image serves n = 3, 1, 16, 3. Its cpp/xclbin-rebuild and OperatorSequence/set_parameters cases go with those APIs, which iron-next does not have. - tests/compilation/factory_kernel_object.py goes with the IRON compilation rules it tested; mlir-aie's KernelObjectCache owns that path now. - gemma4_flm/build.py builds through OperatorImage and the compile cache (NPU_CACHE_HOME defaults to build/cache in the Makefile) and takes argparse flags; test.py reports TTFT and decode rate via record_property. - requirements.txt keeps iron-next's mlir_aie 1.4.4.dev88 pin. Co-Authored-By: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Gemma 4 example (#231) runs in the examples workflow but does not appear on the benchmark site. The site plots only parametrizations with the
benchmarker, and only the metrics a test declares.test_iron_matches_enginehad neither, sokrackan/examples/latest.csvrecords it as a pass withBench=noand no measurements.This PR makes the test report like
llama_3.2_1b:word_problem, andlong_prompt, which spans three 512-token chunks). Each prompt carriesbench.TTFTandTPSwith the same patterns asllama_3.2_1b, so they land in the workflow'sTTFT (mean)andTPS (mean)columns. The values come from flm's/api/generateresponse for the IRON engine.make engineand the stock engine's reference run once per module instead of once per iteration.Run locally the way CI runs it (
pytest -m "not extensive" iron/applications/gemma4_flm --csv-output=..., 5 iterations, Strix): 10 passed in 110 s, aftermake engine.long_prompt(1259 tokens)word_problem🤖 Generated with Claude Code