Skip to content

Report the Gemma 4 example's time to first token and decode rate as benchmarks - #236

Merged
andrej merged 1 commit into
amd:develfrom
andrej:gemma4-bench
Oct 4, 2026
Merged

andrej merged 1 commit into
amd:develfrom
andrej:gemma4-bench

Conversation

@andrej

@andrej andrej commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

The Gemma 4 example (#231) runs in the examples workflow but does not appear on the benchmark site. The site plots only parametrizations with the bench marker, and only the metrics a test declares. test_iron_matches_engine had neither, so krackan/examples/latest.csv records it as a pass with Bench=no and no measurements.

This PR makes the test report like llama_3.2_1b:

  • The test runs once per prompt (word_problem, and long_prompt, which spans three 512-token chunks). Each prompt carries bench.
  • It declares TTFT and TPS with the same patterns as llama_3.2_1b, so they land in the workflow's TTFT (mean) and TPS (mean) columns. The values come from flm's /api/generate response for the IRON engine.
  • Each prompt's tokens must still equal the stock FastFlowLM engine's.
  • make engine and the stock engine's reference run once per module instead of once per iteration.

Run locally the way CI runs it (pytest -m "not extensive" iron/applications/gemma4_flm --csv-output=..., 5 iterations, Strix): 10 passed in 110 s, after make engine.

Prompt Checks TTFT (mean) TPS (median)
long_prompt (1259 tokens) 5/5 2.44 s 24.4
word_problem 5/5 0.68 s 25.4

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

CI performance trends

439cca5 (2026-10-04T16:20:05Z)

Krackan - Operators

Operators dropped: swiglu_prefill_stream

Phoenix - Operators

Operators dropped: flm/gemm

The benchmark site plots only parametrizations with the bench mark, and
only the metrics a test declares. The Gemma 4 test had neither, so CI
recorded it as a pass with no measurements.

The test now runs once per prompt, and each prompt carries the bench mark.
It declares TTFT and TPS with the patterns of llama_3.2_1b and prints the
IRON engine's prefill duration and decode rate from flm's /api/generate
response. Each prompt's tokens must still equal the stock engine's.

make engine and the stock engine's reference run once per module, not once
per iteration: the stock engine is deterministic, so every iteration
compares with one run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@andrej
andrej added this pull request to the merge queue Oct 4, 2026
Merged via the queue into amd:devel with commit 79c04e6 Oct 4, 2026
5 checks passed
@andrej
andrej deleted the gemma4-bench branch October 4, 2026 17:56
hunhoffe added a commit that referenced this pull request Oct 5, 2026
Brings in the FastFlowLM operator family and its Gemma 4 example (#226
prefill attention, #228 decode layer, #229 LM head, #230 flm common, #231
gemma4_flm, #233 the flm GEMM factory), #225's dynamic runtime sequence,
#219's B memtile pool, #232's setup kernels and #236's TTFT and decode-rate
benchmarks. Devel validated these designs on the NPU; this ports those
decisions into iron-next's declared-Operator form, and iron-next's layout
wins wherever the two differ:

- PrefillAttention, PrefillSlidingAttention, DecodeLayer and LMHead are
  declared Operators with In/Out operands whose exported_design wraps
  devel's design functions; their dispatch parameters are DispatchTime
  values. They register in flm/__init__ beside GEMM and DequantBFP.
- OperatorImage takes dispatch scalars as keyword arguments and carries no
  insts when a design's stream is generated per call (#225).
- DecodeLayer's four layer types share the global type's xclbin; each type
  keeps its own stream through configuration().
- The flm designs take the dev88 taplib: fills and drains walk
  TensorAccessPattern views, and DecodeLayer's sequence runs on named
  Flow/PacketFlow unmanaged fills.
- Prefill and layer validate L_end and context lengths before dispatch; a
  misaligned L_end would otherwise hang the NPU.
- dispatch_params.py becomes a declared ChunkCopy with a DispatchTime n,
  checking that one image serves n = 3, 1, 16, 3. Its cpp/xclbin-rebuild
  and OperatorSequence/set_parameters cases go with those APIs, which
  iron-next does not have.
- tests/compilation/factory_kernel_object.py goes with the IRON compilation
  rules it tested; mlir-aie's KernelObjectCache owns that path now.
- gemma4_flm/build.py builds through OperatorImage and the compile cache
  (NPU_CACHE_HOME defaults to build/cache in the Makefile) and takes
  argparse flags; test.py reports TTFT and decode rate via record_property.
- requirements.txt keeps iron-next's mlir_aie 1.4.4.dev88 pin.

Co-Authored-By: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant