Skip to content

Commit b44c615

Browse files
Oppenclaude[bot]
andauthored
feat: guest-side step-profiling markers and per-step function histograms for the recursion verifier (#726)
* feat: recursion profiling + measurement programs Add the measurement/profiling harness for the in-VM STARK verifier: - `empty`-proof and `deserialize-only` bench guests + `sp1/verifier` cross-prover comparison, all exercising the no_std verifier. - Expand the recursion smoke test with PC-histogram, sampled-flamegraph, page-count, cycle-count and per-step-breakdown diagnostics, plus the `make test-profile-recursion` targets and the histogram-aggregation CI script/workflow. - Expose read-only `Executor::memory()`, `Memory::cells()` and `SymbolTable::functions()` accessors and make `flamegraph::demangle` public so the diagnostics can resolve guest PCs to functions. * refactor(prover): drop per-address PC table from recursion profile The top-100 per-address table carried bare PCs with no file:line, so it was not actionable for optimization and the CI aggregator already discarded it. Keep the per-function fold (the view that matters); terminate the aggregator's function-table parse on the trailing rule instead of the removed PC header. * refactor(prover): share setup/progress across recursion diagnostics Extract setup_guest_run (blob build + ELF load + Executor::new) and a log_progress throttled-readout factory, used by the cycle-count, page-count, PC-histogram, sampled-flamegraph and step-breakdown diagnostics. Generalize the PC-histogram runner over guest name + progress stride so the deserialize-only histogram is a one-line caller instead of a near-duplicate. * cargo fmt * refactor(prover): unify recursion execute-only diagnostics Collapse the cycle-count, PC-histogram and step-breakdown diagnostics into one parameterized run_profile(guest, stride, opts, detailed): total cycles print unconditionally, the top-25 functions + per-step breakdown gate on detailed (they share one streamed pass over the same PC stream). Every variant now comes in 1query and multiquery flavours for both recursion and the deserialize-only control. Route execute_outer_and_commit through drive_executor too — the rebased streaming finish() makes its hand-rolled drain loop redundant. * build: enable the deserialize-only recursion guest Add deserialize-only to RECURSION_GUESTS and migrate the guest to the recursion guest's std shape (lambda_vm_syscalls + build-std std), since the old no_std panic handler collided with std. Add getrandom_backend="custom" to its cargo config (transitive getrandom 0.3 needs it) and track its Cargo.lock. The deser control guest now builds and its profile tests run. * build: point profile-recursion make targets at renamed tests * docs: trim recursion smoke-test doc comments * refactor(prover): drop test_host_verify_step_timings The smoke pipelines already host-verify the inner proof, so building with --features stark/instruments surfaces the per-step timings; the dedicated test was just that verify minus the guest run. Documented the flag in the module doc. * Remove the unused SP1 verifier bench program It was never wired into the bench harness or CI (run.sh uses sp1/fibonacci), and its in-VM verifier-cost comparison is superseded by the recursion profile tests in this PR. * cargo fmt * fix ci bug Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> * fix ci bug Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> * ci: gate recursion-profile comment job on profile not being skipped * lint * inline(never) for high-level steps to avoid missing symbols * Revert inline(never) for after_round_1 For some reason that alone appears to inhibit completely the effects of next PRs pre-built commitments and vkey optimization. * fix: reintroduce addr2line Steps detection was misbehaving due to inlined functions not emitting symbols. The solutions were either marking `#[inline(never)]`, which in the case of `replay_rounds_after_round_1` inhibits optimizations. Since we added the dependency, we took more advantage of it and expanded the detailed profile with a per file:line breakdown as well. * Revert "fix: reintroduce addr2line" This reverts commit 3df1e08. * feat: guest-side step-profiling markers for the recursion verifier Replace symbol/DWARF-based verifier-step detection with an explicit addi x0,x0,N marker instruction, immune to inlining. Adds a STEP_DECODE_DONE marker in the recursion guest itself, making the deserialize-only control guest (manual A/B subtraction) redundant. * test: drop recursion smoke-test flamegraph and page-count diagnostics Too much reviewer overhead for their current value. The sampled flamegraph will come back once the executor's flamegraph tooling makes it simple to reimplement; the page-count histogram isn't interesting right now. * refactor: drop accessors only used by the removed diagnostics SymbolTable::functions() and Memory::cells() existed solely for the flamegraph/page-count smoke tests just deleted. * feat: split airs/bus-balance from decode in recursion step profiling Add STEP_AIRS_AND_BUS_BALANCE_DONE marker so the verifier's preprocessed FFT+Merkle commitment build (VmAirs::new) is bucketed separately from postcard decode and from multi_verify's transcript replay. The top-25 cycle table now also tags each row with its verifier step, so e.g. how much of step4:openings is keccak is visible at a glance. Update the CI histogram aggregator to parse and render the new step column. * cargo fmt * fix: per-step top-25 tables instead of a single tagged table The previous split added a step column to one combined top-25 table, losing per-step rank/cum% fidelity. Print the global top-25 (all steps folded together) plus a separate top-25 table per verifier step, so each step's own hottest functions and their cumulative share are visible directly. Update the CI aggregator to parse and render the new multi-table output. * fix: per-step top-25 percentages relative to step cycles, not total Per-step tables previously used the global cycle count as the pct denominator, so a function dominating a cheap step (e.g. 90% of step2:claimed) rendered as a near-zero percentage of the whole run — useless for spotting what dominates within that step. Use each step's own cycle total as the denominator for its table instead; the global table still uses the run's total. Update the CI aggregator to parse and surface the per-step denominator. * lower requirements for comment * fix: NOP/marker-0 collision and step-bucketing latch in recursion profiling decode_step_marker required only dst==0, matching the canonical NOP (addi x0, x0, 0) as marker 0; pin src==0 and imm!=0 per the documented addi x0, x0, N convention. run_profile latched the step bucket at the highest marker ever seen, so multi_verify's per-AIR-table 3,4,5,6 repetition folded every table after the first into bucket 6. Track the latest marker instead. --------- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
1 parent 71a99f1 commit b44c615

14 files changed

Lines changed: 880 additions & 71 deletions

File tree

Lines changed: 180 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,180 @@
1+
#!/usr/bin/env python3
2+
"""Format the recursion-guest per-function profile as a Markdown PR comment.
3+
4+
`test_recursion_profile_1query`/`_multiquery` print a global top-25 functions
5+
table (folded over all verifier steps, % of total run cycles), followed by
6+
one top-25 table per verifier step (% of that step's own cycles, so the
7+
table shows what dominates *within* the step) — e.g. how much of
8+
`step4:openings` is `keccak`. We parse all of those tables and render them
9+
as Markdown.
10+
11+
Top 25 functions by cycle count (aggregated over their PCs, all steps; % of total cycles):
12+
rank cycles % cum % PCs function
13+
1 5335072 24.95% 24.95% 72 <...>::visit_seq::<...>
14+
15+
Top 25 functions by cycle count — step airs_bus_balance (% of this step's 5129138364 cycles):
16+
rank cycles % cum % PCs function
17+
1 5335072 24.95% 24.95% 72 <...>::visit_seq::<...>
18+
19+
Reads the test's captured output from argv[1]; writes the Markdown body to
20+
argv[2] (or stdout).
21+
"""
22+
23+
import re
24+
import sys
25+
from collections import OrderedDict
26+
27+
# A per-function summary row: rank, cycles, pct%, cum%, pcs, function.
28+
FN_ROW = re.compile(
29+
r"^\s*\d+\s+(\d+)\s+([\d.]+)%\s+([\d.]+)%\s+(\d+)\s+(.*\S)\s*$"
30+
)
31+
HEADER_ROW = re.compile(r"^\s*rank\s+cycles")
32+
GLOBAL_TABLE_START = re.compile(
33+
r"Top \d+ functions by cycle count \(aggregated over their PCs, all steps"
34+
)
35+
STEP_TABLE_START = re.compile(
36+
r"Top \d+ functions by cycle count — step (\S+) \(% of this step's (\d+) cycles\):"
37+
)
38+
TOTAL_CYCLES = re.compile(r"Total cycles\s*:\s*(\d+)")
39+
UNIQUE_PCS = re.compile(r"Unique PCs\s*:\s*(\d+)")
40+
EXEC_TIME = re.compile(r"Exec time\s*:\s*(\S+)")
41+
42+
GLOBAL_KEY = "__global__"
43+
44+
45+
def parse(text):
46+
total_cycles = unique_pcs = exec_time = None
47+
# GLOBAL_KEY -> {"denom": int|None, "rows": [...]}, then one entry per
48+
# step tag in first-seen order.
49+
tables = OrderedDict()
50+
current = None
51+
skip_header = False
52+
for line in text.splitlines():
53+
if total_cycles is None and (m := TOTAL_CYCLES.search(line)):
54+
total_cycles = int(m.group(1))
55+
if unique_pcs is None and (m := UNIQUE_PCS.search(line)):
56+
unique_pcs = int(m.group(1))
57+
if exec_time is None and (m := EXEC_TIME.search(line)):
58+
exec_time = m.group(1)
59+
60+
if GLOBAL_TABLE_START.search(line):
61+
current = GLOBAL_KEY
62+
tables[current] = {"denom": total_cycles, "rows": []}
63+
skip_header = True
64+
continue
65+
if m := STEP_TABLE_START.search(line):
66+
current = m.group(1)
67+
tables[current] = {"denom": int(m.group(2)), "rows": []}
68+
skip_header = True
69+
continue
70+
71+
if current is None:
72+
continue
73+
if skip_header:
74+
# The header row right after a table-start line; anything else
75+
# (e.g. a stray blank line) just ends the table early, which is
76+
# fine — an empty table renders as "no rows".
77+
skip_header = False
78+
if HEADER_ROW.match(line):
79+
continue
80+
if m := FN_ROW.match(line):
81+
tables[current]["rows"].append(
82+
{
83+
"cycles": int(m.group(1)),
84+
"pct": m.group(2),
85+
"cum": m.group(3),
86+
"pcs": int(m.group(4)),
87+
"fn": m.group(5),
88+
}
89+
)
90+
else:
91+
current = None
92+
93+
return total_cycles, unique_pcs, exec_time, tables
94+
95+
96+
def short(name, width=90):
97+
return name if len(name) <= width else name[: width - 1] + "…"
98+
99+
100+
def render_table(rows, denom_label):
101+
if not rows:
102+
return "> _no rows_\n"
103+
body = "| Rank | Cycles | % | Cum % | PCs | Function |\n"
104+
body += "|-----:|-------:|--:|------:|----:|----------|\n"
105+
for i, r in enumerate(rows, 1):
106+
body += (
107+
f"| {i} | {r['cycles']:,} | {r['pct']}% | {r['cum']}% | "
108+
f"{r['pcs']} | `{short(r['fn'])}` |\n"
109+
)
110+
last_cum = rows[-1]["cum"]
111+
body += (
112+
f"\n<sub>Each function's cycles are summed over all its program counters "
113+
f"in this table's scope; the top {len(rows)} cover {last_cum}% of "
114+
f"{denom_label}.</sub>\n"
115+
)
116+
return body
117+
118+
119+
def render(total_cycles, unique_pcs, exec_time, tables, title="Recursion guest profile"):
120+
if not tables.get(GLOBAL_KEY, {}).get("rows"):
121+
return (
122+
f"### {title}\n\n"
123+
"> ⚠️ No per-function rows found in the test output — the run may "
124+
"have failed before printing the table. Check the workflow logs.\n"
125+
)
126+
127+
body = f"### {title}\n\n"
128+
if total_cycles is not None:
129+
body += f"**Total cycles:** {total_cycles:,}"
130+
if unique_pcs is not None:
131+
body += f" · **Unique PCs:** {unique_pcs:,}"
132+
if exec_time:
133+
body += f" · **Exec time:** {exec_time}"
134+
body += "\n\n"
135+
136+
global_rows = tables[GLOBAL_KEY]["rows"]
137+
body += f"#### Top {len(global_rows)} functions by cycles (all steps)\n\n"
138+
body += render_table(global_rows, "total cycles")
139+
140+
for step, table in tables.items():
141+
if step == GLOBAL_KEY:
142+
continue
143+
rows, denom = table["rows"], table["denom"]
144+
denom_note = f" of {denom:,} step cycles" if denom is not None else ""
145+
body += (
146+
f"\n<details><summary>Step <code>{step}</code>{denom_note} — "
147+
f"top {len(rows)} functions</summary>\n\n"
148+
)
149+
body += render_table(rows, "this step's cycles")
150+
body += "\n</details>\n"
151+
152+
return body
153+
154+
155+
def main():
156+
import argparse
157+
158+
ap = argparse.ArgumentParser(description=__doc__)
159+
ap.add_argument("log", help="captured test output to parse")
160+
ap.add_argument("-o", "--out", help="write Markdown here instead of stdout")
161+
ap.add_argument(
162+
"-t",
163+
"--title",
164+
default="Recursion guest profile",
165+
help="section heading (e.g. the test/config name)",
166+
)
167+
args = ap.parse_args()
168+
169+
with open(args.log, "r", errors="replace") as f:
170+
text = f.read()
171+
body = render(*parse(text), title=args.title)
172+
if args.out:
173+
with open(args.out, "w") as f:
174+
f.write(body)
175+
else:
176+
sys.stdout.write(body)
177+
178+
179+
if __name__ == "__main__":
180+
main()
Lines changed: 178 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,178 @@
1+
name: Profile Recursion (PR)
2+
3+
# Runs the recursion-guest PC histogram diagnostics (single-query and
4+
# multi-query, in parallel via a matrix) and posts a combined per-function
5+
# profile as a PR comment. Triggered by a `/profile_recursion` comment from a
6+
# repo member, or manually via workflow_dispatch.
7+
8+
on:
9+
workflow_dispatch:
10+
issue_comment:
11+
types: [created]
12+
13+
permissions:
14+
contents: read
15+
pull-requests: write
16+
17+
concurrency:
18+
group: profile-recursion-${{ github.event.issue.number || github.run_id }}
19+
cancel-in-progress: true
20+
21+
jobs:
22+
# One job per configuration; they run in parallel and each uploads a Markdown
23+
# fragment artifact. The `comment` job stitches them into one PR comment.
24+
profile:
25+
# Skip unless: workflow_dispatch, or "/profile_recursion" comment on a PR by a member.
26+
if: >-
27+
github.event_name == 'workflow_dispatch' ||
28+
(github.event_name == 'issue_comment' &&
29+
github.event.issue.pull_request &&
30+
startsWith(github.event.comment.body, '/profile_recursion') &&
31+
contains(fromJSON('["MEMBER","OWNER","COLLABORATOR"]'), github.event.comment.author_association))
32+
runs-on: [self-hosted, bench]
33+
timeout-minutes: 90
34+
strategy:
35+
fail-fast: false
36+
matrix:
37+
include:
38+
- name: single-query
39+
test: single
40+
title: "Single query (blowup=2, 1 query)"
41+
- name: multi-query
42+
test: multi
43+
title: "Multi query (blowup=8, 128-bit)"
44+
steps:
45+
- name: React to comment
46+
if: github.event_name == 'issue_comment' && matrix.name == 'single-query'
47+
uses: actions/github-script@v7
48+
with:
49+
script: |
50+
await github.rest.reactions.createForIssueComment({
51+
owner: context.repo.owner,
52+
repo: context.repo.repo,
53+
comment_id: context.payload.comment.id,
54+
content: 'eyes'
55+
});
56+
57+
- name: Get PR head ref
58+
id: pr-ref
59+
if: github.event_name == 'issue_comment'
60+
env:
61+
GH_TOKEN: ${{ github.token }}
62+
PR_NUM: ${{ github.event.issue.number }}
63+
run: |
64+
SHA=$(gh pr view "$PR_NUM" --repo "$GITHUB_REPOSITORY" --json headRefOid -q .headRefOid)
65+
echo "sha=$SHA" >> "$GITHUB_OUTPUT"
66+
67+
- name: Checkout
68+
uses: actions/checkout@v4
69+
with:
70+
ref: ${{ steps.pr-ref.outputs.sha || github.sha }}
71+
72+
- name: Setup Rust Environment
73+
uses: ./.github/actions/setup-rust
74+
75+
- name: Add cargo to PATH
76+
run: echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
77+
78+
- name: Run recursion PC histogram (${{ matrix.name }})
79+
env:
80+
TEST: ${{ matrix.test }}
81+
run: |
82+
# Self-provision the RISC-V sysroot in a user-writable dir (the default
83+
# /opt path on the bench runner is root-owned); the guest ELF build the
84+
# test triggers picks this up via the Makefile's `SYSROOT_DIR ?=`.
85+
export SYSROOT_DIR="$HOME/.lambda-vm-sysroot"
86+
set -o pipefail
87+
make test-profile-recursion-$TEST 2>&1 | tee /tmp/hist.log
88+
89+
- name: Aggregate into a per-function fragment
90+
if: always()
91+
env:
92+
TITLE: ${{ matrix.title }}
93+
run: |
94+
python3 .github/scripts/aggregate_recursion_histogram.py \
95+
/tmp/hist.log --title "$TITLE" --out "/tmp/fragment-${{ matrix.name }}.md"
96+
cat "/tmp/fragment-${{ matrix.name }}.md" >> "$GITHUB_STEP_SUMMARY"
97+
98+
- name: Upload fragment
99+
if: always()
100+
uses: actions/upload-artifact@v4
101+
with:
102+
name: profile-fragment-${{ matrix.name }}
103+
path: /tmp/fragment-${{ matrix.name }}.md
104+
retention-days: 7
105+
106+
# Stitch the matrix fragments into a single PR comment.
107+
comment:
108+
needs: profile
109+
# always() so partial-matrix failures still post; skip when `profile` was
110+
# skipped (non-/profile_recursion or non-member comment) so this job — and
111+
# the self-hosted bench runner it spins up — doesn't fire on every comment.
112+
if: always() && github.event_name == 'issue_comment' && needs.profile.result != 'skipped'
113+
runs-on: ubuntu-latest
114+
steps:
115+
- name: Get PR head ref
116+
id: pr-ref
117+
env:
118+
GH_TOKEN: ${{ github.token }}
119+
PR_NUM: ${{ github.event.issue.number }}
120+
run: |
121+
SHA=$(gh pr view "$PR_NUM" --repo "$GITHUB_REPOSITORY" --json headRefOid -q .headRefOid)
122+
echo "sha=$SHA" >> "$GITHUB_OUTPUT"
123+
124+
- name: Download fragments
125+
uses: actions/download-artifact@v4
126+
with:
127+
path: fragments
128+
pattern: profile-fragment-*
129+
merge-multiple: true
130+
131+
- name: Assemble comment body
132+
env:
133+
COMMIT_SHA: ${{ steps.pr-ref.outputs.sha }}
134+
run: |
135+
{
136+
echo "## Recursion guest profile"
137+
echo
138+
# Single-query first, then multi-query, then any others.
139+
for frag in fragments/fragment-single-query.md \
140+
fragments/fragment-multi-query.md; do
141+
[ -f "$frag" ] && { cat "$frag"; echo; }
142+
done
143+
echo "<sub>Commit: ${COMMIT_SHA:0:8} · Runner: self-hosted bench</sub>"
144+
} > /tmp/profile_comment.md
145+
cat /tmp/profile_comment.md
146+
147+
- name: Comment on PR
148+
uses: actions/github-script@v7
149+
with:
150+
script: |
151+
const fs = require('fs');
152+
const body = fs.readFileSync('/tmp/profile_comment.md', 'utf8');
153+
154+
const { data: comments } = await github.rest.issues.listComments({
155+
owner: context.repo.owner,
156+
repo: context.repo.repo,
157+
issue_number: context.issue.number,
158+
});
159+
// Reuse our own marker comment so repeated /profile_recursion runs update in place.
160+
const existing = comments.find(c =>
161+
c.user.type === 'Bot' &&
162+
c.body.includes('Recursion guest profile')
163+
);
164+
if (existing) {
165+
await github.rest.issues.updateComment({
166+
owner: context.repo.owner,
167+
repo: context.repo.repo,
168+
comment_id: existing.id,
169+
body,
170+
});
171+
} else {
172+
await github.rest.issues.createComment({
173+
owner: context.repo.owner,
174+
repo: context.repo.repo,
175+
issue_number: context.issue.number,
176+
body,
177+
});
178+
}

‎Makefile‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
11
.PHONY: deps deps-linux deps-macos compile-programs-asm compile-programs-rust compile-bench \
22
compile-programs compile-recursion-elfs clean-asm clean-rust clean-bench clean-shared \
33
clean-recursion-elfs clean test test-asm \
4-
test-rust test-executor test-flamegraph flamegraph-prover \
4+
test-rust test-executor test-flamegraph flamegraph-prover test-profile-recursion test-profile-recursion-single test-profile-recursion-multi \
55
test-fast test-prover test-prover-all test-disk-spill test-math-cuda test-cuda-integration \
66
bench-math-cuda bench-prover bench-prover-cuda build check clippy fmt lint regen-ethrex-fixtures \
77
update-ethrex-fixture-checksums check-ethrex-fixture-checksums
@@ -232,6 +232,14 @@ test-rust: compile-programs-rust
232232
test-flamegraph:
233233
cargo test -p executor --test flamegraph
234234

235+
test-profile-recursion: test-profile-recursion-single test-profile-recursion-multi
236+
237+
test-profile-recursion-single: compile-recursion-elfs
238+
cargo test --package lambda-vm-prover --lib test_recursion_profile_1query -- --ignored --nocapture
239+
240+
test-profile-recursion-multi: compile-recursion-elfs
241+
cargo test --package lambda-vm-prover --lib test_recursion_profile_multiquery -- --ignored --nocapture
242+
235243
# Regenerate the committed ethrex block fixtures (see tooling/ethrex-fixtures).
236244
# Run after bumping the ethrex rev; README checksums are refreshed automatically.
237245
regen-ethrex-fixtures:

‎bench_vs/lambda/recursion/Cargo.toml‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,8 @@ version = "0.1.0"
66
edition = "2024"
77

88
[dependencies]
9-
lambda-vm-prover = { path = "../../../prover", default-features = false }
9+
lambda-vm-prover = { path = "../../../prover", default-features = false, features = [
10+
"profile-markers",
11+
] }
1012
lambda-vm-syscalls = { path = "../../../syscalls" }
1113
postcard = { version = "1.0", features = ["alloc"] }

0 commit comments

Comments
 (0)