Skip to content

fix(autoimplement): bound the decide prompt evidence - #100

Merged
osolmaz merged 16 commits into
mainfrom
fix/autoimplement-prompt-evidence-budget
Sep 17, 2026
Merged

osolmaz merged 16 commits into
mainfrom
fix/autoimplement-prompt-evidence-budget

Conversation

@osolmaz

@osolmaz osolmaz commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Summary

An Autoimplement decision prompt carried raw step results, and one result filled the whole request.
A documentation result held the same 265 KB verification record twice, and the observation brought it
back again, so the assembled prompt grew to roughly 815 KB.
The request no longer fit the model context window, and the run failed after its one recovery attempt.
This change keeps the complete results in run state and bounds only the copy that a prompt shows.

What Changed

Every step result stays exactly as it was recorded, command logs included.
The value that reaches a prompt is now projected: a large payload becomes a size, a digest, and a
short head-and-tail excerpt.
A result type that owns a lot of evidence can also export a view that keeps the fields a routing
decision reads and names bulk payloads by size.

  • Added src/workflows/prompt-evidence.ts. It projects a value against a registry of evidence views
    keyed by the existing versioned schema identifier, bounds long strings, long arrays, and depth,
    and collapses the oldest ledger entries when the prompt budget is tight.
  • Exported changeVerificationEvidence from the change-verification workflow. It keeps the route, the
    reason, the failure findings, and the fingerprints, and drops the raw stdout and stderr text.
  • The decide prompt projects its observation and its recent-attempt list, drops a result that would
    appear twice, and measures the assembled prompt against PROMPT_CEILING_CHARS.
  • Moved the 96,000-character prompt ceiling into the new module, so the Pi agent group and the decide
    prompt share one number.
  • Added a section to docs/WORKFLOWS.md and a plan document under docs/plans/.

The projection runs when a prompt is built. state.steps[].output and state.steps[].prompt keep
their complete values, so a recorded run stays readable and resumable.

Testing

I tested the projection on its own, and I tested it through the real Autoimplement workflow code.

  • npm run check (format, lint, typecheck, build, tests with coverage). The new module reports 100%
    of lines and 92.75% of branches.
  • npm run test:e2e (14 tests, non-destructive, real Pi runtime).
  • npx slophammer-ts@latest dry . and
    npx slophammer-ts@latest check . --only ts.dependency-boundaries-required.
  • The regression test drives the real observe and decide nodes with a result that holds a
    1,000,000-character command log, recorded twice inside the documented plan. On the earlier code the
    assembled prompt was 4,006,002 characters. On this branch it stays under the ceiling, keeps the
    reason, the failure summary, and the fingerprint, and contains none of the log text.
  • One real-model live E2E: deepseek/deepseek-v4-flash through openrouter, result: passed,
    measured model cost $0.0035.

Not tested: the prompt builders of the other built-in workflows still embed raw command output.

Risks

The decide prompt is bounded, and the durable record is untouched, so existing runs stay valid.
The main risk is that a projected field was useful to the decider.

  • The view keeps every field that route selection and the progress fingerprint read, and a test
    asserts that the recorded routes are unchanged.
  • A value the rules cannot represent, including a cycle, a bigint, or a factory function, becomes a
    reference instead of an error.
  • A prompt that is still over the ceiling after projection fails with a message naming the largest line
    and its size, instead of sending a request the model cannot answer.

Follow-ups

The other Autoimplement prompt sites (review, CI, and comment inspection) still embed command output,
which the command-batch limit allows up to 1,000,000 characters per check.
Those results carry no log file path, so bounding them needs a decision about how an agent reads the
full log afterwards.

osolmaz and others added 16 commits September 18, 2026 01:19
Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Keep raw step results in run state. Project the value that a prompt shows.

The decide prompt embedded raw step results. One verification result of
265,600 bytes was recorded twice inside a documented plan and reached the
prompt again through the observation, so the assembled prompt grew past the
model context window and the run failed.

- Add src/workflows/prompt-evidence.ts. A bounded copy replaces a subtree
  that carries a registered schema, bounds long strings, arrays, and depth,
  and collapses the oldest ledger entries when the prompt budget is tight.
- Export an evidence view for the change-verification schema. It keeps the
  route, reason, findings, and fingerprints, and names command logs by size.
- Share PROMPT_CEILING_CHARS between the Pi agent group and the decide prompt.
- Document the rule in docs/WORKFLOWS.md and cover it in tests.

The regression test assembles a 4,006,002-character prompt on the earlier
code and stays under the ceiling after the change.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
…orkspace

Pi Reviewer found that boundLedger measured the entry sizes without the array
brackets and separators that the serialized ledger adds. A ledger a few
characters over budget therefore collapsed every entry instead of only the
oldest ones, against the documented behavior. Count that overhead and add the
regression case, which fails on the earlier code.

Keep the prepared workspace in the change-verification evidence view, so the
decider still sees the repository, branch, and mode.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the depth cap ate the change-verification view's own
fields on the observation path. The observation nests a result two levels
deeper than a ledger entry does, so the per-command summary collapsed to one
reference there while it survived in the ledger. What a model saw of one
result type depended on where the prompt happened to put it.

A view result now opens a fresh depth budget, because a view is a bounded
replacement for a whole subtree. A view also runs once per schema on a path,
which keeps the walk finite. Both cases have tests that fail on the earlier
code, and the decide-prompt regression test now asserts the per-command
summary survives.

Also count the ledger brackets and separators, so a ledger a few characters
over budget collapses its oldest entries instead of all of them.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
…failing

Pi Reviewer found that the projection bounds leaves, arrays, and depth, but not
object width or the size of a whole value, and the Observation line never went
through a size bound. A recorded result whose bulk is many modest values, such
as a file-to-digest map with thousands of entries, could therefore push the
fixed lines over the ceiling, and the decide node threw. Every retry recomputed
the same failure, so the run ended blocked with evidence the workflow could
have shortened.

- Cap the fields of one projected object at EVIDENCE_MAX_FIELDS and name the
  rest in one reference, which bounds a wide map.
- Add boundEvidence, which collapses the largest field of one observation
  object until it fits. The smallest fields, which are the decisive ones, stay
  readable, so the available routes still reach the model.
- The decide prompt now shortens the ledger, then the largest observation
  field, and only reports an error when the fixed lines alone are over.

The new case gives the decide node an observation of 1.6 million characters
whose parts all sit exactly at the projection caps. On the earlier code the
decide node threw with a 1,602,830-character prompt. It now stays bounded and
keeps the available routes.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
…rence

Pi Reviewer found that boundLedger's last-resort branch wrapped every entry
output in a reference without the guard its main loop uses. A projected output
that was already a reference became a reference of a reference, so its digest
named the intermediate reference, its size counted the reference, and its
excerpt was lost. The branch is reachable, because the decide prompt collapses
the ledger to measure how much room the ledger floor needs.

Record the final-revision live E2E in the plan document, and add the case that
fails on the earlier code.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that projectObject copied fields with a direct assignment
into an ordinary object. A recorded own `__proto__` field therefore reached the
prototype setter instead of becoming a field, so the prompt copy dropped it
without a reference and took the payload as its prototype. The same assignment
overwrote a source field named `omittedFields`.

Build the projected object on a null prototype, and give the omitted-fields
reference the next free name. Both cases have tests that fail on the earlier
code. The durable result in run state is unaffected either way.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the backstop error named the largest fixed line, although
the fixed lines are known to fit whenever that branch runs. The overflow is in
the observation or the recent-attempt list, so the message pointed at an
unrelated line and repeated one number.

Report the size of the fixed lines, the observation, and the recent-attempt
list. The case that reaches the backstop records the message shape, including
the part sizes of 95025, 1083, and 2281 characters.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the ledger dropped the terminal node of a control include
but kept its return step, `<mount>/__piw_exit_<exit>`, whose output wraps the
same payload. The comment and the plan promise that one result appears once, so
the newest control include appeared in the observation and again in the ledger,
and the surviving copy was one of the newest entries, which the budget keeps
whole while it collapses older ones.

Drop the return step instead and keep the terminal node, which holds the result
itself. The wrapped copy is also the larger one, so the ledger shrinks. The new
case fails on the earlier filter.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the generic item cap was 20 while a command batch holds up
to 64 checks. A verification result with more than 20 checks therefore lost the
id, outcome, exit code, and log size of every later check from the decide prompt,
which contradicts the change-verification view's own contract and the documented
behavior of the view.

The largest list a registered view carries whole is a command batch, so the
generic item cap now follows MAX_COMMAND_BATCH_ITEMS. A registered view is the
correctness layer for its own result type, and the size budget, not this cap, is
what bounds one prompt.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
…or the ledger correctly

Pi Reviewer found two defects in the previous commit.

The return-step filter matched a mount name only at the start of a node id, so a
nested include, such as documentation/verification/__piw_exit_ready, was still
listed next to the node it returns from, which is the duplication the filter
exists to remove. Match the reserved __piw_exit_ prefix at any depth.

The ledger floor was read from a fully collapsed ledger, but a reference to a
tiny output is larger than the output, so that value could exceed the projected
ledger. The observation was then bounded against an inflated reserve and
collapsed further than the ceiling required. boundLedger now returns the
projected ledger when collapsing would make it larger, which makes its result a
true lower bound.

Both new cases fail on the earlier code.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
…ceiling

Pi Reviewer found that the reserved ledger floor was still not a lower bound. The
smallest shape is not always one of the two extremes: collapsing one large oldest
output shrinks the ledger, while collapsing the tiny outputs around it turns each
of them into a larger reference. boundLedger now measures every shape its walk
builds and returns the smallest one when none fits, so the value the decide
prompt reserves is the room the ledger truly needs at least.

The reviewer also found that PROMPT_CEILING_CHARS bounds the authored prompt
only, because the engine appends its live-control block and the step contract
after a builder returns. Document what the ceiling covers and the bound on the
appended block instead of claiming the whole request.

The new boundLedger case fails on the earlier code.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that a reference reported chars: 0 for a value JSON cannot
represent, while evidenceChars called the same value unbounded. The digest and
the size therefore described different things, and a prompt showed an empty size
for exactly the value the module promises to name.

Give both the digest and the size one source: the value when JSON can represent
it, and the type name when it cannot. evidenceChars reads an unrepresentable
value as unbounded for the same reason, so the two measurements agree.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the ledger still kept the `observe` step, whose recorded
output is the same observation the prompt prints on its own line. The graph runs
observe before decide, so that step was the newest entry, which the budget keeps
whole while it collapses older entries, and the largest evidence blob appeared
twice in one prompt.

Drop the observe step for the same reason the ledger already drops the latest
control attempt. The new case fails on the earlier filter.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
… room

Pi Reviewer found that the first failure path reported that the fixed lines must
be at most the ceiling, which contradicts the size it prints when the lines are
just under it. That path is reached when the lines leave less than the ledger
label plus its newline, so say that instead, and keep the largest line named.

The case that reaches this path records the message.

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
[skip ci]

Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev>
Generated-By: pi 0.85.1 (https://pi.dev)
@osolmaz
osolmaz merged commit c10561f into main Sep 17, 2026
@osolmaz
osolmaz deleted the fix/autoimplement-prompt-evidence-budget branch September 17, 2026 19:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant