fix(autoimplement): bound the decide prompt evidence - #100
Merged
Merged
Conversation
Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Keep raw step results in run state. Project the value that a prompt shows. The decide prompt embedded raw step results. One verification result of 265,600 bytes was recorded twice inside a documented plan and reached the prompt again through the observation, so the assembled prompt grew past the model context window and the run failed. - Add src/workflows/prompt-evidence.ts. A bounded copy replaces a subtree that carries a registered schema, bounds long strings, arrays, and depth, and collapses the oldest ledger entries when the prompt budget is tight. - Export an evidence view for the change-verification schema. It keeps the route, reason, findings, and fingerprints, and names command logs by size. - Share PROMPT_CEILING_CHARS between the Pi agent group and the decide prompt. - Document the rule in docs/WORKFLOWS.md and cover it in tests. The regression test assembles a 4,006,002-character prompt on the earlier code and stays under the ceiling after the change. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
…orkspace Pi Reviewer found that boundLedger measured the entry sizes without the array brackets and separators that the serialized ledger adds. A ledger a few characters over budget therefore collapsed every entry instead of only the oldest ones, against the documented behavior. Count that overhead and add the regression case, which fails on the earlier code. Keep the prepared workspace in the change-verification evidence view, so the decider still sees the repository, branch, and mode. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the depth cap ate the change-verification view's own fields on the observation path. The observation nests a result two levels deeper than a ledger entry does, so the per-command summary collapsed to one reference there while it survived in the ledger. What a model saw of one result type depended on where the prompt happened to put it. A view result now opens a fresh depth budget, because a view is a bounded replacement for a whole subtree. A view also runs once per schema on a path, which keeps the walk finite. Both cases have tests that fail on the earlier code, and the decide-prompt regression test now asserts the per-command summary survives. Also count the ledger brackets and separators, so a ledger a few characters over budget collapses its oldest entries instead of all of them. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
…failing Pi Reviewer found that the projection bounds leaves, arrays, and depth, but not object width or the size of a whole value, and the Observation line never went through a size bound. A recorded result whose bulk is many modest values, such as a file-to-digest map with thousands of entries, could therefore push the fixed lines over the ceiling, and the decide node threw. Every retry recomputed the same failure, so the run ended blocked with evidence the workflow could have shortened. - Cap the fields of one projected object at EVIDENCE_MAX_FIELDS and name the rest in one reference, which bounds a wide map. - Add boundEvidence, which collapses the largest field of one observation object until it fits. The smallest fields, which are the decisive ones, stay readable, so the available routes still reach the model. - The decide prompt now shortens the ledger, then the largest observation field, and only reports an error when the fixed lines alone are over. The new case gives the decide node an observation of 1.6 million characters whose parts all sit exactly at the projection caps. On the earlier code the decide node threw with a 1,602,830-character prompt. It now stays bounded and keeps the available routes. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
…rence Pi Reviewer found that boundLedger's last-resort branch wrapped every entry output in a reference without the guard its main loop uses. A projected output that was already a reference became a reference of a reference, so its digest named the intermediate reference, its size counted the reference, and its excerpt was lost. The branch is reachable, because the decide prompt collapses the ledger to measure how much room the ledger floor needs. Record the final-revision live E2E in the plan document, and add the case that fails on the earlier code. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that projectObject copied fields with a direct assignment into an ordinary object. A recorded own `__proto__` field therefore reached the prototype setter instead of becoming a field, so the prompt copy dropped it without a reference and took the payload as its prototype. The same assignment overwrote a source field named `omittedFields`. Build the projected object on a null prototype, and give the omitted-fields reference the next free name. Both cases have tests that fail on the earlier code. The durable result in run state is unaffected either way. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the backstop error named the largest fixed line, although the fixed lines are known to fit whenever that branch runs. The overflow is in the observation or the recent-attempt list, so the message pointed at an unrelated line and repeated one number. Report the size of the fixed lines, the observation, and the recent-attempt list. The case that reaches the backstop records the message shape, including the part sizes of 95025, 1083, and 2281 characters. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the ledger dropped the terminal node of a control include but kept its return step, `<mount>/__piw_exit_<exit>`, whose output wraps the same payload. The comment and the plan promise that one result appears once, so the newest control include appeared in the observation and again in the ledger, and the surviving copy was one of the newest entries, which the budget keeps whole while it collapses older ones. Drop the return step instead and keep the terminal node, which holds the result itself. The wrapped copy is also the larger one, so the ledger shrinks. The new case fails on the earlier filter. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the generic item cap was 20 while a command batch holds up to 64 checks. A verification result with more than 20 checks therefore lost the id, outcome, exit code, and log size of every later check from the decide prompt, which contradicts the change-verification view's own contract and the documented behavior of the view. The largest list a registered view carries whole is a command batch, so the generic item cap now follows MAX_COMMAND_BATCH_ITEMS. A registered view is the correctness layer for its own result type, and the size budget, not this cap, is what bounds one prompt. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
…or the ledger correctly Pi Reviewer found two defects in the previous commit. The return-step filter matched a mount name only at the start of a node id, so a nested include, such as documentation/verification/__piw_exit_ready, was still listed next to the node it returns from, which is the duplication the filter exists to remove. Match the reserved __piw_exit_ prefix at any depth. The ledger floor was read from a fully collapsed ledger, but a reference to a tiny output is larger than the output, so that value could exceed the projected ledger. The observation was then bounded against an inflated reserve and collapsed further than the ceiling required. boundLedger now returns the projected ledger when collapsing would make it larger, which makes its result a true lower bound. Both new cases fail on the earlier code. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
…ceiling Pi Reviewer found that the reserved ledger floor was still not a lower bound. The smallest shape is not always one of the two extremes: collapsing one large oldest output shrinks the ledger, while collapsing the tiny outputs around it turns each of them into a larger reference. boundLedger now measures every shape its walk builds and returns the smallest one when none fits, so the value the decide prompt reserves is the room the ledger truly needs at least. The reviewer also found that PROMPT_CEILING_CHARS bounds the authored prompt only, because the engine appends its live-control block and the step contract after a builder returns. Document what the ceiling covers and the bound on the appended block instead of claiming the whole request. The new boundLedger case fails on the earlier code. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that a reference reported chars: 0 for a value JSON cannot represent, while evidenceChars called the same value unbounded. The digest and the size therefore described different things, and a prompt showed an empty size for exactly the value the module promises to name. Give both the digest and the size one source: the value when JSON can represent it, and the type name when it cannot. evidenceChars reads an unrepresentable value as unbounded for the same reason, so the two measurements agree. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
Pi Reviewer found that the ledger still kept the `observe` step, whose recorded output is the same observation the prompt prints on its own line. The graph runs observe before decide, so that step was the newest entry, which the budget keeps whole while it collapses older entries, and the largest evidence blob appeared twice in one prompt. Drop the observe step for the same reason the ledger already drops the latest control attempt. The new case fails on the earlier filter. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
… room Pi Reviewer found that the first failure path reported that the fixed lines must be at most the ceiling, which contradicts the size it prints when the lines are just under it. That path is reached when the lines leave less than the ledger label plus its newline, so say that instead, and keep the largest line named. The case that reaches this path records the message. Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
[skip ci] Co-Authored-By: DeepSeek V4.1 Flash · Auto · Novita <noreply@pi.dev> Generated-By: pi 0.85.1 (https://pi.dev)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
An Autoimplement decision prompt carried raw step results, and one result filled the whole request.
A documentation result held the same 265 KB verification record twice, and the observation brought it
back again, so the assembled prompt grew to roughly 815 KB.
The request no longer fit the model context window, and the run failed after its one recovery attempt.
This change keeps the complete results in run state and bounds only the copy that a prompt shows.
What Changed
Every step result stays exactly as it was recorded, command logs included.
The value that reaches a prompt is now projected: a large payload becomes a size, a digest, and a
short head-and-tail excerpt.
A result type that owns a lot of evidence can also export a view that keeps the fields a routing
decision reads and names bulk payloads by size.
src/workflows/prompt-evidence.ts. It projects a value against a registry of evidence viewskeyed by the existing versioned
schemaidentifier, bounds long strings, long arrays, and depth,and collapses the oldest ledger entries when the prompt budget is tight.
changeVerificationEvidencefrom the change-verification workflow. It keeps the route, thereason, the failure findings, and the fingerprints, and drops the raw stdout and stderr text.
appear twice, and measures the assembled prompt against
PROMPT_CEILING_CHARS.prompt share one number.
docs/WORKFLOWS.mdand a plan document underdocs/plans/.The projection runs when a prompt is built.
state.steps[].outputandstate.steps[].promptkeeptheir complete values, so a recorded run stays readable and resumable.
Testing
I tested the projection on its own, and I tested it through the real Autoimplement workflow code.
npm run check(format, lint, typecheck, build, tests with coverage). The new module reports 100%of lines and 92.75% of branches.
npm run test:e2e(14 tests, non-destructive, real Pi runtime).npx slophammer-ts@latest dry .andnpx slophammer-ts@latest check . --only ts.dependency-boundaries-required.observeanddecidenodes with a result that holds a1,000,000-character command log, recorded twice inside the documented plan. On the earlier code the
assembled prompt was 4,006,002 characters. On this branch it stays under the ceiling, keeps the
reason, the failure summary, and the fingerprint, and contains none of the log text.
deepseek/deepseek-v4-flashthroughopenrouter,result: passed,measured model cost $0.0035.
Not tested: the prompt builders of the other built-in workflows still embed raw command output.
Risks
The decide prompt is bounded, and the durable record is untouched, so existing runs stay valid.
The main risk is that a projected field was useful to the decider.
asserts that the recorded routes are unchanged.
reference instead of an error.
and its size, instead of sending a request the model cannot answer.
Follow-ups
The other Autoimplement prompt sites (review, CI, and comment inspection) still embed command output,
which the command-batch limit allows up to 1,000,000 characters per check.
Those results carry no log file path, so bounding them needs a decision about how an agent reads the
full log afterwards.