Skip to content

The impl-loop API-error abort cannot see plain-output errors, and when it fires it records no verdict and files no receipt #916

Description

@yuanhao

The impl-loop API-error abort is deaf to plain-output errors, and when it does fire it records nothing — two defects, creator lane

Backlog, creator lane (scripts/evolve.sh is protected). Deliberately unlabelled.

What it is

scripts/evolve.sh aborts the implementation loop when the impl agent hits an API error:

if grep -q '"type":"error"' "$TASK_LOG" 2>/dev/null; then
    echo "    API error in Task $TASK_NUM. Reverting and aborting implementation loop."
    git reset --hard "$PRE_TASK_SHA"; git clean -fd
    TASK_FAILURES=$((TASK_FAILURES + 1)); API_ERROR_ABORT=true; break
fi

Defect 1 — it cannot see the error it was built for

The detector matches a JSON "type":"error" line. The impl agent runs in plain output mode, where the same failure prints:

error: Rate limited, retry after Some(57715000)ms
⏳ stopped retrying on purpose: the provider says its rate limit resets in ~16h …
API error with no fallback configured. Exiting.

(src/agent_builder.rs:1446, on stderr; run_agent_with_fallback captures 2>&1 | tee, so it is in TASK_LOG.) On Day 195 this fired zero times across five sessions whose every call was refused — measured in run 34590906044. The [Agent stopped: grep two lines below has the same shape and works because that string is what yoagent actually prints.

Defect 2 — when it fires, it leaves no record

The branch resets, counts a failure and breaks. It never calls gasp_task_result … rejected, sets no REVERT_REASON, and files no revert receipt. The graph keeps a task.created with no verdict, and the planner learns nothing next session. Compare the ordinary revert path, which records the verdict with a reason, files a receipt with a class the planner has a rule for, and carries the error tail.

Why it was not extended alongside the empty-diff gate (119f172)

Extending the detector alone would have made Day-195-shaped sessions take this path instead of the gate's — i.e. the less informative one. The gate already produces the better record for that shape: (no changes landed), the reason, and the impl log tail in the receipt. What the abort adds is only "do not start task 2 when the provider is down" — worth having, but not at the cost of the record.

The fix, both halves together

  1. Match the plain-output sentinel as well (grep -qF 'API error with no fallback configured. Exiting.'), with the stated limit that it is a cross-boundary string read off src/ output, exactly like [Agent stopped:.
  2. Give the abort branch the revert path's bookkeeping before it breaks: REVERT_CLASS=" (provider refused)", a reason carrying the log tail, gasp_task_result … rejected, and the receipt — then break so remaining tasks are not attempted against a dead provider.

Extract the predicate as a function (agent_hit_api_error LOGFILE) so tests/harness_logic.sh can drive it against fixtures: the JSON line, the plain sentinel wrapped in ANSI codes, and the near-miss — a log that merely contains the word "Rate limited" in prose must not abort.

Activity

  1. yuanhao commented on Sep 13, 2026

    @yuanhao
    CollaboratorAuthor

    Census: it is eight sites, not one — and the first three are the ones that would have made Day 195 RED

    Every API-error detector in scripts/evolve.sh greps the JSON shape "type":"error", and every agent in the loop runs in plain output mode. Measured:

    line agent what the branch does when it fires
    1384 assessment exit 1 → workflow-level retry
    1754 planner exit 1 → workflow-level retry
    1805 planner corrective retry exit 1 → workflow-level retry
    2100 impl loop reset, TASK_FAILURES++, break — no verdict, no receipt (defect 2 above)
    2411 build-fix loop sets REVERT_REASON, break
    2782 eval-fix loop logs, FIX_NOOP_CAUSE
    2935 evaluator "skipping eval (build+test passed)" — fail-open
    3601 issue-response logs

    On Day 195 none of the eight fired. The consequence runs top-down: had 1384 or 1754 fired, the session would have exited for retry and the run would have been red after its three attempts — which is the honest outcome when the provider is down — instead of "no assessment" → "0 tasks" → fallback → empty diff → evaluator "API error, skipping" → promoted. The empty-diff gate (119f172b) now catches the end of that chain; this is the start of it, and it is the same rule stated eight times, all deaf to the same input.

    The trap in the obvious fix

    The runtime line, as tee'd into $TASK_LOG, is (after the optional colour escape):

      API error with no fallback configured. Exiting.
    

    two leading spaces, then the message, at line start. A bare grep -F on the message would also match a task in which the agent merely reads src/agent_builder.rs — its read_file output lands in the same log, carrying

            eprintln!("{RED}  API error with no fallback configured. Exiting.{RESET}",);
    

    and would abort a healthy task. That is the prose-contamination defect scripts/measure_abstentions.py was built around (an unanchored marker matching my own quoted text). Anchor at line start, allowing only a colour escape before the two spaces:

    ^(\x1b\[[0-9;]*m)*  API error with no fallback configured\. Exiting\.
    

    and keep the JSON match beside it. The source-echo line is the near-miss guard and must be in the fixture.

    Fix shape

    One predicate, agent_hit_api_error LOGFILE, matching either shape, called at all eight sites — one statement of the rule instead of eight copies that are deaf together. Extracted into tests/harness_logic.sh and driven against fixtures: the JSON line; the plain sentinel with and without the escape prefix; the near-miss (a log containing only the eprintln! source echo, and one containing the words "Rate limited" in prose) must not fire; an empty log must not fire. Defect 2 (the impl-loop branch records nothing) still needs the bookkeeping named above before its break.

    Stated limit, in the predicate's doc: it is a cross-boundary string read off src/ output, exactly like the [Agent stopped: grep two lines below site 2100. If src/agent_builder.rs:1446 is ever reworded, every one of the eight goes deaf again and the only symptom is the Day 195 shape. That is the argument for a source-level guard in tests/ that greps the literal out of src/agent_builder.rs and the harness and asserts they agree.

    Sequencing: land after the empty-diff gate has been observed silent on one real session (run 34749358941 is that observation), so a regression in either is attributable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions