Conversation
…ttended rule for routine runs (phase 2 parts 3 and 4) With tools.deferred on, the agents proxy lists a small core plus search_tools and use_tool; every other tool is found by what it does and called by name, on every engine, with nothing depending on tools/list_changed. Routine runs are told nobody is watching: take the most reversible reading, say which, and on a missing sign-in, file, tool or permission stop with a short failure summary (decision 13). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… in routine instructions and delegation briefs (phase 2 part 4) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…se 2 part 4, §16c) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…streak on the routine (phase 2 part 4) A recurring routine that comes due while its previous run is still going now counts the skip on the routine (skippedRuns, lastSkippedAt) instead of silently not running; overlap: "queue" lets the next run wait instead; failureStreak counts consecutive failed runs and a completed run clears it. Parts 3 and 4's built and deferred items are recorded in the spec. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ap, deferred tools (−22% input per call), mention tokens Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Closing at the maintainer’s request: we are deferring the Phase 2/3 harness stack to keep the current release scope smaller and reduce regression risk. This is a scope decision, not a finding that every change here is defective. The branch is being preserved so focused changes can be revisited separately. |
Phase 2, parts 3 and 4 of the harness programme (
docs/plans/2026-09-15-phase-2.md). Stacked on the part 2 PR. Spec, with every item marked built or deferred and why:docs/superpowers/specs/2026-09-15-phase-2-tools-and-turns-design.md.In plain terms
Deferred tool loading (part 3), behind
tools.deferred. With it on, a bot sees six core tools plussearch_toolsanduse_toolinstead of every tool schema on every call. It finds the rest by what they do and calls them by name; nothing depends on an engine honouringtools/list_changed, so it works on every engine. Off by default.Turns that run unattended (part 4). Routine runs are told nobody is watching: take the most reversible reading, say which, and stop with a short failure summary when a sign-in, file, tool or permission is missing (decision 13). Instructions and delegation briefs may carry
@name [[omb:type:id]]references that survive renames and resolve to one line with the id at run time. A recurring routine that comes due while its previous run is still going now counts the skip on the routine instead of running silently over it;overlap: "queue"lets the next run wait; a routine carries its failure streak. A bot's engine preference travels in a package export as setup intent (§16c).Deferred, and why (in the spec)
The tool ladder (#1009: its first commit conflicts in four files on this chain, two of them UI; it stays its own PR to rebase), remote HTTP MCP servers (stdio-only through the registry, the probe and every driver's mount; a per-driver change to verify on real servers), structured
ask_user(no harness tool exists to extend yet; #1184's option card is the surface),cancel_requested/dead_letter, F4's native structured output, and the five-section description lint over every tool (the four new tools follow it).Measured
Proxy suite (a deferred child lists a core of under ten, finds
list_roomsfrom "which rooms exist", calls it throughuse_tool, rejects a guessed name);mentions.test.ts;routines.test.ts(tokens resolved at run time with the definition untouched; a skipped occurrence counted and persisted; a queued second run; the failure streak counting and clearing);package-export.test.ts. The scorecard is unchanged because every switch is off by default.Every engine
All harness-side (the proxy, prompt text, routine records, the export): full on every engine.
Platforms
Checks
pnpm typecheck,pnpm lint,pnpm i18n:checkclean; full vitest suite on this tip: 599 files, 7,182 passed, 54 skipped, 0 failed.Live check (Sep 16, standalone instance, real Claude)
Deferred tools: asked which rooms exist, Claude called
search_tools, thenuse_tool→list_rooms, and answered correctly; with the flag on, input per call fell from 43,634 (main) to 33,968 on T2 turn 1 (−22%), T1–T4 all correct. Mention tokens: a routine's run read "Worker (bot …)" with the resolved references block and completed. Record:docs/bench/scorecard/2026-09-16-phase2-live.md.🤖 Generated with Claude Code