Skip to content

Judge PR evals on cost and prompt cache changes - #618

Merged
AshishKumar4 merged 1 commit into
mainfrom
eval-cache-breaks
Oct 2, 2026
Merged

AshishKumar4 merged 1 commit into
mainfrom
eval-cache-breaks

Conversation

@AshishKumar4

@AshishKumar4 AshishKumar4 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

PR eval comments hide caching wins. #609 removed most cache misses at turn starts, but the comment showed +1 to +3 pp in cache hits, and Bonk called it unchanged.

The cache hit rate is cache reads over all prompt tokens. It also moves with how long a run is and how much new content it reads, so a real caching fix gets lost in run-to-run noise.

Changes

  • comparison.json gets cacheBreakRate per task (of the tokens the previous step had already sent, the share a step sent again instead of reading from cache), plus p-values for cache hits and cache breaks. Each is an exact Mann-Whitney U test on each run's own rate, in the direction the pooled rate moved.
  • The table keeps its columns. A cache hit change is bold when its p-value is below 0.05.
  • Bonk reports each cache change with p < 0.05, and weighs cost (already in each run's trajectory) and cache changes in the verdict. A change counts when it moves the same way on most comparable tasks and the diff explains it.

On #609's results, cache breaks fell on all 3 comparable tasks (p = 0.029, 0.002, 0.075). comparison.ts only builds reports, so no evals rerun.

@github-actions github-actions Bot added the delivery Changes to CI or release delivery label Sep 30, 2026
Comment thread .github/workflows/workshop-evals-pr.yml Outdated
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

Review: 1 finding.

Posted 1 actionable inline comment.

github run

@github-actions

Copy link
Copy Markdown

Preview: pr618-eval-cache-breaks

https://pr618-eval-cache-breaks-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

Review: 1 finding.

The existing inline finding about cache-break rates depending on run length still applies to this revision. No additional actionable issues found.

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 30, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

Review: 1 finding.

The existing cache-break finding still applies: additional fully cached steps lower the reported rate without improving cache behavior, despite the new instructions claiming independence from run length.

No additional actionable issues found.

github run

The cache hit rate also moves with how long a run is and how much new content it reads, so caching fixes were lost in its noise. comparison.json now carries a cache break rate (tokens the previous step sent that a step sent again instead of reading from cache) and per-run Mann-Whitney p-values for cache hits and breaks. The table bolds a significant cache hit change, and Bonk counts cost and cache changes in its verdict.
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. Nothing the evals run changed, so every result is reused.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 100% same inputs — 83% 4.2 25.8
chess 100% same inputs — 97% 11.2 66.9
incident-desk 100% same inputs — 93% 4.3 30.6
worker-logs 100% same inputs — 90% 3.9 20.9

Run · trajectories and raw results

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

🔬 Eval runs review

Performance

All four tasks reuse the same results on main and the PR, each passing 10/10 runs; trajectories, costs, durations and tool errors are identical. Cache hits remain 83% for change-calendar, 97% for chess, 93% for incident-desk and 90% for worker-logs, with unchanged cache-break rates and no comparable independent runs demonstrating a caching gain. Chess is the most expensive task at 66.9 model steps and 11.2 minutes per run, costing $0.0483–$0.0854; implementation and testing dominate its steps, with malformed calls and exact-text edit retries adding avoidable work.

⚪ VERDICT: NO REGRESSION FROM THIS PR

The diff changes comparison reporting and advisory-review instructions, not what the agents execute, so the reused results support neither an agent-performance regression nor an improvement.

Triage

Tool errors

  • editFile: Validation failed for tool "editFile": · change-calendar 1 → 1, chess 7 → 7, incident-desk 7 → 7, worker-logs 3 → 3 · model error — Agents omitted required arguments or used the wrong call shape despite the declared schema; change-calendar trial 2 omitted filename, then successfully retried with client.js.
  • editFile: No matching text was found in the file. · change-calendar 1 → 1, chess 7 → 7, incident-desk 1 → 1 · model error — Agents supplied text that did not match exactly; chess trial 2 added indentation absent from the inline CSS, then used grep and successfully retried with the actual substring.
  • readFile: File does not exist. · change-calendar 4 → 4, worker-logs 2 → 2 · model error — Agents tried reading files before creating them; change-calendar trial 4 read server.js and client.js immediately after creating an empty gadget, then switched to writeFile, despite createGadget documenting that default.
  • editFile: Multiple matches were found. The text to match must be unique. · chess 2 → 2, incident-desk 4 → 4 · model error — Replacement targets lacked enough context; chess trial 1 targeted a repeated gameHistory write, then successfully split the edits using surrounding statements.
  • executeCode: Failed to start Worker: · chess 3 → 3 · model error — Chess trial 1 submitted empty code, and trials 3 and 6 submitted . rather than the required self-contained module with a default export; each recovered by sending valid code.

What to do

  • No agent-side change is required in packages/workshop-backend/src/agent.ts: the existing descriptions and schemas already cover these recovered errors, and this PR leaves them unchanged.

github run

@AshishKumar4
AshishKumar4 merged commit 864de0a into main Oct 2, 2026
19 checks passed
@AshishKumar4
AshishKumar4 deleted the eval-cache-breaks branch October 2, 2026 16:30
bashandbone added a commit to knitli/knitli-os that referenced this pull request Oct 4, 2026
* Describe Google Chat bindings by what they are bound to (cloudflare#606)

An agent holding a Chat Conversation binding (a ChatSpace) took it for an account session and called searchSpaces() on it before recovering. The observation opening every Chat binding was titled "Open a Google Chat session", which reinforced the guess.

- ChatSession, ChatSpace and ChatThread now say which resource each session is bound to, as the Slack workspace, conversation and thread sessions do, and ChatSpace and ChatThread point at post().
- The opening observation names what it opened: a Chat session, conversation or thread.

* Review with GPT-6.1 Sol and opencode 1.18.33 in Bonk (cloudflare#607)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Say a new Gadget starts empty, and name missing files in editFile (cloudflare#600)

In 24 of 40 eval trials, the agent read server.js and client.js right
after creating an empty Gadget, because the system prompt said every
Gadget has both files. The prompt now says to create them, and that a
new Gadget has none unless it came from a blueprint.

editFile checked that the agent had read a file before checking that
the file exists, so a mistyped name got "You must read a file before
you can edit it." It now names the missing file. The check reuses
readFile's lookup, moved into readToolFile, and runs only when the edit
is refused. The editFile description now states the read rule.

* Let Bonk judge whether a PR improves the eval runs (cloudflare#608)

* Let Bonk judge whether a PR improves the eval runs

Bonk's verdict followed the pass-rate test alone, so a PR that only made
the agent's work better came back as "no regression". cloudflare#600 cut the runs
that read files in a new, empty Gadget from 20 of 40 to 2 of 40, and
cloudflare#601 raised cache hits on every task; both got a ⚪.

Bonk now decides from both sides' trajectories as well as
comparison.json: pass rates, wasted steps, tool errors, cache hit rate,
time and what the agent built. A change, better or worse, counts only
when the diff explains it, both sides show it, and it is larger than
run-to-run variation. The Tool errors and Prompt cache sections now show
both sides, and the cache section reports rises as well as falls.

* Call a mixed eval outcome inconclusive, not a regression

The regression rule came first and fired on any change for the worse, so a PR that also counted a change for the better could never reach the inconclusive case meant for it.

* Read the Bonk model from the BONK_MODEL repository variable

The three Bonk workflows each named their model twice, and the eval review had fallen behind on gpt-5.6-sol. BONK_MODEL is set to openai/gpt-6.1-sol.

* Bump dompurify from 3.4.15 to 3.4.16 (cloudflare#614)

Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.15 to 3.4.16.
- [Release notes](https://github.com/cure53/DOMPurify/releases)
- [Commits](cure53/DOMPurify@3.4.15...3.4.16)

---
updated-dependencies:
- dependency-name: dompurify
  dependency-version: 3.4.16
  dependency-type: direct:development
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Suggest GPT-6.1 Sol and Claude Sonnet 5.5 and bump pi to 0.99.1 (cloudflare#610)

* Upgrade pi to 0.99.1

* Add GPT-6.1 Sol and Claude Sonnet 5.5 to the suggested models

* Deflake the scheduler callback concurrency test (cloudflare#617)

The test required four 10ms callbacks to overlap by chance; on a loaded
runner the startHook and authorization round trips could outlast that
window. Hold callbacks until the concurrency bound has been reached once.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Resume agent turns after auto-approval, and fix retryAgent and Scheduler docs (cloudflare#599)

* Document that retryAgent throws while an agent is running

The implementation rejects with the agent-running error, as the other chat mutators do; the doc
said it silently did nothing.

* Show executeCode minting scheduler callbacks via env.GADGET[restore]

The Scheduler's agent-facing types told agents to call ctx.restore() from executeCode, which
targets the executeCode worker and throws that it implements no [restore](). The example also
called a bare scheduler, the Workers global, rather than env.SCHEDULER.

* Resume agent turns whose awaited actions an auto-approval drain decides

Only a manual approveAction resumed a turn suspended on awaitDecision. "Always approve" (the
chat card, Activity panel and Connections rule toggle) and approveAction's cascade apply actions
through the drain alone, so when the drain decided a turn's last awaited action -- e.g. the first
Gmail archive or a vetted MCP tool call -- the action applied and the agent never resumed or
learned of it.

Both drains now resume the chats whose awaited actions were pending on that gatekeeper. A drain
requested while one is running now shares its promise, so awaiting drain() waits for the rerun
it requested instead of returning before the action is applied. No caller awaits a drain, so a
failed resume is logged rather than ending the loop.

An auto-approvable awaited action queued behind a manual gate, which the in-order drain stops at,
also let its turn run on while the action sat unapplied. submitAction now suspends that turn too,
and rejectAction drains as approveAction does, so rejecting the gate applies the queued action
and resumes the turn instead of leaving both waiting on a drain that nothing starts.

* Resume only chats whose awaited actions the drain approved

The drain-and-resume path checked every chat with an awaited action pending on the gatekeeper,
including ones the drain left pending, so a chat whose older action stayed on a manual gate could
be resumed from its latest turn's already-approved actions.

* Migrate Worker configs to cloudflare.config.ts (cloudflare#597)

* Author Worker configs in cloudflare.config.ts

Each Worker's config is now a TypeScript cloudflare.config.ts in the
@cloudflare/config format, built from a shared factory in
scripts/worker-config.ts that holds the compatibility date, the
observability block, the capnweb-validate build and the Text module
rules the configs used to copy.

Wrangler's --experimental-new-config covers only dev, build, deploy and
versions, and rejects --config, so it can't drive the multi-worker dev
server, wrangler types, the vitest pools or the integration harness.
The wrangler.jsonc beside each config therefore stays, committed but
generated by scripts/generate-worker-configs.ts, and every existing
consumer reads it unchanged. Settings the new format has no field for
(build, rules, KV preview_id, assets.directory) go in a named
`wrangler` export, and Durable Object migrations stay tagged in a
`migrations` export, because the release manifest ships them to the
deploy service verbatim.

The generated files parse identical to the hand-written ones minus
$schema, and the golden manifest is unchanged. `pnpm configs:check`
fails on a stale file, a hand-written wrangler.jsonc with no source,
or an export the generator doesn't read; it runs in `pnpm lint` and
CI's lint job, and `pnpm dev-server` regenerates before starting.

* Share the Workers compatibility date with the vitest pools

The Worker vitest pools copied the deployed compatibility date by hand.
They now import COMPATIBILITY_DATE from @gadgets/scripts/worker-config,
so a date bump reaches the runtime the tests model too. Flags stay
per-pool, since several pools deliberately differ from the deployed
Workers. The library pools (observability, gatekeeper-kit,
typed-storage) keep their literal: they model no deployed Worker.

* Point docs and the gatekeeper skill at cloudflare.config.ts

The write-gatekeeper skeleton now shows a cloudflare.config.ts built
from the shared factory, and the skill registers a new gatekeeper in
workshop-backend's env with bindings.worker(). AGENTS.md describes the
generated wrangler.jsonc and the configs:generate/configs:check pair,
and the READMEs that pointed at wrangler.jsonc for a flag or var now
point at its source.

* Trim worker-config duplication left by the migration

- Drop dev-router's enable_ctx_exports flag, the default since
  2025-11-17 (wrangler dev warned it was redundant).
- Add DEFAULT_GATEKEEPER_WRANGLER for the capnweb-validate build plus
  .txt/.svg Text rules that 13 gatekeepers spelled out; their generated
  wrangler.jsonc is unchanged.
- The six vitest pools that model their deployed runtime exactly read
  compatibilityDate and compatibilityFlags from their cloudflare.config.ts
  instead of copying them. The router and backend pools, whose flags
  differ on purpose, keep theirs.

* Hide superseded models from pickers while they still resolve (cloudflare#611)

* Hide superseded models from pickers while they still resolve

SUGGESTED_MODELS is also the AI Gateway registry, so stored model ids
(chat authors, spawner configs, preferences) must keep resolving. A
hidden model is left out of getModelList() but resolveModel() still
finds it, and an external message to a chat on a hidden model stays on
it instead of falling back to the preferred model.

* Keep dialog actions within the viewport (cloudflare#619)

The "Verify your access" dialog lists one card per connection, and Kumo's
Dialog sets no max-height, so with enough connections its "Verify and
open" button was pushed below the fold on desktop and only zooming out
reached it. The Cloudflare account picker (one card per account) and the
Blueprint settings dialog (a textarea) could overflow the same way.

All three now use the bounded layout BlueprintModal adopted in cloudflare#436: a
capped flex-column dialog with a fixed header and footer and a scrolling
body. The mobile bottom-sheet rule for responsive-dialog still overrides
the top offset and max-height on phones.

* Public-API kernel integration tests (round 5) (cloudflare#620)

* Stop proposing a removed gadget in chats that pinned it

* Add round-5 kernel e2e scenarios

* Harden round-5 kernel tests and share fixture control helpers

* Assert only intended behaviour; keep the concurrent-approval race as an it.fails repro

* Bump @cloudflare/workers-types (cloudflare#622)

Bumps the cloudflare-toolchain group with 1 update in the / directory: [@cloudflare/workers-types](https://github.com/cloudflare/workerd).


Updates `@cloudflare/workers-types` from 5.20260903.1 to 5.20260924.1
- [Release notes](https://github.com/cloudflare/workerd/releases)
- [Changelog](https://github.com/cloudflare/workerd/blob/main/RELEASE.md)
- [Commits](https://github.com/cloudflare/workerd/commits)

---
updated-dependencies:
- dependency-name: "@cloudflare/workers-types"
  dependency-version: 5.20260924.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: cloudflare-toolchain
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump the github-actions group with 3 updates (cloudflare#621)

Bumps the github-actions group with 3 updates: [voidzero-dev/setup-vp](https://github.com/voidzero-dev/setup-vp), [actions/upload-artifact](https://github.com/actions/upload-artifact) and [actions/download-artifact](https://github.com/actions/download-artifact).


Updates `voidzero-dev/setup-vp` from 1.17.0 to 1.21.1
- [Release notes](https://github.com/voidzero-dev/setup-vp/releases)
- [Commits](voidzero-dev/setup-vp@v1.17.0...3754dd7)

Updates `actions/upload-artifact` from 4.6.2 to 7.0.1
- [Release notes](https://github.com/actions/upload-artifact/releases)
- [Commits](actions/upload-artifact@v4.6.2...043fb46)

Updates `actions/download-artifact` from 5.0.0 to 8.0.1
- [Release notes](https://github.com/actions/download-artifact/releases)
- [Commits](actions/download-artifact@634f93c...3e5f45b)

---
updated-dependencies:
- dependency-name: voidzero-dev/setup-vp
  dependency-version: 1.21.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-actions
- dependency-name: actions/upload-artifact
  dependency-version: 7.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: github-actions
- dependency-name: actions/download-artifact
  dependency-version: 8.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump yjs from 13.6.32 to 13.6.33 in the collab-yjs group (cloudflare#626)

Bumps the collab-yjs group with 1 update: [yjs](https://github.com/yjs/yjs).


Updates `yjs` from 13.6.32 to 13.6.33
- [Release notes](https://github.com/yjs/yjs/releases)
- [Commits](yjs/yjs@v13.6.32...v13.6.33)

---
updated-dependencies:
- dependency-name: yjs
  dependency-version: 13.6.33
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: collab-yjs
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump the react-and-ui group with 3 updates (cloudflare#624)

Bumps the react-and-ui group with 3 updates: [@cloudflare/kumo](https://github.com/cloudflare/kumo/tree/HEAD/packages/kumo), [@tanstack/react-router](https://github.com/TanStack/router/tree/HEAD/packages/react-router) and [@tanstack/router-plugin](https://github.com/TanStack/router/tree/HEAD/packages/router-plugin).


Updates `@cloudflare/kumo` from 2.13.2 to 2.14.0
- [Release notes](https://github.com/cloudflare/kumo/releases)
- [Changelog](https://github.com/cloudflare/kumo/blob/main/packages/kumo/CHANGELOG.md)
- [Commits](https://github.com/cloudflare/kumo/commits/@cloudflare/kumo@2.14.0/packages/kumo)

Updates `@tanstack/react-router` from 1.170.33 to 1.170.39
- [Release notes](https://github.com/TanStack/router/releases)
- [Changelog](https://github.com/TanStack/router/blob/main/packages/react-router/CHANGELOG.md)
- [Commits](https://github.com/TanStack/router/commits/@tanstack/react-router@1.170.39/packages/react-router)

Updates `@tanstack/router-plugin` from 1.168.36 to 1.168.40
- [Release notes](https://github.com/TanStack/router/releases)
- [Changelog](https://github.com/TanStack/router/blob/main/packages/router-plugin/CHANGELOG.md)
- [Commits](https://github.com/TanStack/router/commits/@tanstack/router-plugin@1.168.40/packages/router-plugin)

---
updated-dependencies:
- dependency-name: "@cloudflare/kumo"
  dependency-version: 2.14.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: react-and-ui
- dependency-name: "@tanstack/react-router"
  dependency-version: 1.170.39
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: react-and-ui
- dependency-name: "@tanstack/router-plugin"
  dependency-version: 1.168.40
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: react-and-ui
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump the editor-codemirror group across 1 directory with 4 updates (cloudflare#625)

Bumps the editor-codemirror group with 4 updates in the / directory: [@codemirror/commands](https://github.com/codemirror/commands), [@codemirror/state](https://github.com/codemirror/state), [@codemirror/view](https://github.com/codemirror/view) and [@lezer/highlight](https://github.com/lezer-parser/highlight).


Updates `@codemirror/commands` from 6.11.0 to 6.11.1
- [Changelog](https://github.com/codemirror/commands/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/commands/commits)

Updates `@codemirror/state` from 6.7.4 to 6.7.6
- [Changelog](https://github.com/codemirror/state/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/state/commits)

Updates `@codemirror/view` from 6.43.11 to 6.43.13
- [Changelog](https://github.com/codemirror/view/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/view/commits)

Updates `@lezer/highlight` from 1.2.3 to 1.2.4
- [Changelog](https://github.com/lezer-parser/highlight/blob/main/CHANGELOG.md)
- [Commits](https://github.com/lezer-parser/highlight/commits)

---
updated-dependencies:
- dependency-name: "@codemirror/commands"
  dependency-version: 6.11.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: editor-codemirror
- dependency-name: "@codemirror/state"
  dependency-version: 6.7.6
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: editor-codemirror
- dependency-name: "@codemirror/view"
  dependency-version: 6.43.13
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: editor-codemirror
- dependency-name: "@lezer/highlight"
  dependency-version: 1.2.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: editor-codemirror
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump the remaining-npm group with 7 updates (cloudflare#627)

* Bump the remaining-npm group with 7 updates

Bumps the remaining-npm group with 7 updates:

| Package | From | To |
| --- | --- | --- |
| [isomorphic-git](https://github.com/isomorphic-git/isomorphic-git) | `1.41.9` | `1.42.2` |
| [yaml](https://github.com/eemeli/yaml) | `2.9.0` | `2.9.1` |
| [motion](https://github.com/motiondivision/motion) | `13.2.0` | `13.4.2` |
| [diff3](https://github.com/axosoft/diff3) | `0.0.3` | `0.0.4` |
| [vitest-evals](https://github.com/getsentry/vitest-evals/tree/HEAD/packages/vitest-evals) | `0.16.1` | `0.17.0` |
| [@oxlint/plugins](https://github.com/oxc-project/oxc/tree/HEAD/npm/oxlint-plugins) | `1.73.0` | `1.85.0` |
| [zod](https://github.com/colinhacks/zod) | `4.5.4` | `4.6.5` |

Updates `isomorphic-git` from 1.41.9 to 1.42.2
- [Release notes](https://github.com/isomorphic-git/isomorphic-git/releases)
- [Commits](isomorphic-git/isomorphic-git@v1.41.9...v1.42.2)

Updates `yaml` from 2.9.0 to 2.9.1
- [Release notes](https://github.com/eemeli/yaml/releases)
- [Commits](eemeli/yaml@v2.9.0...v2.9.1)

Updates `motion` from 13.2.0 to 13.4.2
- [Changelog](https://github.com/motiondivision/motion/blob/main/CHANGELOG.md)
- [Commits](motiondivision/motion@v13.2.0...v13.4.2)

Updates `diff3` from 0.0.3 to 0.0.4
- [Commits](https://github.com/axosoft/diff3/commits)

Updates `vitest-evals` from 0.16.1 to 0.17.0
- [Release notes](https://github.com/getsentry/vitest-evals/releases)
- [Commits](https://github.com/getsentry/vitest-evals/commits/v0.17.0/packages/vitest-evals)

Updates `@oxlint/plugins` from 1.73.0 to 1.85.0
- [Release notes](https://github.com/oxc-project/oxc/releases)
- [Changelog](https://github.com/oxc-project/oxc/blob/main/CHANGELOG.md)
- [Commits](https://github.com/oxc-project/oxc/commits/oxlint_v1.85.0/npm/oxlint-plugins)

Updates `zod` from 4.5.4 to 4.6.5
- [Release notes](https://github.com/colinhacks/zod/releases)
- [Commits](colinhacks/zod@v4.5.4...v4.6.5)

---
updated-dependencies:
- dependency-name: isomorphic-git
  dependency-version: 1.42.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: remaining-npm
- dependency-name: yaml
  dependency-version: 2.9.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: remaining-npm
- dependency-name: motion
  dependency-version: 13.4.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: remaining-npm
- dependency-name: diff3
  dependency-version: 0.0.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: remaining-npm
- dependency-name: vitest-evals
  dependency-version: 0.17.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: remaining-npm
- dependency-name: "@oxlint/plugins"
  dependency-version: 1.85.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: remaining-npm
- dependency-name: zod
  dependency-version: 4.6.5
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: remaining-npm
...

Signed-off-by: dependabot[bot] <support@github.com>

* Keep diff3 at 0.0.3 and @oxlint/plugins at vite-plus's pin

diff3 0.0.4 has a known issue, so revert to 0.0.3 and have Dependabot
skip that version (later releases are still offered).

@oxlint/plugins is types-only and must equal the version vite-plus pins
(scripts/oxlint-plugin.test.ts); 1.85.0 failed that check against
vite-plus 0.2.8's =1.73.0. Revert it and leave it to move by hand with
vite-plus, since Dependabot always picks the newest release rather than
the pinned one.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Nathan Disidore <nathan@cloudflare.com>

* Bump node-html-parser from 7.1.0 to 9.0.4 (cloudflare#629)

Bumps [node-html-parser](https://github.com/taoqf/node-fast-html-parser) from 7.1.0 to 9.0.4.
- [Release notes](https://github.com/taoqf/node-fast-html-parser/releases)
- [Changelog](https://github.com/taoqf/node-html-parser/blob/main/CHANGELOG.md)
- [Commits](taoqf/node-html-parser@v7.1.0...v9.0.4)

---
updated-dependencies:
- dependency-name: node-html-parser
  dependency-version: 9.0.4
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Bump pako and @types/pako (cloudflare#630)

Bumps [pako](https://github.com/nodeca/pako) and [@types/pako](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/pako). These dependencies needed to be updated together.

Updates `pako` from 2.2.0 to 3.0.2
- [Changelog](https://github.com/nodeca/pako/blob/master/CHANGELOG.md)
- [Commits](nodeca/pako@2.2.0...3.0.2)

Updates `@types/pako` from 2.0.4 to 3.0.0
- [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases)
- [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/pako)

---
updated-dependencies:
- dependency-name: pako
  dependency-version: 3.0.2
  dependency-type: direct:production
  update-type: version-update:semver-major
- dependency-name: "@types/pako"
  dependency-version: 3.0.0
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Explain why the code view is read-only outside a conversation (cloudflare#633)

Code edits are made on a chat's branch, so the Code view locks editing
when no conversation is selected -- e.g. a workspace freshly created
from a blueprint, or one reopened from the sidebar. Nothing said so: the
header just read "Viewing" and the New file button was silently
disabled, which users read as the editor breaking until they sent a new
prompt.

Show "Select or start a conversation to edit" as a banner above the
editor and as the disabled New file button's tooltip.

* Bump vite-plus to 1.0.0 (cloudflare#632)

vite-plus 1.0.0 is the only VoidZero package that can move within the
24h minimumReleaseAge window without breaking the toolchain:

- vite stays 7.3.6: Vite 8's Oxc leaves Stage-3 decorators unlowered,
  so every workerd suite importing a `@validateRpc()` class fails with
  "SyntaxError: Invalid or unexpected token" (re-verified on 8.3.1; see
  the catalog comment).
- vitest stays ^4.1.11: @cloudflare/vitest-pool-workers 0.22.0 peers
  vitest ^4.1.0.
- @vitejs/plugin-react stays ^5.2.0: 6.x requires vite ^8.

Migration for 1.0:

- Override `vite@*` instead of bare `vite`, so vite-plus keeps its
  `vite` -> @voidzero-dev/vite-plus-core alias instead of having vite 7
  substituted under `vp`.
- Move task `env`/`input`/`output` under `cache` (vite-task#749), in the
  package configs and the shared task builders.
- Import the oxlint plugin types and RuleTester from
  `vite-plus/lint/plugins{,-dev}` and drop the separately pinned
  @oxlint/plugins dependency, its drift test and dependabot ignore.
- oxlint 1.85: use toSorted() on three freshly built arrays, and turn
  off react/globals for workshop-frontend tests, whose hook probes
  assign to an outer `let` on purpose.

* Bump the build-toolchain group across 1 directory with 2 updates (cloudflare#635)

Bumps the build-toolchain group with 2 updates in the / directory: [esbuild](https://github.com/evanw/esbuild) and [terser](https://github.com/terser/terser).


Updates `esbuild` from 0.28.1 to 0.28.2
- [Release notes](https://github.com/evanw/esbuild/releases)
- [Changelog](https://github.com/evanw/esbuild/blob/main/CHANGELOG.md)
- [Commits](evanw/esbuild@v0.28.1...v0.28.2)

Updates `terser` from 5.49.2 to 5.51.2
- [Changelog](https://github.com/terser/terser/blob/master/CHANGELOG.md)
- [Commits](terser/terser@v5.49.2...v5.51.2)

---
updated-dependencies:
- dependency-name: esbuild
  dependency-version: 0.28.2
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: build-toolchain
- dependency-name: terser
  dependency-version: 5.51.2
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: build-toolchain
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* Refactor / cleanup: Pull storage schemas out into their own files. (cloudflare#639)

This is the first, small step to break up overseer.ts, and probably the most obvious.

All of the typed-storage collections and TS types are now defined under a `storage-schema` subdirectory.

I generally think of APIs and storage schemas as defining the architecture of a piece of software. Everything in between is just glue, and can change more easily.

Consolidating storage schemas into one place makes it easy to tell what storage schemas are changing in any given PR.

I also moved migrations into this new directory.

Claude suggested several small related cleanups as well which I told it to go ahead and do.

* Add the eval user model only for direct access (cloudflare#634)

Since cloudflare#611, addModel refuses an id a gateway model would shadow, and the eval target added one in gateway mode, where the gateway already serves the eval model. So every eval trial failed at setup.

* Gmail gatekeeper: Fully simulate actions (cloudflare#638)

* Plan for improving gmail simulation.

* Make gmail gatekeeper simulate label changes.

This is part 1 of plans/gmail-simulation.md.

* Make gmail gatekeeper simulate sends.

This is part 2 of plans/gmail-simulation.md.

* Judge PR evals on cost and prompt cache changes (cloudflare#618)

The cache hit rate also moves with how long a run is and how much new content it reads, so caching fixes were lost in its noise. comparison.json now carries a cache break rate (tokens the previous step sent that a step sent again instead of reading from cache) and per-run Mann-Whitney p-values for cache hits and breaks. The table bolds a significant cache hit change, and Bonk counts cost and cache changes in its verdict.

* Send the system prompt as a static and a dynamic block (cloudflare#609)

pi sends the leading system message as one block, so a change to the project-specific text missed the cache from the start of the prompt. The handle now splits it after the static text, with a cache breakpoint there: on Anthropic by moving the system block's breakpoint, and on OpenAI GPT-5.6+ with an explicit prompt_cache_breakpoint.

* Upload preview secrets to the Preview base config (cloudflare#648)

The Workers API stopped accepting `preview_defaults`, the field the pinned draft Wrangler writes for `wrangler preview secret bulk`, so every preview deploy failed at the first worker with secrets. The secrets now go through `wrangler preview base-config secret bulk` on the workspace's own Wrangler, which writes `previews_base_config`.

Everything else stays on the draft build: neither sibling preview bindings nor per-preview KV and R2 provisioning is in a released Wrangler, checked against 4.138.0 and 4.147.0.

* Public-API kernel integration tests (round 6) (cloudflare#647)

Cover user journeys a kernel refactor could silently break:

- the agent's system prompt follows the workspace's gadgets, bindings,
  ambient connection and standard formats, and never another chat's
  pending binding
- a pasted link (capsule) becomes a binding for that chat only
- a chat attachment reaches the model and stays in later turns
- one accept commits every gadget a chat built, each with its own code
- gadget console logs reach subscribers labelled draft or mainline, stop
  on dispose, and never reach a use collaborator
- a gadget's LLM binding runs on its bound model; a blueprint install
  uses the installer's model
- admin instance instructions and format hints reach the agent, not the
  user
- disconnecting an account from the Connectors page removes and revokes
  it

Adds systemPromptOf() to the mock model for reading recorded prompts.

Also makes bare `.rejects.toThrow()` assertions on RPC promises able to
fail. vitest's `.rejects` calls a callable subject, and a Cap'n Web
RpcPromise is callable: calling it pipelines a call on the result, which
rejects with "'' is not a function." when the original call succeeded,
so those assertions passed whatever the call did. They now match the
refusal message; two were also wrong, since writeValue resolves once the
write is submitted for approval, and now assert that instead.

* Restart a workspace whose loop counter is exhausted (cloudflare#640)

* Restart a workspace whose loop counter is exhausted

The Workers runtime refuses a Durable Object call once the loop
counter behind it is spent ("Subrequest depth limit exceeded. This
request looped back into the Workers runtime too many times."). A
workspace object's outgoing channels can end up holding a spent
counter with nothing recursing, and from then on every call it makes
to a user object is refused until the instance is replaced.

When one of the workspace's own user-object calls is rejected that way
(a call through a wrapped stub, the last-active bump, or the outputs
sync), the workspace now schedules the existing access restart. At
most one restart per instance, and none in an instance's first 60
seconds. Errors thrown by gadget, agent or gatekeeper-facet code are
never consulted.

* Let admins manage a deployment's AI Gateway models via admin panel (cloudflare#616)

A deployment could only change which models its AI Gateway offers by patching SUGGESTED_MODELS. That patch goes stale whenever the catalog changes, and some models can't go in the public catalog at all.

In AI Gateway mode, /admin gets a new Models tab:

Enable / test providers and set default reasoning for the deployment and more!

* feat(google-gatekeeper): begin Google Chat DMs and group chats from the account binding (cloudflare#649)

* Start Google Chat direct messages and group chats from the account binding

The whole-account Chat session gains two methods:

- searchPeople(query) searches the connected account's Workspace
  directory (domain profiles only, never contacts), recording each page
  as an observation.
- sendDirectMessage(people, text) sends to one person (a DM) or 2-49
  people (a group chat with exactly them). An existing conversation is
  an ordinary send. Otherwise the message is queued as a new
  "chatStartConversation" action kind, separately auto-approvable from
  sends, and every person must be named by email and resolve to a
  directory profile: outsiders and yourself are refused before anything
  is queued.

Applying a start creates the conversation idempotently (spaces.setup
with a requestId, remembered once created), reuses a group chat that
appeared since queueing, and verifies membership before posting, since
Google silently drops anyone who blocks the caller from a new group
chat. Until committed, the message carries a temporary
pending:space:{id}; edits queued against it carry over into the real
conversation, replies are refused. Undo deletes the message.

The account resource adds chat.spaces.create (not chat.spaces) and
directory.readonly, so existing account connections re-consent.

* Fix Google Chat sends whose new conversation exists before they post

A send that starts a direct message or group chat is queued under a
temporary pending:space name. If applying it set up the conversation but
the post then failed, two things went wrong:

- ChatSpace.listMessages() for the new conversation left the queued send
  out, so an agent could conclude it was never sent and send it again.
  listForSpace and resolveMessage now resolve the temporary name to the
  conversation, which also lets the send's capability pass that
  conversation's scope check. The send reports the real spaceId once its
  conversation exists, and the agent-facing docs say so.
- A failed member check after a successful setup, such as a 503, left the
  send marked as possibly sent, so it could not be rejected until a retry
  succeeded. The setup's mark is now cleared as soon as the conversation
  exists: an empty conversation shows nobody anything.

A retry after the conversation is recorded still skips setup, so a send
whose post may have landed stays unrejectable. A new test pins that.

* Export gatekeeper-kit's SingleFlight and use it for Chat conversation setup

SingleFlight coalesces concurrent work by key and releases each flight
once it settles. It was internal to the kit; it is now the
./single-flight subpath, listed in the README inventory and noted in the
design record.

The Google Chat gatekeeper used a hand-rolled map of promises to make
sends to the same people, applied at once, share one conversation setup.
It now uses SingleFlight, whose own tests cover joining and release.

* Re-check a new Google Chat conversation's members before retrying an unsent post

A send that starts a group chat records the conversation once it is set
up. If the post was then refused outright, a retry reused that
conversation without checking who was in it, so someone who joined in
between would receive the approved message. A retry that definitely
didn't post now refuses when the members have changed; one whose post
may have landed still goes back to the same conversation, so it stays
idempotent.

The directory lookup that confirms each person in a new conversation
read one page of matches and called anyone not on it outside the
organization. When Google reports further pages it now says it couldn't
confirm them instead.

* Check that Google Chat sends reach exactly the approved people

A new conversation's send checked only that nobody approved was missing,
and several paths could still reach someone who wasn't:

- A replayed spaces.setup returns the group as it is now, so after a
  lost setup response someone added since would receive the message.
  The send is now refused when the conversation holds anyone extra, and
  stays rejectable.
- A retry after a post that may have landed skipped the member check.
  Every retry that reuses the recorded conversation now checks it first,
  leaving the attempt mark as it is, so an uncertain send stays
  unrejectable.
- findGroupChat trusted a matching member count, but Google's fallback
  for a block can offer a group with someone else in it. Each requested
  person is now confirmed, by user ID against the member list or by
  email through members.get, and a membership Google reports as
  NOT_A_MEMBER doesn't count.
- peopleIn read only the first page of members. It now follows every
  page, since Google may return fewer members than asked for.

* Simplify Google Chat conversation starts

A send that creates its conversation is queued under a temporary
pending:space: name. That name was resolved at five call sites; the
store now resolves it once, in get() and list(), so every reader sees
the real conversation once it exists. The guard that keeps such a send
from being replied to keyed on the temporary spaceId, which missed the
window after setup but before the post; it now checks the stored send.

openConversation returns as soon as Google's lookup finds the group
chat, since the lookup already confirmed exactly who is in it, rather
than listing the members a second time. Its post-setup check is one
early return and one throw instead of a nested ternary.

The directory search reads its response through readGoogleJson, as the
other People API call does, so it is size-bounded and logs Google's
reasons; the #request/#fetchJson split that existed only for it is
gone. The 400/404 "no such user" check shared by three lookups is one
predicate, and #onlyIn no longer reads a membership for a users/{id}
reference that is already absent.

The tests drop unused backend state and share the group-chat setup and
Chat-app membership they repeated.

* Address Codex P2 review findings on sync PR

- VoiceSettingsDialog: clear the saved speaker when switching the
  conversation TTS model, so validation resolves the new model's default
  voice instead of rejecting the whole update.
- overseer #drainAutoApprovalsAndResume: snapshot non-agent awaited actions
  too and fan out per approved action as approveAction does, so a drain that
  auto-applies a user/OpenAPI-caller action resumes the turn waiting on it.
- overseer loop-limit recovery: route every User-DO call's rejections through
  recovery -- wrapUserDo at 14 construction sites (get/getByName) plus
  explicit restartIfLoopLimited in the catches covering only User-DO calls.
  Documents the invariant on restartIfLoopLimited.
- user getModelReasoning: borrow behavesLike levels for added models the
  runtime doesn't know, matching gatewayCatalogModel, so the composer shows
  their effort selector.

Tests: 2 dialog cases, drain-resume (negative-controlled), behavesLike
borrow/no-borrow/own-wins, 8 loop-limit routes (spot-negative-controlled);
openFakeOverseer gains a ctx passthrough for first-open coverage.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Nathan Disidore <nathan@cloudflare.com>
Co-authored-by: Maximo Guk <62088388+Maximo-Guk@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ashish Kumar Singh <ashishsingh@cloudflare.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Kenton Varda <kenton@cloudflare.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

delivery Changes to CI or release delivery

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants