Repository navigation
Merge cloudflare/cloudflare-os main (6eb11120) into p5 - #5
Merged
Merged
Conversation
* Add descriptionIsComplete to ActionDescription A gatekeeper sets the flag to assert that the description reproduces, verbatim, every piece of workspace-originated content the action will write or send. Absent means a summary or truncated text. A workspace that has read restricted data will accept only complete actions; this commit only adds the field, the submit-time gate follows separately.
…flare#562) The max-cell-count export test guarded worksheet batching only through vitest's default 5s timeout. Unbatched export still passes locally (~4.2s), while batched export exceeds 5s on loaded 4-CPU CI runners. Count the chunks fed to CompressionStream instead, and give the test a 30s timeout so runner load cannot fail it.
Upgrade pi to 0.87.1 to add Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra, Sol, and Luna to Workshop’s suggested models, with catalog-backed cost estimates and the provider settings needed to run them correctly. Keep GPT-6’s 1.05M context window and 128K output limit, using 272K as the preferred compaction budget. Quick requests use low effort on Claude models that require adaptive thinking. pi 0.87 now expects agent instructions and tool descriptions in the conversation, and asks Workshop whether to continue before signaling the end of a turn. We updated the shared agent loop to supply that context and capture pending approvals in time to save them with the turn. This preserves guidance for every model and lets agents pause for approval or a new connection with their work recorded.
) When you open a workspace someone shared with you, a dialog asks you to confirm that your own accounts can read the data it uses. This fixes two ways to get stuck in it. Clicking Connect, re-authenticate, or grant access disables that button until the sign-in finishes. If you close the pop-up, decline consent, or sign in with an account you've already connected, the dialog is never told, so the button stays disabled and you can't try again. This PR only disables the button while the sign-in is being started, which is how the other connect dialogs already work. The downside is that clicking again while a pop-up is still open starts a second sign-in, same as in those dialogs. If a gatekeeper refuses an account whose credentials are fine, usually because that account can't see the document or calendar the workspace uses, the dialog still shows the account as Ready and tells you to re-authenticate, which won't help. This PR drops the Ready badge for that account, and the message now also suggests asking the workspace owner to share the data. The re-authenticate button is still there, just less prominent, because some gatekeepers refuse on an auth error without marking the account as expired. This doesn't change what the server does when you sign in with an account you've already connected. That needs a separate fix.
* Add typed ActionField values to action and observation descriptions ActionDescription.fields carries the values an approver reviews as data (inline, text, json, list, file), so surfaces can show them literally instead of relying on Markdown fences. descriptionIsComplete now covers description and fields together; a file field counts as shown only when its bytes come unchanged from the provider. Suggested by Nathan in review of cloudflare#541. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Emit structured fields from the action description builder ActionDescriptionBuilder now collects ActionFields instead of fenced Markdown blocks, leaving description to the gatekeeper's prose. Values a field cannot show exactly are rerouted to escaped JSON (text with non-emoji invisibles or bare CRs now stays complete this way), truncation is recorded as shownBytes/totalBytes, and the new file() method names bytes by size and digest with a provider/agent origin. Gmail attachments and Confluence uploads become file fields; the kit action set and MCP sessions forward fields to the approval queue. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Show action fields literally on every approval surface A new ActionFields component renders ActionDescription.fields after the description's prose, never through Markdown: inline values and list items as code, text and JSON in scrollable blocks with a syntax caption (and a CRLF note when the value has CRLF breaks), files as a card with name, type, size and SHA-256, plus truncation and omission captions. The chat's pending, blocking and expanded action cards, observations, and the Activity review and history rows show the fields; the Activity review row and the notifications popover show a field count when collapsed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Preserve whitespace in list items and file names The kit marks a list item, file name or media type complete even with runs of spaces, tabs or edge whitespace, but they rendered with the default white-space, which collapses them: a recipient display name with two spaces read as one. Render them with whitespace-pre-wrap like text fields. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Build file fields from their known members only A caller's object can carry properties FileDescription does not admit; spread last, a stray label, kind or truncated overrode the builder's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Show file names with invisible characters escaped in the card A file name or media type with a control, zero-width or bidi character was marked complete yet led the card raw, so invoice<RLO>fdp.exe read as invoiceexe.pdf. The card now shows such a value as an escaped JSON string with a note, which makes the kit's separate escaped-name field redundant, so file() no longer adds it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When an approved MCP tool call failed after being sent, the action store closed it for good, then also refused the discard the Workshop offers for a failed apply. The approval stayed pending with no way to clear it, and since the agent awaits MCP calls, its chat stayed blocked. Discarding a failed call now succeeds. The failure stays on record for the Gadget to collect, and the call can no longer be sent.
* Support reading tables in Google Docs * Preserve links in Google Docs tables * Harden Google Docs Markdown edits * fix google doc paragraph boundary replay * Group Google Docs Markdown write requests Clear inherited text style with one request over the inserted text, and share paragraph style and bullet removal requests across adjacent paragraphs that need the same change, so plain multi-paragraph edits cost two requests. Remove bullets before resetting indentation in every path. Derive the request cap from the Markdown complexity limits. * Preserve unchanged blocks and list depth in Google Docs Markdown Align rewritten blocks with their sources by longest common subsequence, so an unchanged paragraph between inserted blocks keeps its title style or markerless list. Pad skipped list levels in table HTML with markerless items. Model the indent Google leaves after removing bullets, and assert the contextual insertion preview. * Bound Google Docs block alignment for bulk rewrites Match unchanged leading and trailing blocks linearly, build each source block's text once, and run the full alignment only while the changed middle stays under a fixed pair budget. A larger middle pairs by position when its counts match, so a rewrite near the block limit no longer builds a quadratic table. * Rewrite multi-paragraph Google Docs edits paragraph by paragraph A replacement spanning several paragraphs used to delete the range and insert it again, so every paragraph inherited one surviving paragraph's list and style. Edits mixing a list with surrounding prose were refused. Pair each source paragraph with the block that rewrites it, then leave unchanged paragraphs alone, rewrite edited ones in place, delete removed ones, and insert new ones beside a paragraph of the same shape so they inherit its list. Requests run from the end of the range backwards in one batch. A paragraph merged into the tab's undeletable final break has its own style restored and any inherited bullet removed. * refactor: simplify the logic * Fix Google Docs separator inserts, cell rules, and large alignments Inserting a paragraph before paragraph text with a trailing blank-line separator wrote an extra empty paragraph, so the applied document no longer matched its preview. Table cells dropped text beside a horizontal rule. Beyond the alignment limit, a rewrite's changed middle had no anchors, so unchanged paragraphs could lose their title style or list. They are now anchored by blocks unique on both sides. * Apply the Google Docs replacement its approval shows Approval wrote the raw requested Markdown while the preview, simulation, and replay used its canonical form, so an escaped edit that previewed as a no-op still rewrote the text. Approval now writes the canonical form. The canonical form can start a fragment with a blank-line separator after paragraph text, which wrote two paragraph breaks. A separator beside paragraph text now collapses to one break at either end. * Write Google Docs paragraph structure exactly as previewed Blank runs parsed as one break too few, so edits containing empty paragraphs wrote the wrong number of them. Parsing is now the inverse of rendering (each pair of blank lines is one empty paragraph), and inline fragments size their edge breaks against the surrounding newline runs. Structural edits whose preview disagreed with the written document: - deleting a paragraph break (merges, dropping empty paragraphs) threw; it now takes the whole-block rewrite - odd separator runs are rounded to a renderable length, and separators are no longer measured from inside a newline run - block rewrites size leading newlines against the separator before them - paragraphs split off a list item or heading no longer inherit its style - a rewrite creates bullets once over its final layout, so paragraphs turned into list items form one list Bumps MARKDOWN_RENDERING_VERSION to 5. * Append plain Google Docs paragraphs as previewed An append inserted its paragraphs after the tab's last one without resetting them, so text appended after a heading, title, or list item inherited that style or bullet while the approval showed plain text. Appends now reset the new paragraphs and leave the existing one alone, as inserts after a paragraph already do. The append preview now joins through the same canonical boundary rules as replacements, so a list item appended to a same-type list previews as the one list Google produces. The test model only creates list state when bullets are added, matching Google, so removing bullets from a text tab no longer drops its links. * Assert Google Docs approval values as structured fields Since cloudflare#565 the description builder emits each shown value as an ActionField rather than a fenced block in the description, so the tests that checked approval text now read the submitted fields.
…iews the full text (cloudflare#487) Once a workspace has read data marked containsRestrictedData, main refuses every action outright. This replaces that with a stop-gap which leans heavily on human approvals. For gatekeepers that have observed restricted, every action pends for manual approval and can never be auto-approved, and the approver is shown the action's full text with a notice that they are responsible for checking it contains no restricted data.
* Replace the eval tasks with four multi-turn tasks The three tasks were single-request contracts in toy domains, and the one that looked saturated (project-doc, 8/10) was failing on a verifier that rejected a literal reading of its own prompt. None was long-horizon, only one had a second turn, and nothing checked whether the agent's words matched the data. Four tasks replace them, all multi-turn, verified through the Gadget's RPC against references the verifier owns: - change-calendar: a maintenance-window calendar built to an exact contract, written up as a Document from its own data, changed under new rules, committed and reloaded, then asked a question whose answer is in the seeded schedule. - worker-logs: a Workers request-log analyser fed a seeded day of events with one planted bad hour; filters and rankings are added, then the agent is asked which Worker had the worst hour, in a fixed reply shape checked against the data. - incident-desk: an on-call desk where twenty responders acknowledge at once and exactly one may win, then escalation and metrics. - chess: a complete engine written without a library, then PGN, then draw detection, each checked differentially against chess.js on curated positions and seeded random games. Either en passant FEN convention is accepted. The verifier now sees the agent's chat replies for the question turns, and the verification budget grows to four minutes: these checks make hundreds of RPC calls a turn. * Regenerate the lockfile after the rebase The rebase onto main left pnpm-lock.yaml referencing a vitest 4.1.11 entry that no longer existed, so vp install failed with ERR_PNPM_LOCKFILE_MISSING_DEPENDENCY. Rebuilt from main's lockfile; the only difference is the chess.js dependency this branch adds. * Close the verifier gaps the review found Each check now fails for the wrong implementation it was meant to catch. change-calendar: conflict ids are compared element by element, not through a joined string; the plan needs one distinct bullet per window; the seeded-windows check also asks conflicts() over a range that overlaps both api-gateway windows, so the RPC is exercised after the rule change and after commit. incident-desk: every incident turn 1 left is compared on service, severity, summary, owner, timestamps and escalations, not just id and status; a race winner's owner must be the responder, and UNKNOWN_INCIDENT and ALREADY_RESOLVED must carry owner: null; escalation is exercised on an acknowledged incident and must keep its owner; metrics are checked on a service holding an open, an acknowledged and a resolved incident, and the empty service on all five fields. worker-logs: failures now include 501 and 504; three worker+route pairs share a planted slow tail so slowestRoutes has to break the p95 tie by worker then route, with a module-load guard that the data still does; turn 2 resets and re-ingests to prove reset() survived the edit. chess: a refused move must also leave the game's PGN unchanged once PGN exists; six plausible six-field FENs must be refused; Black castling, en passant and promotion have fixtures the oracle is checked to offer; the post-draw random game compares the draw fields. * Check the game record after a rejected PGN and the winning open's severity A refused loadPgn must leave pgn() as it was, not only the position. The one open() that wins the simultaneous-opens race must have stored the severity it was sent, not a default. * Verify each check from the contract, not from one correct gadget One pass over the four verifiers with three rules: after a mutation, assert the return and re-read the whole record; give every parser one input shape a lazy parser gets wrong; for rules that interact, one case where the order matters. Each check stays satisfiable by any correct reading of its prompt: the offset-timestamp probe asserts only that scheduling fails, since the prompt fixes no code for it, and K+NN vs K is left out because the prompt does not define insufficient material and engines disagree. incident-desk: the turn-one predicate covers inc-1 after resolve too; escalation must change severity and the count and nothing else, and is read back after every call including the refused one; race-open's severity is checked against the attempt its summary names; the opens a check depends on fail the check when refused. chess: compareHere also reads fen(), so the stored position is checked wherever moves and status are, including after each special move; the PGN self-comparison strips tag pairs so a Gadget that adds a Date tag is not failed for it; the repetition and fifty-move moves are asserted. change-calendar: the touching window is mw-100 so conflicts() has to sort; two new windows in one turn 22 hours apart; a +05:00 start that is outside hours in UTC must fail; a bullet may state its length as hours and minutes. worker-logs: the planted slow tail sits in one colo, so a colo-filtered ranking differs from the day's; one summary range has endpoints inside an hour. * Check state after refusals, grandfathered windows and loadFen resets incident-desk: a refused escalation of a resolved incident must leave it unchanged, and one of an unknown id must not create it. change-calendar: seed a six-hour billing window, valid under turn 1's cap and over turn 2's, so a migration that drops windows the new rules would refuse fails "everything already scheduled stays". chess: after the start position has occurred three times, loading it again must not report a repetition, and the loaded fifty-move position must match the oracle before the move that draws. * Run each eval task on its own runner A run took as long as the sum of two lanes of slowest trials, because two eval files shared one runner at a time: chess 316 s plus change-calendar 135 s made the last run 451 s. Each eval file of each measured revision now gets its own runner, so a run takes as long as its slowest task, and a runner hosts 10 Workshops instead of 20. A plan job lists the eval files at each revision from a sparse checkout, and an assemble job per revision joins the per-task reports. It fails unless it has exactly the planned files, since a missing task would otherwise read as removed and be cached that way. The push baseline goes through the same jobs, so main and pull requests are measured the same way, and the cache key records the layout so older baselines are measured again rather than compared against. A failed candidate no longer stops a complete baseline from being cached. * Measure evals on GPT-6 Luna Makes GPT-6 Luna, which cloudflare#556 added to the suggested models, the model eval baselines are measured on. * Probe DUPLICATE_ID with a window that breaks no other rule The probe resubmitted mw-101 unchanged, so it also overlapped mw-101, and the prompt does not say whether DUPLICATE_ID or OVERLAP wins. GPT 6 Luna answered OVERLAP in 3 of 10 trials, a defensible reading the check failed. The probe now reuses mw-101's id at a time that overlaps nothing. * Close the gaps in Bonk's last two reviews chess: turn 3 checks that malformed FENs are still refused, and reaches insufficient material by a capture as well as by loading; PGN export must replay to the played position, not only the same moves; imports include games ending 0-1 and 1/2-1/2. Before turn 3 defines draws, only a dead position may list no legal moves. incident-desk: a turn-two open reads back exactly as submitted. change-calendar: no bullet anywhere in the plan, including before the first heading, may list a window outside the week. worker-logs: events arrive in a seeded random order, so ordering by first appearance fails. * Show eval pass rates as bars and mark each change's direction The comparison comment opens with every cohort's pass rate as a ten-cell bar, green for passed and red for failed, compared or not, and marks each pass delta green, red or grey. The Bonk review gets a coloured verdict, marks deltas as rising or falling, and copies the bars from the comparison instead of drawing its own, so its numbers stay the comparison's. * Lead the eval results comment with a verdict and failure details The comment now opens with one verdict (improved, regressed, unchanged or inconclusive) and the files the evals exercise that the PR changed. The score table shows pass bars for every task and colours a pass-rate change only when it is significant. Under the table, each task with failures lists its failing checks, the first failing evidence, its most common tool errors, and infrastructure errors apart from the agent's. Bonk reuses the computed verdict instead of deciding its own. * Keep the eval results comment short Show file names rather than paths, one tool error per task rather than five, and no raw evidence line; comparison.json keeps the evidence for Bonk. Drop the red marker on check counts, which went up on noise. * Key the eval baseline by the files the evals exercise A baseline was cached per commit, so every main commit needed its own measurement even when it changed nothing the evals depend on. Key it by the blob ids of the workflow's trigger paths instead: commits that leave them alone share one measurement, whichever PR or push made it. Bonk's diff now starts at the PR's base, since the baseline may be older. Record why each eval file keeps its own runner: all four on one runner ran it out of memory. * Have Bonk explain why each task failed Bonk's review stopped at 'nothing is comparable' on any PR that changes the tasks, and otherwise named one difference. It now gives one line per task with failed runs: what the agent did wrong, one cause (system prompt, tool design, harness bug, verifier or model error) and a fix, reading the system prompt and tool definitions to tell them apart. Key the baseline by commit again. Keying it by the trigger paths' content reused baselines across main commits that changed code the evals run but the paths do not list, such as files overseer.ts imports and the lockfile. Keep the note on why each eval file has its own runner. * Post eval results at the end of the PR, replacing earlier ones The results comment was edited in place, so after a few pushes it sat far above the conversation, and every run added another Bonk review. Each run now deletes this workflow's earlier results comments and Bonk's earlier eval reviews, then posts the new results; Bonk's review follows. Bonk's code reviews and people's comments are left alone. Let Bonk grep and list files as well as read them: a run's trajectories are about 25,000 lines, too long to page through. * Share one Workshop across an eval file's trials Each trial started its own Workshop, about 340 MB, so 40 trials ran a 16 GB runner out of memory and every task needed its own runner. Now an eval file starts one Workshop per model and its trials share it, each with its own user and workspace. With all four files running at once on one runner, the probe peaked at 7.7 GB with no infrastructure errors in 40 trials, so Vitest now runs every file at once. The Workshop closes after the file's last trial and before the network filter comes off, so no worker outlives the filter. * Reuse each task's eval result until something it runs changes Every push measured both sides of the comparison, and every relevant merge measured main again, even when nothing a task runs had changed. Now each task's result is stored under a key of the files its run executes: the Worker's inputs, the harness, the shared helpers, the task file, and the workflow's eval settings. A run measures only the tasks with no stored result, in at most one job per side, and a key both sides share is measured once. The comparison and Bonk's review run again only when their own code or a task's key changed, or on a manual re-run. The rules that decide whether a result is stored move into results.ts, which is part of every key, as is the validator script. Whether a task is comparable, and which files differ, now come from the keys instead of a git diff, so a task's id must match its file name. A newer push waits for the run in progress instead of cancelling it, and the results comment is posted before the earlier ones are deleted. * Keep test edits out of the report key and say results are reused Editing a test under scripts/evals changed the report key, so it reposted the same results and ran Bonk again. Test files now change no key. The results header said each task ran once per commit, but a task's result can come from an earlier run with the same inputs, so the header now says so. The key's comment now states the rule the code follows: files that can change a result, not every file a run executes. * Simplify the eval results comment and restructure Bonk's review The results comment is now the verdict and one table. Each row gives the task's score on both sides, the change, Fisher's exact test, and the average minutes, cost and steps per run. The file list, the commit line, the pass-rate tiles and the per-task failure breakdown are gone, and the eval keys no longer list changed files, since only that line read them. Headers are short and values do not break, so the table fits a pull request comment. Bonk's review now gives performance, failure modes with their cause, whether this PR caused each one, and what to do, then the verdict in caps. A failure is this PR's only when the code at fault changed, and REGRESSION needs a significant fall in comparison.json. The heading is main's "Eval runs review" again, so reviews that main posted are still replaced. * Show scores as percentages and close two verifier gaps The results table shows each score as a pass rate, for example 90% → 70%, next to the change in percentage points. The plan document's own title must now be the name the prompt asks for. Turn 2 of incident-desk must give back every turn-1 record exactly, including who won each race and when; it may only add the escalation count. Each trial keeps its turn-1 board by its Desk, because trials share the module. All ten real trials of each task already pass both checks. * Pool turn one's incident records instead of keying them by Workpiece id Workpiece ids are workspace-local, so every trial's Desk is id 0 and the ten concurrent trials overwrote each other's turn-one snapshot: eight of ten turn-two checks compared against another trial's board. Turn one now adds each record to one pool, and turn two requires each of its records to be in it. A record holds its own trial's timestamps and race winners, so one that changed matches no trial's copy. Replayed on the real boards of run 36057910489: the old code flags all seven records in 8 of 9 intact trials, the new code flags none, and it still flags a swapped owner, a 1 ms timestamp shift, reset timestamps and a dropped record. * Keep a closing session connected until its workspace's abort lands Closing a trial's session deletes its workspace, and deleteSelf() aborts the workspace DO about 100 ms later. close() disposed its stubs at once, so the abort found no client left. When the DO has started a Gadget facet, that segfaults local workerd. Miniflare restarts it, and every other session on the same Workshop loses its socket: turns reconnect, but a verifier call in flight fails and counts as the agent's failure. With one Workshop per eval file, most trial closes crashed it (33 times in one 40-trial run). close() now waits for the Workshop to drop the session, which the abort causes, before disposing. A drop that never comes fails close() instead of passing silently. The target.ts comment said this crash had not happened; it now states what sharing a Workshop costs. A model-free repro on Node 24 goes from 10 crashes and 24 failed in-flight calls in 24 closes to none. close() takes about 110 ms. * Move the eval workflow's plan, join and post steps into TypeScript These steps were about 120 lines of bash and jq inside the workflow, with no tests. scripts/evals/workflow.ts now holds them as subcommands, and workflow.test.ts covers the decisions a plausible bug would break: which stored results are reused, which revision measures a shared key, when the posted comments are current, and which comments a new one replaces. On valid input the behaviour is the same. On real stored results and on crafted commits, plan wrote byte-identical outputs in every scenario and join wrote byte-identical results files. Run against a stand-in for the GitHub API, post built the same comment and deleted the same comments. The measurement settings are unchanged, so every stored result stays valid. On malformed input the script no longer guesses: an artifact with no recorded run is not reused, a clean list that does not parse fails the join, and a comment with no body is skipped. Only Node built-ins are imported, so the plan job still needs no install step. eval-keys.ts now exports evalKeys(); its command line is unchanged. * Remove the JSON reader layer from workflow.ts workflow.ts had one reader per JSON shape (field, stringField, numberField, listOf, revisionOf, sourceOf, artifactOf, commentsOf, readJson) that copied each input into a typed object before use. Now one isRecord guard narrows each input where it is read, and a step checks only the fields it uses. download() is inlined, and a directory check replaces the set of downloaded artifacts. A malformed artifact or comment from GitHub is now skipped instead of failing the step. Plan, join and post outputs stay byte-identical to the old bash in every scenario. * Keep Workshop crashes out of stored eval results A check that failed because the shared Workshop crashed counted as the agent's failure, and later runs reused it from the stored results. WorkshopAgentSession now counts dropped connections. When a check fails while one drops, EvalVerifier.collect throws, so the harness records a run error, which is never stored. A timeout is no longer the only way verification ends early, so its check ids become *.incomplete. * Leave test files under evals/ out of task keys Test files elsewhere already stay out of the keys. One next to the tasks would change all four keys and force a full re-measure.
…are#567) workerd 1.20260921.1 includes cloudflare/workerd#7370, which fixes a facet channel use-after-free. The lockfile had workerd 1.20260831.1 and 1.20260801.1, both older than the fix. This only affects local runs and the wrangler that bundles releases.
* Declare miniflare once in the pnpm catalog The three gatekeepers that depend on miniflare and the vitest-pool-workers override each pinned 5.20260921.1-alpha by hand. Point them all at a single catalog entry so a bump is one edit. * Declare zod and typescript6 once in the pnpm catalog zod (^4.5.4) and the typescript6 alias (npm:typescript@6.0.3) were each pinned by hand in several packages. Point them at single catalog entries so a bump is one edit and the packages can't drift apart. * Drop the pi 0.87.1 release-age exemption The four @earendil-works 0.87.1 packages were published 2026-09-22, well past the 24-hour minimumReleaseAge gate, so the one-time exception is no longer needed. The lockfile passes the supply-chain policy without it.
The pool loads the whole backend on the first call into the Durable Object, so the first test paid for that load against its timeout: about 10s alone, and over 30s when the full suite runs many files at once. Import the backend when the file loads instead, where no per-test timeout applies, and drop the 30s timeout override.
* Allow spawned agents to create and manipulate worktrees. Previously spawned agents were limited to just describeBinding and executeCode, on the theory that they shouldn't be modifying gadgets, but now that we have worktrees, things have changed. * Add `env.GIT` special binding that allows access to git storage. `env.GIT.newWorktree(commitId)` creates a new, transient `Worktree` instance, which can be used to inspect and create commits. This is particularly useful for gadget code that needs to manipulate commits (agents should use the `createWorktree` tool instead). * Extend `env.GIT` to allow reading metadata about a commit. This does not require checking out a worktree first. * Require full 40-digit commit IDs in `env.GIT` and `createWorktree` tool. Otherwise, it's too easy to guess commit prefixes in order to fish for commits that weren't intended to be exposed. We treat commit IDs as capabilities to read the content. * Add `Worktree.structuredDiff()` method. It's like `diff()` but returns the result as a data structure, useful e.g. to display a UI. * Add tests for grep tool on worktrees that haven't been modified. (Originally this commit actually fixed the bug, but the bug has since been fixed upstream, so the rebase now just adds tests.) The tool assumed that worktrees are always pinned, but that is no longer true as of the worktreees-ui changce. * Add profile setting for preferred git email address. This will be used when creating commits on behalf of the user. * Addess Bonk review comments: - Fully reserve `env.GIT` binding name in new envs (but don't break existing chats / gadgets that may have this binding name already). - Properly report mode changes in `structuredDiff()`. * Address more Bonk review comments: - `commit()` no longer tries to walk the tree when no changes have been made (and thus won't fail due to the tree not being faulted in). - Newly-created spawners can no longer include `GIT` in their seed bindings.
Previously, an agent could only describe its own direct bindings, not bindings on a gadget, but the latter might be needed in order to write code. Also in this change, describeBinding's output is now stored in the chat log and visible in the UI.
…t. (cloudflare#575) * UI bugfix: Only auto-switch to code view when there's no gadget UI yet. When an agent writes some code, we were always automatically switching the user from the gadget UI view to the code view. Even when the gadget had a working UI already. And even when the code the agent edited wasn't even the gadget's code. The auto-switching behavior exists solely for the first-prompt experience when starting a new gadget, so that the user can see that stuff is happening. We should never switch the user away from a working gadget UI to show them code. So, this switchin is now gated on whether or not the gadget has any UI yet. If it's still showing the "no UI" placeholder, then switching is OK. * Address AI review comment.
This is stuff I needed to configure Opus 5.5 when it was still early access. * Allow configuring extra headers on manually-configured AI models, needed e.g. to set `cf-access-token` for AI Gateway. * Allow configuring context window and output limit on manually-configured AI models. * Add ability to edit and clone existing model configs.
These files are delivered to the agent as-is in order to describe the API. The agent has no ability to follow imports. (Something we should fix eventually.)
Agents receive gatekeeper `types.d.ts` (and the other agent-facing declarations) as verbatim text, and nothing resolves their imports (cloudflare#581). Add `gadgets/self-contained-agent-types`, which flags import and export-from declarations, inline import types, `import = require`, and `/// <reference path|types>` directives unless the module is `cloudflare:workers`, which the agent's environment provides. Enable it for every gatekeeper `*types.d.ts`, mcp-shared's base types, and workshop-backend's `*-binding.d.ts`. Google's docs/drive types keep their sibling imports via `allow`, since `type-bundle.ts` strips exactly those and concatenates the imported declarations into the same bundle.
* Fix stale integration-testing docs - The fixture gatekeeper does run capnweb-validate (build:test-gatekeeper). - Drop the wrangler version from the workerd/wrangler coupling note. * Pin diff3 to 0.0.3 so updating from mainline works diff3@0.0.4 assigns the undeclared `ed` in onp.js, which throws in the Worker's strict ES-module bundle. updateChatFromMainline therefore failed with `ReferenceError: ed is not defined` whenever a file changed on both sides. 0.0.3 is the version isomorphic-git depends on and the one src/diff3.d.ts already describes; its diff3Merge() has the same signature and return shape. Fixes cloudflare#554. * Add public-API e2e tests for chat changes - scriptedModelRouter() routes scripted model requests by Workers AI account id, so concurrent agent tests can share one NetworkInterceptor and each consume only their own script. - workshop-changes.test.ts drives the real Workshop over Cap'n Web: accepting agent changes, stale merges and updating from mainline (clean and conflicting), reverting agent steps and releasing a pending gadget's name, two editors on one chat draft, clients behind on generation across merge/revert/discard, and draft vs mainline gadget code across merge and revert. * Drop chat-changes tests covered by the e2e suite Removes the concurrent-transform, agent-running rejection, stale-merge and destructive-bump cases, and trims the dedupe case to its sequence checks; workshop-changes.test.ts now exercises those through the public API. ChatChangeRecord no longer needs exporting. * Add public-API e2e tests for action approval and history - The test gatekeeper can arm a one-shot applyAction() failure per account (/control/fail-next-apply), consumed before it throws so a retry applies. - workshop-agent-actions.test.ts runs each case concurrently on its own scripted model: rejecting a held write leaves the agent stopped; approving two writes applies each and resumes the agent once, and re-approval is refused; a failed apply stays pending until the user retries or rejects; and an approval after a workspace restart (collaborator removal) applies and resumes exactly once. - workshop-action-history.test.ts pages every filter over a 56-record history, resolves an action between pending pages, streams live entries, and checks the reconnect replay against the listed change times, including that a record unchanged since before the watermark is not replayed. * Drop action-log tests covered by the e2e suite Removes the live-only, epoch/cutoff replay, newest-first, type-filter and resolution-between-pages cases, and the open-gadget-rpc listActions smoke; workshop-action-history.test.ts now exercises those through the public API. * Add public-API e2e tests for capability and security boundaries - workshop-use-role: a "use" collaborator is refused every Overseer and GadgetClient method outside the UI surface (exhaustive, type-checked tables), sees only mainline UI, and gets inert action history. - workshop-restricted-web-fetch: webFetch is refused once the workspace has read restricted data, and no request leaves the sandbox. - observer-exclusion: excludeObservers blocks only collaborators whose role's verification scope covers the connection, leaves no trace when it blocks, and stops blocking once the collaborator is removed. The fixture gatekeeper's readValue now forwards excludeObservers. * Drop use-role and exclusion tests covered by the e2e suite Each removed case's assertions are covered by workshop-use-role or observer-exclusion, confirmed by mutating overseer.ts and watching the e2e tests fail. Race, hook, spawner, lost-access, unknown-observer-id and ownerInvitesOnly revocation cases have no public equivalent yet and stay, as does the use-bound block for its no-teardown check. * Add public-API e2e test for turn recovery after a workspace restart
Pin Cloudflare-Studio/ask-bonk@45777ce (v0.5.0), and deny opencode paths outside the checkout instead of prompting, which nothing can answer in CI.
…nd (cloudflare#584) * Measure eval runs on the pull request's merge commit The eval workflow compared pull_request.base.sha with the raw head. A branch cut before a change on main then showed that change as its own: cloudflare#570, cut before cloudflare#525, showed the old eval tasks as added and the new ones as removed, and Bonk's diff showed the same. base.sha can also be older than the main GitHub merges into. Plan the merge commit against its first parent, and have the review job run from that parent and diff the same pair. * Count only eval code as an eval task's definition Any non-test file in integration-tests, or a workshop-evals manifest, changed every task's definition, which blanks the comparison: cloudflare#578 (README, mock model, fixture gatekeeper), cloudflare#568 (the Worker-inputs table) and cloudflare#569 (zod catalog specs). Only the eval package's TypeScript defines or scores a trial. A change to the rest still reruns the tasks. * Show which checks failed in the eval comment The comment showed only each task's pass rate, so a drop gave no hint where trials failed. It now lists, in a collapsed table, each check that failed on either side, counted over the trials that reached its turn. A turn the agent didn't finish shows as agent.<outcome>, since it has no checks to fail. Trials that failed for infrastructure reasons no longer count against the score, which shows them instead, for example "89% (1 run error)". Such rows were already left out of the verdict, so no verdict changes. A value both sides share is now shown once instead of as "X -> X", so a reused result reads as one set of numbers. * Stop waiting for the session drop in close() close() waited for the Workshop to drop the session after deleting the workspace, because an abort that found no client left crashed local workerd 1.20260831.1. cloudflare#567 moved to workerd 1.20260921.1, which fixed it (likely cloudflare/workerd#7370). With the wait removed, the crash probe from cloudflare#525 crashed workerd in 10 of 12 Gadget closes on cloudflare#567's parent and in none on cloudflare#567. On this commit it crashed in none of 48. The settleRestart doc gave the crash as its reason. Its callers are tests that check a change did not restart the workspace, so the doc now says that.
* Add kernel e2e harness hooks for hooks and external replies - Share withOwnerWorkspace, restartWorkspace, streamGeneration, waitForIdleChat and WorkpieceRecorder via rpc-client; listConnectedAccounts takes a ConnectedAccountsFilter. - Fixture TestThing.watch() binds a hook; /control/fire-hook and /control/hook-state drive it. - Fixture records delivered external-message replies (/control/gadget-responses). * Refuse new connections to an ambient gatekeeper an admin disabled * Add public-API e2e tests for connections, hooks and external messages - workshop-connections: accepting a connection request binds the gatekeeper the user picked and resumes the agent; denying leaves it stopped; removing a connection strands its pending action, drops its bindings, refuses its open sessions and blocks re-binding it. - workshop-hooks: an enabled hook fires into a restarted workspace and its write pends for approval, while disabled, deleted and administratively disabled gatekeepers' hooks refuse, even though the fixture never forgets its initiator. - external-message-round-trip: an external message gets exactly one reply, a reused idempotency key starts nothing, and deleting the chat mid-reply sends the terminal text. * Add public-API e2e tests for callable agents and worktrees - workshop-callable-agents: a call made while the callable agent's turn runs waits, and both calls land in order, each answered once, across a workspace restart. - workshop-worktrees: a worktree is invisible to other chats; its commits advance the head, a merge auto-commits the overlay into pinBase, a revert rolls the head back without touching pinBase, and deleting the chat removes the worktree. * Add public-API e2e tests for admin policy, sharing, deletion and blueprints - workshop-admin-policy: only admins get the AdminApi; disabled, enabled and optional gatekeeper modes and disabled resources gate listing, provisioning, disconnecting and new connections, while an existing session keeps working. - workshop-sharing: removing a sharer keeps or drops the people they shared with, and losing a build path falls back to the owner's direct use grant. - workshop-lifecycle: only the owner deletes a workspace, and later opens report enumerable access-denied and not-found codes. - workshop-blueprints: an installed blueprint binds the installer's account, not the publisher's, and reproduces the source gadget's files. * Add public-API e2e tests for provider errors and chat replay - A provider failure leaves the chat idle, a retry answers the prompt once, and a busy chat refuses new messages. - Resubscribing during a running turn or after a workspace restart replays exactly the messages the client's history lacks. * Drop restricted-action latch tests covered by the e2e suite sensitive-observations and auto-approval-policy already exercise these cases through the public API, confirmed by mutating overseer.ts and auto-approval.ts: latched actions pend and never auto-approve, approval still applies while latched, and an incomplete description pends rather than being refused. The push refusal and removed-connection cases stay.
* test(integration): connection requests wait for every decision and survive restarts * test(integration): stopping an agent keeps its work; external conversations stay separate * test(integration): approvals route per account; disabling auto-approval restores consent * test(integration): merged bindings and enabled hooks re-verify live viewers * test(integration): republished blueprints reach new installs; shared outputs follow access * test(integration): model edits keep secrets usable without returning them
* test(integration): per-step model usage and a fixture connect flow - A scripted model step can set its own token usage; unset steps keep the default. - The fixture gatekeeper's connectAccount() now runs a real flow: GET /connect/<label> completes the callback with the account of that label and serves gatekeeper-kit's handoff page. - Fixture accounts record their revocations, read back through /control/revocation-count. - The harness sets PUBLIC_BASE_URL, which only the handoff target origin reads. - gatekeeper-kit joins the Worker inputs, since the fixture now bundles it. * test(integration): chat deletion, attachments, compaction and rule draining - workshop-lifecycle: an attachment is readable only through its own chat, a discarded upload can't be sent, and a deleted chat's history is empty. - workshop-changes: deleting a chat drops its provisional gadget, its binding and name claims, and its attachment. - workshop-chat-recovery: a chat over its context budget compacts before the next turn, and history pages across the checkpoint with no gap or overlap. - auto-approval-policy: enabling a rule applies the eligible writes already pending and leaves manual ones pending. * test(integration): move gateway cost and user-directory coverage to the public API - ai-gateway-cost: a chat is charged the AI Gateway log's cost, read once more after a first 404, in place of the catalog estimate. - workshop-user-directory: search finds a user by display name, excludes the caller, follows a rename, and honours the admin switch without dropping the index; closing signups refuses new accounts while existing ones still log in. signUp takes an optional display name for it. Replaces workshop-backend's ai-gateway-cost test and user-directory-rpc's indexing case. user-directory-rpc keeps its cache expiry case, which needs the in-isolate clock: an open capability must re-read the search policy once its cache expires, whether it was disabled or re-enabled. Two old checks are dropped on purpose: that creating an account alone doesn't index it (every creation path authenticates at once, so it only pinned where indexing happens), and that a policy is still served from cache inside the TTL (a performance detail). * test(integration): move connect-handoff redemption coverage to the public API connect-handoff: a flow's account is added only when the initiating session redeems its ticket, once; another user's session, malformed input and another flow's nonce are refused, and a refused redemption spends both the ticket and the nonce and revokes the staged grant. Drops the backend cases these replace (unknown, malformed, redeemed or foreign tickets, another flow's nonce, a malformed nonce); the expired-nonce, at-rest, expiry, failure and reconnect cases stay.
…e#592) This avoids a race condition where if a message had just arrived between when the thread was inspected and when it was archived, the new message could be missed.
* Rename @gadgets/backend-utils to @gadgets/observability Everything the package holds is observability (logger, observability context, tracing, error reporting), and metrics are about to join it. No behavior change beyond the error-reporting log component, which follows the package name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add @gadgets/observability/metrics: activity/v1 Analytics Engine writer A typed codec for the gadgets_metrics_v1 dataset. Each event is written once to its actor's index, the canonical copy to count (an event with no user is attributed to sys_unattributed), and once more per owner or workspace it names, so per-entity queries read their own index. The column layout lives in the import-free metrics-schema.ts, so dashboards can resolve column names from it. A no-op without the optional METRICS binding. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Record Workshop product events as activity metrics recordAnalytics now also writes each event to @gadgets/observability/metrics through an exhaustive mapping. New events cover gadget workpieces (created when one becomes permanent, by direct creation, blueprint instantiation, or merging a chat's changes; proposed when a chat creates one provisionally), workspaces created by an external message, and accounts created through a sign-in gatekeeper, none of which were recorded before. gadget_opened moves from the server into Overseer.open(), which knows the workspace owner. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oudflare#598) * Keep the deleting session's WebSocket when a workspace is deleted deleteSelf() restarts the workspace DO about 100ms later (scheduleAccessRestart). The abort drops every session's notifyClosed stub uncalled, which AuthenticatedApiImpl reads as a lost DO and answers by closing that session's whole WebSocket. The deleter kept its socket only if its own dispose made it to the DO and back inside those 100ms, so under load, or for a user far from the DO, deleting a workspace dropped every other call and subscription on the connection. The deleter has nothing to reopen, so deleteSelf() now reports its session closed once the deletion has committed, well ahead of the abort. The abort still severs every other session holding the workspace, and the deleter's own stub still dies with the DO; only the forced reconnect is skipped. This was the integration-test flake "Peer closed WebSocket: 3000 RPC session was shut down by disposing the main stub" in workshop-blueprints.test.ts. * test(integration): the deleter's session survives its workspace's restart workshop-lifecycle now holds the deleted workspace's stub through the restart and reads the user's list over the same session, in place of reconnecting and logging in again. Before the kernel fix this failed every time with "Peer closed WebSocket: 3000". workshop-blueprints drops the workarounds that disposed each workspace right after deleting it to try to beat the abort: the race they narrowed is gone. * Close the deleting session before scheduling the restart deleteSelf() fired notifyClosed() without waiting, after the restart was already scheduled, so correctness rested on the call landing inside the restart's ~100ms delay. An abort in the same turn as the call loses it and still closes the deleter's WebSocket. Await it instead, before scheduling the restart. The rejection is swallowed: inside blockConcurrencyWhile it would reset the DO mid-delete, and a gone client needs no close anyway.
* Remove redundant observation logs in gmail gatekeeper. The gmail gatekeeper was logging observations for *internal* reads that don't actually represent the agent/gadget caller observing anything. This is redundant and unnecessary. * Also remove redundant observations from GitHub gatekeeper. Similar to previous commit. * Update write-gatekeeper skill to elaborate on when observations need to be authorized.
* Trace agent turns for the Agents dashboard Cloudflare's Agents dashboard shows each agent's sessions, turns, model calls, tool calls and token use. It reads them from Workers trace spans with fixed names and gen_ai.* attributes. The Agents SDK emits those spans for Think and the AI SDK. Our agent runs on pi, so we emit them ourselves: - invoke_agent for each turn, in runAgent. It also carries the turn's log fields (operation, gadgetId, chatId, modelId), so you can filter traces by the same names as logs. - chat for each model request, in the model handle. Its input token count includes cache reads and writes, which pi counts apart from input. - execute_tool for each tool call, including calls pi rejects before running them. The span leaves out a tool name the model made up. - tool_approval when an agent's action waits for a manual decision, and when a user or a rule approves it or a user denies it. No span records prompts, responses, tool arguments or results. A turn the user stops is marked canceled, not failed. A request that fails without an HTTP error status, such as a refusal, is marked _OTHER. Model calls outside a turn, such as a gadget's model binding, also get chat spans, without agent identity. Logs written inside runAgent now carry agentName. invoke_agent replaces the agent.run span, so the tracing helper in @gadgets/observability goes too. httpStatusFromError now takes one request's response metadata, so a chat span reads the status of its own request. A span belongs to the Durable Object invocation that started the turn. The turn keeps all its spans while that request stays open, as a browser session does. A turn that an external message, a callback or a restart started can lose spans when the next request arrives. * Cover turn setup in the invoke_agent span The invoke_agent span opened in runAgent, after the overseer's turn setup: gadget cleanup, the usage check and model selection. So a turn the usage limit blocked, or one that failed before it reached a model, had no span, and setup time was not part of the turn. The overseer now opens the span around the turn's setup and run. It adds the model once it chooses one, and marks a usage-limit block as error.type usage_limit. The span still ends before the turn's teardown, because the teardown can start the chat's next turn, which would otherwise nest inside this one. The turn body in #runAgentTurnWithContext moves into the traced callback, so most of its diff is indentation. * Match the stubbed model provider by origin in tests CodeQL flagged the fetch stub's url.startsWith() check as incomplete URL sanitization. The check only routes test requests, but an origin comparison is the exact match the stub needs.
Bumps [pako](https://github.com/nodeca/pako) and [@types/pako](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/pako). These dependencies needed to be updated together. Updates `pako` from 2.2.0 to 3.0.2 - [Changelog](https://github.com/nodeca/pako/blob/master/CHANGELOG.md) - [Commits](nodeca/pako@2.2.0...3.0.2) Updates `@types/pako` from 2.0.4 to 3.0.0 - [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases) - [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/pako) --- updated-dependencies: - dependency-name: pako dependency-version: 3.0.2 dependency-type: direct:production update-type: version-update:semver-major - dependency-name: "@types/pako" dependency-version: 3.0.0 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…lare#633) Code edits are made on a chat's branch, so the Code view locks editing when no conversation is selected -- e.g. a workspace freshly created from a blueprint, or one reopened from the sidebar. Nothing said so: the header just read "Viewing" and the New file button was silently disabled, which users read as the editor breaking until they sent a new prompt. Show "Select or start a conversation to edit" as a banner above the editor and as the disabled New file button's tooltip.
vite-plus 1.0.0 is the only VoidZero package that can move within the
24h minimumReleaseAge window without breaking the toolchain:
- vite stays 7.3.6: Vite 8's Oxc leaves Stage-3 decorators unlowered,
so every workerd suite importing a `@validateRpc()` class fails with
"SyntaxError: Invalid or unexpected token" (re-verified on 8.3.1; see
the catalog comment).
- vitest stays ^4.1.11: @cloudflare/vitest-pool-workers 0.22.0 peers
vitest ^4.1.0.
- @vitejs/plugin-react stays ^5.2.0: 6.x requires vite ^8.
Migration for 1.0:
- Override `vite@*` instead of bare `vite`, so vite-plus keeps its
`vite` -> @voidzero-dev/vite-plus-core alias instead of having vite 7
substituted under `vp`.
- Move task `env`/`input`/`output` under `cache` (vite-task#749), in the
package configs and the shared task builders.
- Import the oxlint plugin types and RuleTester from
`vite-plus/lint/plugins{,-dev}` and drop the separately pinned
@oxlint/plugins dependency, its drift test and dependabot ignore.
- oxlint 1.85: use toSorted() on three freshly built arrays, and turn
off react/globals for workshop-frontend tests, whose hook probes
assign to an outer `let` on purpose.
…udflare#635) Bumps the build-toolchain group with 2 updates in the / directory: [esbuild](https://github.com/evanw/esbuild) and [terser](https://github.com/terser/terser). Updates `esbuild` from 0.28.1 to 0.28.2 - [Release notes](https://github.com/evanw/esbuild/releases) - [Changelog](https://github.com/evanw/esbuild/blob/main/CHANGELOG.md) - [Commits](evanw/esbuild@v0.28.1...v0.28.2) Updates `terser` from 5.49.2 to 5.51.2 - [Changelog](https://github.com/terser/terser/blob/master/CHANGELOG.md) - [Commits](terser/terser@v5.49.2...v5.51.2) --- updated-dependencies: - dependency-name: esbuild dependency-version: 0.28.2 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: build-toolchain - dependency-name: terser dependency-version: 5.51.2 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: build-toolchain ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…loudflare#639) This is the first, small step to break up overseer.ts, and probably the most obvious. All of the typed-storage collections and TS types are now defined under a `storage-schema` subdirectory. I generally think of APIs and storage schemas as defining the architecture of a piece of software. Everything in between is just glue, and can change more easily. Consolidating storage schemas into one place makes it easy to tell what storage schemas are changing in any given PR. I also moved migrations into this new directory. Claude suggested several small related cleanups as well which I told it to go ahead and do.
Since cloudflare#611, addModel refuses an id a gateway model would shadow, and the eval target added one in gateway mode, where the gateway already serves the eval model. So every eval trial failed at setup.
* Plan for improving gmail simulation. * Make gmail gatekeeper simulate label changes. This is part 1 of plans/gmail-simulation.md. * Make gmail gatekeeper simulate sends. This is part 2 of plans/gmail-simulation.md.
The cache hit rate also moves with how long a run is and how much new content it reads, so caching fixes were lost in its noise. comparison.json now carries a cache break rate (tokens the previous step sent that a step sent again instead of reading from cache) and per-run Mann-Whitney p-values for cache hits and breaks. The table bolds a significant cache hit change, and Bonk counts cost and cache changes in its verdict.
pi sends the leading system message as one block, so a change to the project-specific text missed the cache from the start of the prompt. The handle now splits it after the static text, with a cache breakpoint there: on Anthropic by moving the system block's breakpoint, and on OpenAI GPT-5.6+ with an explicit prompt_cache_breakpoint.
The Workers API stopped accepting `preview_defaults`, the field the pinned draft Wrangler writes for `wrangler preview secret bulk`, so every preview deploy failed at the first worker with secrets. The secrets now go through `wrangler preview base-config secret bulk` on the workspace's own Wrangler, which writes `previews_base_config`. Everything else stays on the draft build: neither sibling preview bindings nor per-preview KV and R2 provisioning is in a released Wrangler, checked against 4.138.0 and 4.147.0.
Cover user journeys a kernel refactor could silently break: - the agent's system prompt follows the workspace's gadgets, bindings, ambient connection and standard formats, and never another chat's pending binding - a pasted link (capsule) becomes a binding for that chat only - a chat attachment reaches the model and stays in later turns - one accept commits every gadget a chat built, each with its own code - gadget console logs reach subscribers labelled draft or mainline, stop on dispose, and never reach a use collaborator - a gadget's LLM binding runs on its bound model; a blueprint install uses the installer's model - admin instance instructions and format hints reach the agent, not the user - disconnecting an account from the Connectors page removes and revokes it Adds systemPromptOf() to the mock model for reading recorded prompts. Also makes bare `.rejects.toThrow()` assertions on RPC promises able to fail. vitest's `.rejects` calls a callable subject, and a Cap'n Web RpcPromise is callable: calling it pipelines a call on the result, which rejects with "'' is not a function." when the original call succeeded, so those assertions passed whatever the call did. They now match the refusal message; two were also wrong, since writeValue resolves once the write is submitted for approval, and now assert that instead.
* Restart a workspace whose loop counter is exhausted
The Workers runtime refuses a Durable Object call once the loop
counter behind it is spent ("Subrequest depth limit exceeded. This
request looped back into the Workers runtime too many times."). A
workspace object's outgoing channels can end up holding a spent
counter with nothing recursing, and from then on every call it makes
to a user object is refused until the instance is replaced.
When one of the workspace's own user-object calls is rejected that way
(a call through a wrapped stub, the last-active bump, or the outputs
sync), the workspace now schedules the existing access restart. At
most one restart per instance, and none in an instance's first 60
seconds. Errors thrown by gadget, agent or gatekeeper-facet code are
never consulted.
…loudflare#616) A deployment could only change which models its AI Gateway offers by patching SUGGESTED_MODELS. That patch goes stale whenever the catalog changes, and some models can't go in the public catalog at all. In AI Gateway mode, /admin gets a new Models tab: Enable / test providers and set default reasoning for the deployment and more!
…he account binding (cloudflare#649) * Start Google Chat direct messages and group chats from the account binding The whole-account Chat session gains two methods: - searchPeople(query) searches the connected account's Workspace directory (domain profiles only, never contacts), recording each page as an observation. - sendDirectMessage(people, text) sends to one person (a DM) or 2-49 people (a group chat with exactly them). An existing conversation is an ordinary send. Otherwise the message is queued as a new "chatStartConversation" action kind, separately auto-approvable from sends, and every person must be named by email and resolve to a directory profile: outsiders and yourself are refused before anything is queued. Applying a start creates the conversation idempotently (spaces.setup with a requestId, remembered once created), reuses a group chat that appeared since queueing, and verifies membership before posting, since Google silently drops anyone who blocks the caller from a new group chat. Until committed, the message carries a temporary pending:space:{id}; edits queued against it carry over into the real conversation, replies are refused. Undo deletes the message. The account resource adds chat.spaces.create (not chat.spaces) and directory.readonly, so existing account connections re-consent. * Fix Google Chat sends whose new conversation exists before they post A send that starts a direct message or group chat is queued under a temporary pending:space name. If applying it set up the conversation but the post then failed, two things went wrong: - ChatSpace.listMessages() for the new conversation left the queued send out, so an agent could conclude it was never sent and send it again. listForSpace and resolveMessage now resolve the temporary name to the conversation, which also lets the send's capability pass that conversation's scope check. The send reports the real spaceId once its conversation exists, and the agent-facing docs say so. - A failed member check after a successful setup, such as a 503, left the send marked as possibly sent, so it could not be rejected until a retry succeeded. The setup's mark is now cleared as soon as the conversation exists: an empty conversation shows nobody anything. A retry after the conversation is recorded still skips setup, so a send whose post may have landed stays unrejectable. A new test pins that. * Export gatekeeper-kit's SingleFlight and use it for Chat conversation setup SingleFlight coalesces concurrent work by key and releases each flight once it settles. It was internal to the kit; it is now the ./single-flight subpath, listed in the README inventory and noted in the design record. The Google Chat gatekeeper used a hand-rolled map of promises to make sends to the same people, applied at once, share one conversation setup. It now uses SingleFlight, whose own tests cover joining and release. * Re-check a new Google Chat conversation's members before retrying an unsent post A send that starts a group chat records the conversation once it is set up. If the post was then refused outright, a retry reused that conversation without checking who was in it, so someone who joined in between would receive the approved message. A retry that definitely didn't post now refuses when the members have changed; one whose post may have landed still goes back to the same conversation, so it stays idempotent. The directory lookup that confirms each person in a new conversation read one page of matches and called anyone not on it outside the organization. When Google reports further pages it now says it couldn't confirm them instead. * Check that Google Chat sends reach exactly the approved people A new conversation's send checked only that nobody approved was missing, and several paths could still reach someone who wasn't: - A replayed spaces.setup returns the group as it is now, so after a lost setup response someone added since would receive the message. The send is now refused when the conversation holds anyone extra, and stays rejectable. - A retry after a post that may have landed skipped the member check. Every retry that reuses the recorded conversation now checks it first, leaving the attempt mark as it is, so an uncertain send stays unrejectable. - findGroupChat trusted a matching member count, but Google's fallback for a block can offer a group with someone else in it. Each requested person is now confirmed, by user ID against the member list or by email through members.get, and a membership Google reports as NOT_A_MEMBER doesn't count. - peopleIn read only the first page of members. It now follows every page, since Google may return fewer members than asked for. * Simplify Google Chat conversation starts A send that creates its conversation is queued under a temporary pending:space: name. That name was resolved at five call sites; the store now resolves it once, in get() and list(), so every reader sees the real conversation once it exists. The guard that keeps such a send from being replied to keyed on the temporary spaceId, which missed the window after setup but before the post; it now checks the stored send. openConversation returns as soon as Google's lookup finds the group chat, since the lookup already confirmed exactly who is in it, rather than listing the members a second time. Its post-setup check is one early return and one throw instead of a nested ternary. The directory search reads its response through readGoogleJson, as the other People API call does, so it is size-bounded and logs Google's reasons; the #request/#fetchJson split that existed only for it is gone. The 400/404 "no such user" check shared by three lookups is one predicate, and #onlyIn no longer reads a membership for a users/{id} reference that is already absent. The tests drop unused backend state and share the group-chat setup and Chat-app membership they repeated.
…flare#656) * Stream git packs through consumePack instead of buffering them Mounting a large repo as a worktree exceeded the Overseer's 128 MB memory limit. consumePack collected the whole pack, inflated every object, and deflated each one again before storing anything, so a TypeScript- or vscode-size mount held 173-224 MiB of buffers at once. decodePackStream replaces decodePackBytes. It decodes the pack in one pass, keeping only each entry's oid by offset, and resolves every delta base through resolveBase. It reads the stream with a BYOB reader in 64 KiB chunks, enforces the pack size cap as bytes arrive, and hashes exactly the bytes before the trailer. A failed decode cancels the source. Every existing malformed-pack check is kept. MAX_DELTA_DEPTH is gone because nothing recurses any more. consumePackFromGatekeeper stores blobs of up to 1 MiB as they arrive. It holds commits, trees, tags and oversized blobs (a later delta may name one as its base) and stores them in one transaction once the trailer verifies, so a failed pack never leaves a commit that fetchCommit would treat as mounted. #storeVerifiedObject now handles the oversized case for both put() and consumePack. A delta must now follow its base. That holds for every pack git upload-pack sends a fetch with no haves, which is the only kind the gatekeepers request. Real depth-1 mount packs (react, next.js, TypeScript, vscode, and a 52 MiB internal repo) went through the real code on Durable Object SQLite in workerd. Every object was stored, or measured if oversized. A pack with a corrupted trailer left none of its commits or trees. * Use native node:zlib for loose git objects Mounting a vscode fork on a preview exceeded the Overseer's 30 s per-invocation CPU limit. In local workerd, deflating every stored object with pako in encodeLooseObject was the largest single cost of consumePackFromGatekeeper (about 4 of 10 s for vscode). encodeLooseObject and decodeLooseObject now use workerd's native node:zlib (deflateSync/inflateSync, available under nodejs_compat). The output is standard zlib at the same default level, so stored records stay readable by isomorphic-git and by records written before. decodeLooseObject views inflateSync's Buffer as a plain Uint8Array, so the payloads it hands out stay plain arrays. The pack decoder keeps pako, which it needs to find where each entry's zlib stream ends. consumePackFromGatekeeper in local workerd, three runs each, real Overseer DO storage, workerd process CPU: - vscode (23,289 objects): 10.3 s -> 7.4 s - TypeScript (65,040 objects): 18.6 s -> 13.3 s Every object was stored or measured, oids recompute, isomorphic-git reads a sample byte-for-byte, and stored bytes and SQLite size are unchanged. Whether a vscode mount now fits the production limit is not yet verified. * Store a pack's held objects one per transaction A vscode mount still reset the Overseer on the preview, now on CPU: the invocation running consumePack used 32.5 s against the 30 s default. Run on Cloudflare against the real vscode mount pack, in a throwaway Worker's SQLite Durable Object, consumePackFromGatekeeper used 23.7-98 s of CPU, and 3 of 7 runs were reset on a storage timeout. Most of the excess was the one transaction around the held commits and trees: each metadata put inside it opens a nested transaction, and ~29k of those in one transaction cost 14-18 s there. Locally the whole transaction takes under 1 s. Held objects are now stored one per transaction, like the small blobs, with commits last. No await separates the stores, so they still reach disk together. A store that throws rolls back only its own object, and with commits last no commit is stored before the pack's trees. A new test covers a tree the cache refuses (mode 100664, which git only reports as informational) arriving after its commit, as in the packs git sends. Same production round, interleaved: vscode 13.1-14.5 s (was 23.7-27.4 s), every object stored. TypeScript-size packs (65k objects) still take about 38 s (was 48.6 s), over the default limit. * Reject a pack entry whose size varint overflows The decoder checks each entry's declared size against maxObjectSize before inflating it, but a size varint with enough zero continuation digits overflows its multiplier to Infinity, and 0 * Infinity leaves a NaN size that the comparison lets through. On a 1,024-byte cap such an entry inflated 8 MiB before failing. The varint loop now rejects the entry once the next digit would be worth more than the cap, since no size within the cap needs it. * Deflate loose git objects at level 1 Deflating each object in encodeLooseObject is the largest single CPU cost of storing a pack, and CPU is what still stops large mounts. On this repo's 1,378 tracked files of up to 64 KiB, level 1 deflated 1.53x faster than the default level 6 and stored 11% more bytes. In local workerd a vscode first pull went from 7.3 s to 6.2 s with 12% more stored bytes, and a TypeScript one stored 9% more. Any zlib level is a valid loose object: oids hash the inflated bytes, isomorphic-git and decodeLooseObject read every level, and rows already stored at level 6 stay readable. Push packs are re-compressed by buildPackBytes, so they are unaffected. The extra bytes are kept for as long as the workspace is, since nothing repacks the store. * Leave objects a gatekeeper already stored untouched A pull sends no haves, so a retried pull, or one for another commit of an already-mounted repository, delivers mostly objects the store already holds. #storeVerifiedObject rewrote each of them: a fresh deflate, an object put and a metadata put, for no change. An object that is already present, measured and proven on this gatekeeper's remote is now left as it is. Nothing it would write can differ: the payload is fixed by the oid, its referents were recorded and its pending-push marks propagated when it was first stored, and marking walks descend into present objects themselves. Another gatekeeper's copy is still recorded as that gatekeeper's proof. In local workerd, delivering the same vscode pack again went from 6.7-7.9 s to 2.8-3.2 s. * Test consumePack on a pack streamed over Workers RPC The consumePack tests hand the decoder a byte stream. A gatekeeper sends a default stream from another Worker, which the decoder's BYOB reader refuses when handed one directly ("This ReadableStream does not support BYOB reads"); it works only because Workers RPC delivers it as a byte stream. A new test sends a real pack that way, from a Worker Loader worker in 100-byte chunks, and checks every object is stored with its bytes intact. * Correct the GitCache.consumePack contract doc It said consumePack is exactly equivalent to decoding the pack and put()ting each object. Since packs are streamed that no longer holds: a delta must follow its base, and a pack that fails can leave some of its objects stored, though never a commit without its trees. The doc also now says what the result holds and that an object too large to store is left out of it rather than thrown on. * Report why a git fetch failed, not how its stream ended When the fetch response fails -- a server ERR line, the transfer limit, a truncated response -- consumePack sees the pack stream it reads across RPC only end early, and rejects with "disconnected prematurely". That is what gitPull reported, so an agent mounting a repo too large to fetch was told the connection dropped, and retried. demuxGitFetchResponse takes an optional onFailure callback, and pullGitObjectsIntoCache rethrows the failure it reports in place of consumePack's rejection. A consumePack failure of its own still surfaces when the fetch itself succeeded.
…loudflare#657) * Let an admin state an added model's image input and reasoning levels An added AI Gateway model the runtime has no entry for gets provider defaults, and "Behaves like" borrows another model's flags wholesale. An added model can now carry two optional stated facts: whether it takes images, and which reasoning levels it accepts. Per field the order is the runtime's own entry, then what is stated, then what the model it behaves like lends, then the provider defaults.
Node stores a symlink target with backslashes on Windows, and the importer reports ignored paths with the platform separator, so three assertions that spell those paths with forward slashes fail there. Compare against the platform form instead; on POSIX the expectations are unchanged. Co-authored-by: snowyukitty <270071858+snowyukitty@users.noreply.github.com>
…fter 8 hours (cloudflare#661) * fix(gatekeeper-github): refresh expiring user tokens instead of dying after 8 hours GitHub enables expiring user tokens by default for OAuth apps registered since 2026-08-14. The gatekeeper kept only the access token, so every connection on such an app failed with Bad credentials eight hours in. UserAccount now stores its grant through the kit's CredentialCoordinator, with the code exchange and refresh going through the kit's OAuthClient: an expiring grant is refreshed shortly before it expires, concurrent reads share one refresh, both rotated tokens are stored together, and GitHub's bad_refresh_token marks the grant dead and notifies the Workshop once. Non-expiring grants, including ones stored in the old layout, are served as before. A 401 is now adjudicated against the token the request presented: a token a refresh or reconnect replaced mid-request fails as retryable instead of marking the account expired, and a rejection of the current token refreshes past it where the grant can. Fixes cloudflare#641 * fix(gatekeeper-github): fence revoke before awaiting GitHub, re-arm legacy expiry latch Review follow-ups: - revoke() now moves the credential fence before its first await, so a refresh landing mid-disconnect is discarded and its tokens revoked instead of being stored and then deleted unrevoked. - Migrating a pre-refresh grant re-arms the expiry latch: that layout latched before delivering, so a failed delivery left a dead account showing as connected. - prepareReconnect no longer resets the latch by hand; every credential replacement re-arms it. * test(gatekeeper-github): race the in-flight refresh from the test, not the fake's handler * fix(gatekeeper-github): replay configurator and observer reads past an in-flight token refresh A refresh invalidates the token a request already carries, so a configurator lookup or hasRepoAccess check overlapping one failed with GitHub's 401 although the replacement token works. Both now run through withAccountApi, which adjudicates the rejected token with the account and, for these replay-safe reads, reruns the call once with the replacement. A rejection the account confirms as grant death now also notifies the Workshop from these paths. * fix(gatekeeper-github): replay a review apply's follow-up reads past a token refresh applyAction(postReview) creates the review, then reads its comments back before recording the action applied. A refresh that replaced the token one of those reads carried failed the apply after GitHub had accepted the review, leaving the action pending, so approving it again posted a second review. Those reads (and the review-comment sync they can fall back to) now go through #readApi, which reruns a read once with the replacement token. Mutations stay non-replayable. * test(gatekeeper-github): dispose the review helper's param stubs Under the 2026-09-04 compatibility date (rpc_params_dup_stubs semantics) a caller keeps ownership of stubs it passes as params, so it must dispose them.
…cloudflare#659) A blueprint now encodes a git packfile instead of a Yjs snapshot. When you update a blueprint, people who have created gadgets from the old blueprint are now given the option to update it. For now, this does not happen automatically -- you just see a dot on the blueprints button and from there can choose to update. (We can consider implementing auto-update later perhaps.) If local changes have been made since the blueprint was first instantiated, an agent is spawned to review the merge and fix any conflicts. You can also choose to switch blueprints when updating. This allows you to e.g. switch from an original blueprint to some other user's customized version of it. This change also changes the way merges are handled, both for the purpose of merging updates from blueprints, but also for merging a chat branch with changes that have landed in a gadget's mainline from other chats: merges are no longer encoded as OTs but rather as a git commit, which allows handling much larger diffs.
Providers cache a prompt by its prefix, so a request that doesn't start with the whole previous request pays again for everything after the first difference. The test runs two turns, one with a tool step, and checks that each request keeps every field but its messages unchanged and starts its messages with those of the request before.
…ads (cloudflare#668) Supabase, Linear, Spotify, ZoomInfo and Home Assistant set expiredNotified before awaiting credentialsExpired(), so one failed callback (a Workshop restart mid-deploy is enough) silenced the reconnect prompt for good, and in four of them the RPC error replaced the auth error the caller saw. They now notify through the kit's notifyCredentialsExpiredOnce, which latches only after delivery, and re-arm with clearCredentialExpiryLatch. Home Assistant and the MCP gatekeeper parsed their connect form with formData() before the account checked the nonce, so anyone holding a connect URL could make the Worker buffer a body as large as the plan allows into a 128 MB isolate. Both now read it through readTextCapped (16 KiB), which accepts a Request as well as a Response. gatekeeper-kit: - connect-handshake adds isConnectAttempted/markConnectAttempted on the connectAttempted key the internal no-nonce connect flows already store, and USAGE states the at-most-once complete() rule they enforce. - connectMutationError's JSDoc notes that a form page served with Referrer-Policy: no-referrer, as htmlResponse pages are, submits with Origin: null (checked in Chromium), so a guarded form needs same-origin. Also clarifies the Cloudflare gatekeeper's omitted-scope comment (no behavior change), and updates the kit plan: status, the complete-once rule, the Origin: null constraint on §4.3 and tokenAuth, package-name fixes, and the candidate leaves reviewed and not built. Records the kit's admission rule in its AGENTS.md and plan §1: a leaf abstracts behavior several gatekeepers share under one contract, and an unusual gatekeeper gets plain TypeScript seams, not a provider-specific variant. Plan entries that leaned on a single consumer now follow it: ironclad's split-key handshake gets no kit variant, OAuth discovery and client registration wait for a second consumer outside the MCP SDK, and scope helpers wait for a second gatekeeper on Google's grant model.
GPT-5.6 and later keep a separate prompt cache for each prompt_cache_key, and pi sets one per chat, so every new chat wrote the tools and static system prompt again. Drop pi's key on these models, so chats share one cache, as they already do on Anthropic. Older models route by the key, so they keep it. The project-specific part of the system prompt now starts with a random salt stored per workspace, so nobody without that prompt can probe the shared cache for a project's prompt or chats. Requests without it, such as compaction, keep pi's per-chat key.
* Retry a model request that fails before it streams anything A transient provider failure, such as a stream that ends early or a 5xx, ended the turn with an error. runAgent now runs the pass again, up to 2 times, when pi calls the failure retryable and the request streamed nothing, so the client has no partial output to withdraw. * feat: retry transient model failures even after partial output A transient failure is now retried whether or not the failed request already streamed text, reasoning or tool calls, so background agents with no one to press Retry recover from mid-stream drops too. The failed request still persists nothing; the retry starts from the last saved step. Before retrying, the agent emits a new "streamReset" stream event. The chat UI handles it like an error message: it drops the streamed text and reasoning, tool-call cards, the active-file marker and all edit previews, so the retry's output isn't appended to the failed attempt's. * Type the retry test and reset the stream only before a retry The retry test now reaches OverseerImpl through a typed view, as agent-tracing.test.ts does, instead of any. runAgent emits streamReset when it actually retries, so a failure that ran out of retries ends with the error message alone, as the event's docs say. --------- Co-authored-by: Cloudflare OS PRs <kenton@cloudflare.com>
On Anthropic, keep the head of an agent's request (tools and static system prompt) in the cache for 1 hour instead of 5 minutes. Every chat of the agent's kind shares it, so a chat that resumes after a pause of more than 5 minutes reads it instead of writing it again. The chat's own messages keep 5 minutes, because a 1-hour write costs 2x instead of 1.25x on every new token. One-shot calls (compaction summaries, chat and gadget titles, binding names) now send cacheRetention: "none". Their prompts are sent once, so a cache write only added its premium: 1.25x on Anthropic and on GPT-5.6 and later, where "none" selects explicit caching with no breakpoints. A gadget's model binding keeps caching, since a gadget may send the same system prompt many times, and so does the admin's gateway model test, so that a model that rejects the cache fields fails the test rather than the first agent turn.
…are#684) Cloudflare's Web Search API is reached through the Workers AI binding's websearch() method, which shipped in workerd 1.20260924.1. Our generated worker-configuration.d.ts files predate it (they were last generated with workerd 1.20260831.1), and local runs used workerd 1.20260921.1, so the method was neither typed nor present in local dev. This bumps wrangler to 4.147.0 and miniflare to 5.20261001.0-alpha, which bring workerd 1.20261001.1, raises @cloudflare/workers-types to wrangler 4.147's peer range (^5.20261001.1), and regenerates every worker-configuration.d.ts with pnpm types:generate.
* feat(overseer): mint gatekeeper facet stubs through ctx.restore() A gatekeeper facet reached through a bare ctx.facets.get() stub cannot call its own ctx.restore(), so it cannot mint a persistent stub to itself for a push handler to deliver events through. Route every gatekeeper facet stub through the Overseer's ctx.restore(), as gadget facets already are; [restore]() checks only that the connection still exists, so such stubs stop restoring once it is removed. addGatekeeper keeps the raw facet for its pre-publication describe(). * docs(agent): fix the persistent-stub example's import; clarify restore params and `self` The system prompt's restore example imported `Greeter` from cloudflare:workers, which exports no such name and collides with the class the example declares, and omitted the `RpcTarget` the class extends, so copying it failed to load. Also say that a Gadget may hold many persistent stubs, each restored from its own params, which is how one Gadget tells several hooks apart. Restore params may hold persistent stubs, and `self` is one, so code that receives it stores it as it is. Agents applied the subscription advice to `.dup()` a received callback to `self`, which calls a reserved method and yields an unstorable RpcPromise, and read "must be serializable" as ruling `self` out of params. "e.g. to a subscription method" is gone: it invited subscribing `self` as a hook whose deliveries carry live capabilities it cannot store.
Upstream main at 6eb1112, 78 commits since the merge base bfe217f. Every PointFive carry is kept. Conflicts and their resolutions: - packages/workshop-backend/package.json: upstream bumped @types/pako, dompurify and esbuild beside the lines where the gadget client bundler adds esbuild-wasm and pdfjs-dist. Upstream's three versions plus the bundler's two dependencies. - packages/workshop-backend/src/agent.ts: upstream rewrote the system prompt's "two main files" sentence (cloudflare#600) on the line where the bundler adds "(plus any other files they import)". Upstream's sentence with that clause kept: "A Gadget is defined by two main files, client.js and server.js (plus any other files they import). Create them with writeFile ...". - packages/workshop-backend/src/overseer.ts: upstream moved UserAiModelRecord and WorkspaceOutputEntry to ./storage-schema/user-storage (cloudflare#639) on the import line beside the bundler's ./gadget-bundle import. The bundler's import plus upstream's narrowed ./user import. - pnpm-lock.yaml: regenerated by pnpm install from the conflicted file, never hand-merged. It keeps both sides' locked direct dependencies (pdfjs-dist stays 6.3.289); pnpm re-resolved a few transitive build dependencies under capnweb-validate (unplugin 3.4.0) and dropped the registry's deprecation notes. pnpm install --frozen-lockfile passes on it. Each resolved file differs from upstream's side only by the PointFive lines, and from p5's side only by upstream's own changes. Auto-merged without conflict: ai-gateway.ts, ai-models.ts, env.d.ts, wrangler.jsonc, GadgetUI.tsx. wrangler.jsonc is generated from cloudflare.config.ts upstream (cloudflare#597); the esbuild.wasm Data rule survives here as text and moves into cloudflare.config.ts in the next commit.
…nfig.ts wrangler.jsonc is generated from cloudflare.config.ts since cloudflare#597, and the generator emits `rules` only from the config's `wrangler` export. The rule that loads esbuild.wasm as bytes for the gadget client bundler lived only in the hand-edited wrangler.jsonc, so `configs:check` failed on it and `configs:generate` (which `pnpm dev-server` runs) dropped it. Without it esbuild.wasm loads as a compiled WebAssembly.Module, which the Worker Loader in production refuses ("Unable to deserialize cloned data"). ModuleRule now models Data rules beside Text; the rule and its reasoning move into workshop-backend's cloudflare.config.ts, and the generated wrangler.jsonc carries it.
… only exports A bare `export` imports nothing, so the build could only hand back its own bytes. Sending it to the bundler anyway made a Worker Loader necessary to serve such a UI, which broke upstream's integration tests that boot without one (the use-role and sharing UI-bundle reads). mightImport now asks for an `import`, or an `export ... from`; a new test uses a loader that throws if called. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge cloudflare/cloudflare-os main into p5
Target: upstream
mainat6eb111202d926e759946a3cbbe21e4638e4a737c(2026-10-06, #677), refetched last.78 commits since the merge base
bfe217f5, none of them merges.Upstream #642 (gadget UIs import other modules) is not in the target: still an open draft, none of its 8 commits
reachable from it. Upstream #659 (blueprints as git,
dcedc16a) is.The fork's
mainis synced to the target before this PR.git log --no-merges upstream/main..origin/p5lists the five carries plus3d81a4af, which retired2726b96.Conflicts
Trial
git merge-tree --write-tree origin/p5 upstream/main: four files, as expected.workshop-backend/package.jsonesbuild-wasm ^0.28.2,pdfjs-dist ^6.3.289@types/pako ^3.0.0,dompurify ^3.4.16,esbuild ^0.28.2workshop-backend/src/agent.ts(system prompt)workshop-backend/src/overseer.ts(imports)./gadget-bundleimport beside the old./userimportUserAiModelRecord,WorkspaceOutputEntrytostorage-schema/user-storage./gadget-bundleline plus upstream's./userlinepnpm-lock.yamlpnpm installafterpackage.jsonAuto-merged:
ai-gateway.ts,ai-models.ts,env.d.ts,wrangler.jsonc,GadgetUI.tsx.Per carry
1177d69key aliasai-gateway.ts4 (#611, #639, #616, #657);ai-models.ts8;env.d.ts1 (#568, adjacent import rename)0cdc5f7bundleragent.ts12,overseer.ts16,package.json10,wrangler.jsonc1 (#597), lockfile 21loadGadgetWorker's (code version plus chat sequence) and still matches it after #659. Hazard: after #597wrangler.jsoncis generated fromcloudflare.config.ts. The esbuild.wasm Data rule survives as text, but the generator emitsrulesonly from the config'swranglerexport, sopnpm configs:check(part ofpnpm lint, the fork CI) fails, andconfigs:generate(run bypnpm dev-server) drops the rule.workshop-backend/cloudflare.config.tswrangler.rules; widenModuleRuleinscripts/worker-config.ts(it modelsTextonly) to allowData; regenerate. Retires with #642, not here.2815b4cdownloadsGadgetUI.tsx1 (#575, a no-UI callback)allow-downloads.8686b38attachmentsprepareChatAttachment.ts02726b96sync workflow3d81a4af.Trust boundary
access.tsuntouched;CF_ACCESS_ISS/CF_ACCESS_AUDhandling unchanged;server.tschanges are model and blueprint APIsVITE_CF_ACCESS_MODEstill inlined fromuseAuth.ts; vite-plus 1.0 moves it undercache.env#isAdminandADMINSuntouched; the list stays deployment-controlledaiGateway.providers, which becomes a floor; default reasoning and compaction; switch off users' own models; test calls billed to the gateway. Defaults keep today's behaviour (users' own models on, models.dev suggestions off).fields,descriptionIsComplete(#541, #565); persistent stubs throughctx.restore(#677: every gatekeeper facet stub now routes through the Overseer's restore);consumePackreturns oids (#656). Restricted mode (#487): after a restricted observation, actions pend for manual approval instead of being refused. Only Google's BigQuery marks restricted data (off here); PointFive gatekeepers mark none.@gadgets/backend-utilsrenamed@gadgets/observability(#568); newaction-description,single-flightexports; capped connect-form reads (#668)cloudflare.config.ts(#597); DO migrations unchanged (v0 to v3); compatibility date unchanged; optionalMETRICSAnalytics Engine binding, a no-op unbound (#568)gatekeepers[])GoogleChatGatekeeperImpl.deploy-inputs.json: the same two secrets, no new input; setup steps add the Chat and People APIs and a configured Chat app. Three new resource types (Chat Account, Chat Conversation, Chat Thread), on by default (/admin stores a disable list). Six new scopes:chat.spaces.readonly,chat.messages,chat.memberships.readonlyfor all three; the account type addschat.users.readstate.readonly,chat.spaces.create,directory.readonly. Chat sends, edits, reactions and new conversations can be set to Always approve (Gmail sends cannot). The account type cannot be shared.Fix forward
Rollback target: the live gitlink
4d8b6ce4. Once deployed, these do not undo with it:<id>/<commitId>;.gadgetarchive version 2No other storage version or migration found: the range's only new
version.putis Overseer's v5, and DO migrationschange only in Google.
After the deploy, check
The key alias (a Workshop chat on Claude), the bundler (a multi-file gadget, one importing
@gadget/pdf), a gadgetdownload, a large PNG attachment, /admin → Models listing the deployment's providers, each PointFive gatekeeper
(brain, snowflake, r2, d1) through a gadget, and Google with its Chat resource types off.
Not confirmed
defaultkey when the alias has no key for a provider: Cloudflare's BYOKpage does not say. It matters only for a provider an admin turns on in /admin.
configs:checkfailing on the moved rule: by reading the generator, not by running it.On this branch
The merge was rehearsed on a local copy of
p5, and its four conflict resolutions were replayed on the realp5(ffa0563a) with git rerere: identical tree. One commit on top moves the esbuild.wasm Data rule intoworkshop-backend/cloudflare.config.ts(wideningModuleRuletoData), soconfigs:checkpasses andconfigs:generatekeeps the rule. Locally on the merged tree: the four carries' fork tests pass (key alias 8/8, bundler 19/19, downloads and attachments 4/4),pnpm lint:check,configs:check, and the workshop-backend and workshop-frontend builds.🤖 Generated with Claude Code