Skip to content

Share OpenAI's prompt cache across chats - #615

Open
AshishKumar4 wants to merge 1 commit into
mainfrom
prompt-cache-deployment-key
Open

AshishKumar4 wants to merge 1 commit into
mainfrom
prompt-cache-deployment-key

Conversation

@AshishKumar4

@AshishKumar4 AshishKumar4 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

On GPT-5.6 and later, OpenAI keeps a separate prompt cache for each prompt_cache_key, and pi sets one per chat. So every new chat wrote the tools and static system prompt again, about 7k tokens at 1.25x. In #609's eval run, the first request of all 40 trials read 0 cached tokens.

This PR drops pi's key on those models, so chats share one cache across the OpenAI organization, as they already do across a workspace on Anthropic. Older models route by the key, so they keep it.

A shared cache lets anyone on the same organization tell from cached token counts whether a prompt prefix was sent recently. So the project-specific part of the system prompt now starts with a random salt, created on a workspace's first agent turn and stored with it. Only agent turns, which carry the salt, drop the key; other requests on a chat, such as compaction, keep pi's per-chat key. Only the static part can still be probed: the tools, the base prompt and the deployment's admin instructions.

In this PR's last eval run, first requests that read the cache went from 0/10 to 10/10 on chess and incident-desk, and to 9/10 on worker-logs.

@github-actions github-actions Bot added the kernel Changes to the Workshop kernel label Sep 30, 2026
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

Copy link
Copy Markdown

Preview: pr615-prompt-cache-a4635ee0

https://pr615-prompt-cache-a4635ee0-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 30, 2026
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 8028d18 to a8769fc Compare October 1, 2026 19:09
@AshishKumar4 AshishKumar4 changed the title Share OpenAI's prompt cache across a deployment's chats Share OpenAI's prompt cache across chats Oct 1, 2026
@ask-bonk

ask-bonk Bot commented Oct 1, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 1, 2026

@kentonv kentonv left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems reasonable.

Also, if we are actually concerned about cache-probing attacks, it strikes me that there's a better solution: Place an unguessable random string at the beginning of the project-specific part of the system prompt, selected on a per-project basis.

I'm not personally all that worried about it, but this seems so easy that we might as well do it?

@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from a8769fc to 8b5f03e Compare October 2, 2026 14:49
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 8b5f03e to 9abee69 Compare October 2, 2026 16:19
Comment thread packages/workshop-backend/src/ai-models.ts Outdated
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

Review: 1 finding.

Posted one actionable inline comment.

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 9abee69 to 43bd84d Compare October 2, 2026 16:41
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@AshishKumar4
AshishKumar4 added this pull request to stack #644 October 2, 2026 16:45
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 43bd84d to 2a664c2 Compare October 2, 2026 16:45
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

Base automatically changed from prompt-cache-system-blocks to main October 2, 2026 16:56
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 2a664c2 to 9597812 Compare October 2, 2026 16:56
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Comment thread packages/workshop-backend/src/agent.ts Outdated
let systemMessage: SystemMessage = {
role: "system", content: systemPromptSlots[0], sections: {environment: systemPromptSlots[1]},
role: "system", content: systemPromptSlots[0],
sections: {environment: `Workspace ID: ${hooks.getWorkspaceId()}\n\n${systemPromptSlots[1]}`},

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using the workspace ID is not an ideal choice here because:

  • We have to think about whether we want the agent to know the workspace ID.
  • It's possible that people who do not have access to the workspace know the ID.

Both are minor points, but we could also derive an entirely random value here, store it, and use that.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I chose workspace ID primarily to avoid adding much code and complexity; we basically need a random value that is stored and persisted in the workspace and accessible across all chats, and workspaceID seemed to fit the bill well.
Having a separate value would add to the schema, and a new argument to every function in the chain inbetween

@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 9597812 to 744a0cc Compare October 5, 2026 15:55
@ask-bonk

ask-bonk Bot commented Oct 5, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 100% → 80% −20 pp p = 0.47 86% → 88%
+2 pp
3.4 → 4.5 25.0 → 24.6
chess 90% → 88% (2 run errors) candidate run errors — 97% 9.2 → 8.8 61.9 → 53.3
incident-desk 100% → 100% (1 run error) candidate run errors — 95% 4.3 → 4.7 38.3 → 32.3
worker-logs 100% → 89% (1 run error) candidate run errors — 92% → 93% 4.0 24.2 → 20.4
Failed checks
Task Check Failed
change-calendar t1 agent.timedOut 0/10 → 1/10
change-calendar t4 names-the-booked-windows-the-new-rules-reject 0/10 → 1/9
chess t1 agrees-with-the-oracle-on-the-hard-positions 1/10 → 1/8
chess t1 agrees-with-the-oracle-on-perft-positions 1/10 → 1/8
chess t1 agrees-with-the-oracle-through-random-games 0/10 → 1/8
worker-logs t1 ingests-resets-and-summarises-per-worker 0/10 → 1/9
worker-logs t1 hourly-buckets-cover-every-hour-including-empty-ones 0/10 → 1/9
worker-logs t1 ranges-are-half-open-and-filter-by-worker 0/10 → 1/9

Run · trajectories and raw results

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 5, 2026
@ask-bonk

ask-bonk Bot commented Oct 5, 2026

Copy link
Copy Markdown

🔬 Eval runs review

Performance

The only comparable task, change-calendar, fell from 10/10 to 8/10 passes, within noise (p=0.47); candidate run errors prevent comparing the other three tasks. Its cache hits rose from 86.4% to 88.0% and breaks fell from 7.8% to 7.4%, neither significant; completed runs cost $0.0235–$0.0342 versus main’s $0.0217–$0.0299, while mean duration increased from 3.4 to 4.5 minutes. Provider authorization errors lost the most runs, and chess’s exact-text editing detours wasted the most steps, including 30 unmatched edits versus main’s eight.

⚪ VERDICT: NO REGRESSION FROM THIS PR

Removing the per-chat cache key demonstrably enables first-request cache reads, but comparison.json establishes no improvement beyond noise, and the failed trajectories do not implicate that rewrite or the workspace salt.

Triage

Failure modes

  • Provider authorization rejection · chess 2/10, incident-desk 1/10 · harness bug · this PR: no — chess trials 2 and 8 stopped in turns 3 and 2 after successful edits; incident-desk trial 3 passed turn 1 but received a 401 instead of answering turn 2. These are infrastructure errors, not failed gadget implementations; the diff does not change credentials.
  • Activation never finishes · change-calendar 1/10 · harness bug · this PR: no — trial 2, turn 1 ended after three parallel writeFile calls, with no final reply or tool error, until the 420-second activation timeout. The underlying stall is not exposed, and activation handling is unchanged.
  • Incorrect calendar diagnosis · change-calendar 1/9 · model error · this PR: no — trial 9, turn 4’s executeCode fetched every booking, but the final reply incorrectly marked mw-101 and mw-103 as TOO_CLOSE; their separation exceeds 24 hours. Main’s agents calculated the rejected set rather than relying on this mistaken date arithmetic.
  • Kings move like sliding pieces · chess 1/8 · model error · this PR: no — trial 10, turn 1’s writeFile(server.js) put kings through the queen-style sliding loop before adding castling, producing illegal and duplicate moves. The agent inspected the file but never exercised the rules with executeCode; main also had one rules-engine failure, although its defect was lost castling metadata.
  • Oversized SQL inserts · worker-logs 1/9 · system prompt · this PR: no — trial 3, turn 1’s writeFile(server.js) batched 100 events into 600 SQL bindings, incorrectly claiming this stayed below SQLite’s limit. Verification rejected ingestion, leaving later summaries empty; main’s successful implementation inserted one row per statement. The unchanged storage guidance omits the environment’s bound-parameter limit.
  • Verification connection loss · worker-logs 1/10 · harness bug · this PR: no — trial 6 completed turn 2’s edits and reply, but verification encountered WebSocket connection failed; three other turn-2 checks passed. comparison.json classifies this as an infrastructure error, not evidence of lost stored data.

Tool errors

  • editFile: No matching text was found · chess 8 → 30, incident-desk 3 → 4, worker-logs 1 → 1 · model error — agents quoted stale or incorrectly escaped source. Chess trial 6, turn 2 submitted twelve overescaped replacements together, all rejected, before correcting the escaping.
  • editFile: Validation failed · change-calendar 5 → 4, chess 1 → 1, incident-desk 5 → 2, worker-logs 5 → 2 · model error — calls omitted required fields or supplied wrapper objects instead of tool arguments. Change-calendar trial 4, turn 1 used .replace instead of replacement, then recovered with the correct field.
  • editFile: Multiple matches were found · chess 10 → 4, incident-desk 1 → 3, worker-logs 0 → 1 · model error — replacements lacked unique surrounding context. Chess trial 5, turn 1 retried the same ambiguous fragment even after grep showed both matches, then succeeded with longer snippets.
  • readFile: File does not exist · incident-desk 2 → 0, worker-logs 2 → 0 · model error — main’s agents tried reading server.js and client.js immediately after creating empty gadgets, despite the prompt explaining that new gadgets have no files.
  • executeCode: Failed to start Worker · chess 0 → 3 · model error — candidate calls contained an unescaped apostrophe in a PGN string or only : or }. Corrected JavaScript succeeded; these were malformed agent calls, not Workshop startup failures.

What to do

  • Rerun after stabilizing provider authorization and the Workshop connection so chess, incident-desk, and worker-logs become comparable; no corrective PR code change is supported by this run.
  • Optional, unrelated follow-up in packages/workshop-backend/src/agent.ts: document the SQL bound-parameter limit beside the storage guidance and recommend per-row inserts inside transactionSync.
  • Optional, unrelated follow-up in packages/workshop-backend/src/agent.ts: add explicit editFile recovery guidance to re-read unmatched text, preserve literal escaping, and include surrounding context for duplicate matches.

github run

GPT-5.6 and later keep a separate prompt cache for each prompt_cache_key, and pi sets one per chat, so every new chat wrote the tools and static system prompt again. Drop pi's key on these models, so chats share one cache, as they already do on Anthropic. Older models route by the key, so they keep it.

The project-specific part of the system prompt now starts with a random salt stored per workspace, so nobody without that prompt can probe the shared cache for a project's prompt or chats. Requests without it, such as compaction, keep pi's per-chat key.
@AshishKumar4
AshishKumar4 force-pushed the prompt-cache-deployment-key branch from 744a0cc to 82a3c50 Compare October 5, 2026 19:45
@kentonv

kentonv commented Oct 5, 2026

Copy link
Copy Markdown
Member

LGTM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kernel Changes to the Workshop kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants