Share OpenAI's prompt cache across chats - #615
AshishKumar4 wants to merge 1 commit into
Conversation
|
LGTM! |
Preview:
|
0857ace to
8028d18
Compare
|
LGTM! |
8028d18 to
a8769fc
Compare
|
LGTM! |
kentonv
left a comment
There was a problem hiding this comment.
Seems reasonable.
Also, if we are actually concerned about cache-probing attacks, it strikes me that there's a better solution: Place an unguessable random string at the beginning of the project-specific part of the system prompt, selected on a per-project basis.
I'm not personally all that worried about it, but this seems so easy that we might as well do it?
a8769fc to
8b5f03e
Compare
|
LGTM! |
8b5f03e to
9abee69
Compare
|
Review: 1 finding. Posted one actionable inline comment. |
9abee69 to
43bd84d
Compare
|
LGTM! |
43bd84d to
2a664c2
Compare
|
LGTM! |
2a664c2 to
9597812
Compare
|
LGTM! |
| let systemMessage: SystemMessage = { | ||
| role: "system", content: systemPromptSlots[0], sections: {environment: systemPromptSlots[1]}, | ||
| role: "system", content: systemPromptSlots[0], | ||
| sections: {environment: `Workspace ID: ${hooks.getWorkspaceId()}\n\n${systemPromptSlots[1]}`}, |
There was a problem hiding this comment.
Using the workspace ID is not an ideal choice here because:
- We have to think about whether we want the agent to know the workspace ID.
- It's possible that people who do not have access to the workspace know the ID.
Both are minor points, but we could also derive an entirely random value here, store it, and use that.
There was a problem hiding this comment.
I chose workspace ID primarily to avoid adding much code and complexity; we basically need a random value that is stored and persisted in the workspace and accessible across all chats, and workspaceID seemed to fit the bill well.
Having a separate value would add to the schema, and a new argument to every function in the chain inbetween
9597812 to
744a0cc
Compare
|
LGTM! |
Eval resultsVerdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.
Failed checks
|
🔬 Eval runs reviewPerformanceThe only comparable task, change-calendar, fell from 10/10 to 8/10 passes, within noise (p=0.47); candidate run errors prevent comparing the other three tasks. Its cache hits rose from 86.4% to 88.0% and breaks fell from 7.8% to 7.4%, neither significant; completed runs cost $0.0235–$0.0342 versus main’s $0.0217–$0.0299, while mean duration increased from 3.4 to 4.5 minutes. Provider authorization errors lost the most runs, and chess’s exact-text editing detours wasted the most steps, including 30 unmatched edits versus main’s eight. ⚪ VERDICT: NO REGRESSION FROM THIS PRRemoving the per-chat cache key demonstrably enables first-request cache reads, but comparison.json establishes no improvement beyond noise, and the failed trajectories do not implicate that rewrite or the workspace salt. TriageFailure modes
Tool errors
What to do
|
GPT-5.6 and later keep a separate prompt cache for each prompt_cache_key, and pi sets one per chat, so every new chat wrote the tools and static system prompt again. Drop pi's key on these models, so chats share one cache, as they already do on Anthropic. Older models route by the key, so they keep it. The project-specific part of the system prompt now starts with a random salt stored per workspace, so nobody without that prompt can probe the shared cache for a project's prompt or chats. Requests without it, such as compaction, keep pi's per-chat key.
744a0cc to
82a3c50
Compare
|
LGTM |
On GPT-5.6 and later, OpenAI keeps a separate prompt cache for each
prompt_cache_key, and pi sets one per chat. So every new chat wrote the tools and static system prompt again, about 7k tokens at 1.25x. In #609's eval run, the first request of all 40 trials read 0 cached tokens.This PR drops pi's key on those models, so chats share one cache across the OpenAI organization, as they already do across a workspace on Anthropic. Older models route by the key, so they keep it.
A shared cache lets anyone on the same organization tell from cached token counts whether a prompt prefix was sent recently. So the project-specific part of the system prompt now starts with a random salt, created on a workspace's first agent turn and stored with it. Only agent turns, which carry the salt, drop the key; other requests on a chat, such as compaction, keep pi's per-chat key. Only the static part can still be probed: the tools, the base prompt and the deployment's admin instructions.
In this PR's last eval run, first requests that read the cache went from 0/10 to 10/10 on chess and incident-desk, and to 9/10 on worker-logs.