Repository navigation
Test that each agent request extends the one before it - #663
Merged
Merged
Conversation
Providers cache a prompt by its prefix, so a request that doesn't start with the whole previous request pays again for everything after the first difference. The test runs two turns, one with a tool step, and checks that each request keeps every field but its messages unchanged and starts its messages with those of the request before.
|
LGTM! |
Preview:
|
ndisidore
approved these changes
Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Anthropic and OpenAI cache a prompt by its prefix. A request reads the cache up to the first point where it differs from an earlier request, and pays again for everything after that. So each agent request in a chat must start with all of the request before it.
This PR adds an integration test for that rule. It runs two turns against the scripted model, one of them with a tool call. For each request, the test checks that:
messagesare the same as in the previous request, andmessagescome first.The test compares serialized JSON, so a key in a different order also fails it.
Check: I added
Date.now()to the system prompt, and the test failed on the first request of turn 2.Known gap: a turn that changes the workspace still breaks the prefix, because the system prompt lists the gadgets and their files. #665 fixes that and adds that case to this test.