The ledger comment, two test docstrings and the trace test catch up with room_agent - #547
Conversation
…om_agent The comment on _OPEN said twenty-seven test fakes stood in for call_agent across thirteen files — a census from before #546, when tests patched call_agent at scattered sites. Every test's fake now goes through one seam, room_agent().turn.call_agent, so the number named nothing a reader could still verify. The comment now points at that seam instead of a count. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both files said their scenario is invisible at run_turn.call_agent — the name of the seam before #546 split RoomAgent out of run_turn. The fake moved to room_agent().turn.call_agent along with everything else on the turn, and the old name in the docstring no longer points at anything a reader could grep for. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…t an absence test_a_passing_verdict_leaves_no_refusal_trace only checked that no refusal record showed up on LOGGER_NAME, so it stayed green under a wrong logger name exactly as easily as under the right one: no refusal is expected either way, refusal or not. Falsified by pointing LOGGER_NAME at a name nothing logs to — 7 of the file's other 8 tests went red, this one did not. _timed() logs an INFO line on this same logger for every turn, including the passing one, so the test now raises the capture level and asserts that record showed up naming its own session. Falsified again the same way: all 8 go red now. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
| #: nothing in exactly the runs a regression would show up in. Each request owns its own copy, | ||
| #: so two sessions answering at once cannot add to each other's total. | ||
| #: not an argument threaded through `call_agent`: every test's fake for that function goes | ||
| #: through one seam, `room_agent().turn.call_agent`, and a parameter it never touches is a |
There was a problem hiding this comment.
room_agent().turn.call_agent is neither the only seam tests fake nor the only one this ledger counts. the_room_agent_is (tests/turn_harness.py:88) installs at four — turn, analyst, classifier, judge — and 30 test files pass one of the other three, e.g. the_room_agent_is(monkeypatch, classifier=agent) at tests/test_internalization_room_coverage_settle.py:94. Those calls are in flight under a ledger too: counted_for wraps the classifier at background.py:75 and the analyst read at background.py:151. Nothing executes here, but a reader who takes this at face value sizes the seam at one. Pick one, I'm fine with any: name room_agent() as the door and say four agents sit behind it, scope the sentence to the turn path, or drop the seam count and keep the argument — a fake never touches a threaded parameter whichever seam it sits at.
There was a problem hiding this comment.
Read the whole diff — the ledger comment, the two docstrings and the trace test. One thing to raise, and it is on the comment: the seam count it puts in place of the stale one is not what the tree has either. The trace test holds up: _refusal_records filters on WARNING, so dropping the capture to INFO does not loosen the absence it already checked, and _timed (validated_turn.py:208) does write the session id on that logger. I could not run the suite from here — I read it rather than ran it. That is the whole list: one finding, nothing else.
…th through it little-henok caught that the rewrite named room_agent().turn.call_agent as though it were the only seam a test fake sits at and the only one this ledger counts, when the_room_agent_is installs at four — turn, analyst, classifier and judge — and every one of them runs under an open ledger. The comment now names room_agent() itself as the door and the four agents that sit behind it, with no site to miscount. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Second pass. Read the whole diff at 18e67a5, and git diff 0aed8afc..HEAD, which is the _OPEN comment and nothing else. The one item the first pass left is closed at head — checked against tests/turn_harness.py and room_agent.py. Nothing new: the lines added since that SHA carry no bug the fix brought with it. Not covering the suite — I did not run it, so the falsified trace test in the description is unverified from here. That was the whole ledger, one item, and there is nothing else.
Three texts left behind by ENG-918. The comment on
usage._OPENcounted test fakes from before the provider existed; two test docstrings still namedrun_turn.call_agentas the seam; and the trace test's "a passing verdict leaves no refusal trace" only asserted an absence, so it stayed green with the engine logging under the wrong name while the seven tests beside it went red.The comment says what holds now, the docstrings name
room_agent, and the passing-verdict test asserts on the record the engine does write, under the engine's own logger, before asserting what it does not.Verificação: 3 commits, um por item; o teste de rastro falsificado com o nome de logger errado (agora vermelho, antes verde); ruff, mypy e a suíte do recorte tocado verdes; revisora (brejo-lupa) sem achados.
🤖 Generated with Claude Code