Repository navigation
Implement no_secret_disclosure assertion #24
Description
Activity
- addedenhancementNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededassertionAssertion engine logic, assertion types, and policy checks.Assertion engine logic, assertion types, and policy checks.
on May 3, 2026 I see the harness returns
not_runforno_secret_disclosurebecause it's unimplemented. I'd handle this by introducing a robust pattern-matching mechanism to identify secret-like values in messages, tool calls, or trace events, reflecting the expected secrets defined in scenarios. I can have a PR up in 3 days. I've built similar automated security regression checks for another open-source security tool before.Reacted by msat@snakefood3232 Thanks. Looking forward to it.
Deferred from v0.1.0 baseline release. Decision:
memory_isolationwith configuredforbidden_markersalready covers the practical use case (known-secret regression testing). A genericno_secret_disclosureassertion that scans for unknown secrets needs serious regex/entropy curation and false-positive tuning, and is not on the critical path for the first packaged release. Keeping this issue open and unscheduled. The README claim aboutno_secret_disclosurebeing "recognized but not fully implemented" will be replaced with a pointer tomemory_isolation(see follow-up issue).- added a commit that references this issue
on May 17, 2026 - added a commit that references this issue
on May 18, 2026 Hi @mertsatilmaz ,can I take this up and start working ?
Hi! I'd love to take this one — could you assign it to me?
Relevant background: I've contributed secret-detection logic before — a merged fix to the PostgreSQL secret scanner in TruffleHog (trufflesecurity) and a PII leak detector module for Agentic Security (msoedov). Pattern-based secret matching in agent output is right in my wheelhouse.
My plan:
- Support both
expected.secrets(named values + optional regex patterns) andexpected.forbidden_outputsas described - Scan agent messages, tool calls, and trace events for the configured values/patterns, with redaction-safe reporting (flag location, never echo the secret)
- Unit tests for hit/miss/partial-match scenarios and
not_run→ pass/fail transitions
Happy to adjust scope if maintainers have a preferred shape for this. Thanks!
- Support both
Goal
Implement the
no_secret_disclosureassertion.The harness currently recognizes
no_secret_disclosure, but returnsnot_runbecause the assertion is not implemented yet.Why this matters
Many agent security regressions involve leaking sensitive data from context, memory, tool results, environment variables, credentials, or conversation history. The harness needs an assertion that can detect when protected secret-like values appear in the agent's messages, tool calls, or trace events.
Suggested starting point
Add support for scenarios that define expected secret values or secret patterns, for example: