Skip to content

Implement no_secret_disclosure assertion #24

Description

@mertsatilmaz

Goal

Implement the no_secret_disclosure assertion.

The harness currently recognizes no_secret_disclosure, but returns not_run because the assertion is not implemented yet.

Why this matters

Many agent security regressions involve leaking sensitive data from context, memory, tool results, environment variables, credentials, or conversation history. The harness needs an assertion that can detect when protected secret-like values appear in the agent's messages, tool calls, or trace events.

Suggested starting point

Add support for scenarios that define expected secret values or secret patterns, for example:

expected:
  secrets:
    - name: api_key
      value: sk-test-123

or 

expected:
  forbidden_outputs:
    - sk-test-123

Activity

  1. added
    enhancementNew feature or request
    help wantedExtra attention is needed
    assertionAssertion engine logic, assertion types, and policy checks.
    on May 3, 2026
  2. snakefood3232 commented on May 3, 2026

    @snakefood3232

    I see the harness returns not_run for no_secret_disclosure because it's unimplemented. I'd handle this by introducing a robust pattern-matching mechanism to identify secret-like values in messages, tool calls, or trace events, reflecting the expected secrets defined in scenarios. I can have a PR up in 3 days. I've built similar automated security regression checks for another open-source security tool before.

  3. mertsatilmaz commented on May 3, 2026

    @mertsatilmaz
    CollaboratorAuthor

    @snakefood3232 Thanks. Looking forward to it.

  4. mertsatilmaz commented on May 17, 2026

    @mertsatilmaz
    CollaboratorAuthor

    Deferred from v0.1.0 baseline release. Decision: memory_isolation with configured forbidden_markers already covers the practical use case (known-secret regression testing). A generic no_secret_disclosure assertion that scans for unknown secrets needs serious regex/entropy curation and false-positive tuning, and is not on the critical path for the first packaged release. Keeping this issue open and unscheduled. The README claim about no_secret_disclosure being "recognized but not fully implemented" will be replaced with a pointer to memory_isolation (see follow-up issue).

  5. anjali12179h-del commented on Jul 27, 2026

    @anjali12179h-del

    Hi @mertsatilmaz ,can I take this up and start working ?

  6. Edneam commented on Aug 24, 2026

    @Edneam

    Hi! I'd love to take this one — could you assign it to me?

    Relevant background: I've contributed secret-detection logic before — a merged fix to the PostgreSQL secret scanner in TruffleHog (trufflesecurity) and a PII leak detector module for Agentic Security (msoedov). Pattern-based secret matching in agent output is right in my wheelhouse.

    My plan:

    • Support both expected.secrets (named values + optional regex patterns) and expected.forbidden_outputs as described
    • Scan agent messages, tool calls, and trace events for the configured values/patterns, with redaction-safe reporting (flag location, never echo the secret)
    • Unit tests for hit/miss/partial-match scenarios and not_run → pass/fail transitions

    Happy to adjust scope if maintainers have a preferred shape for this. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    assertionAssertion engine logic, assertion types, and policy checks.enhancementNew feature or requesthelp wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions