Skip to content

feat(gates): catch at analysis time what three model passes caught in review - #6

Merged
MyAlterLego merged 2 commits into
mainfrom
feat/analysis-time-evidence-gates
Jul 25, 2026
Merged

MyAlterLego merged 2 commits into
mainfrom
feat/analysis-time-evidence-gates

Conversation

@MyAlterLego

Copy link
Copy Markdown
Contributor

A full-surface run of a ~3,600-file TypeScript monorepo was authored by one model, adversarially reviewed by a second, then audited by a third. The review corrected 48% of 96 findings (22 refuted, 14 downgraded). The audit then found two whole unmodelled control actions, a systemic authority defect, and a seeded super-admin account.

Every one of those errors was findable at analysis time. This PR makes them findable mechanically, instead of depending on the next reviewer being strong enough — because the review kept doing work the analysis should have done.


The evidence that motivated each gate

What went wrong Why it survived Now caught by
4 findings claimed a control was absent; all 4 were false Read the use of a flag, never the declaration or the mount stpa controls + Prioritize absence check
A belief was credited as a working control when 2 of 3 code paths read the role from the caller's token sourcedFrom named a hop, not an origin stpa evidence trust roots
Two generic RPC bridges missed by two passes and a review A :param route collapsed into one REST control action dynamic-dispatch-route modality
Prose said "44 tombstones"; the grid held 66 A number typed by hand, invalidated by a later correction stpa evidence derived numbers
0 of 96 verdicts were upgrades Reviewers briefed only to refute stpa verify upgrade duty
A reviewer's correction was reverted by the apply step Nothing checked that verdicts reached the artifact stpa verify exit 7

New: stpa controls (exit 8)

ControlStructureScan.ts already emits a guards category. In that run it produced 2,508 guard candidates and the analysis consumed zero of them.

For two of the four false absence claims, the scan had already flagged a guard in the very file the finding called unguarded — including WWW-Authenticate inside the auth middleware the finding said did not exist. The evidence was on disk the whole time. Nothing looked at it.

This does the one comparison nobody was doing. Nothing to fill in by hand — a gate that costs busywork gets bypassed with --warn-only. On a collision it prints the guard lines and gives three exits: the guard is the control (withdraw the finding), the guard doesn't cover this path (cite why), or it's noise (cite and dismiss). All three require having read the line.

It cannot prove an absence. It proves you did not look, which is the actual failure mode.

New: stpa evidence (exit 9)

Trust roots. Every process-model belief must name where its evidence bottoms out: database · iam · network · human · attacker-input · unverified. A run credited "roles are re-read from the database on every request"; true for loginRequired, false for identityHasRole()/hasClaim(), which decrypt the role out of the caller's own token — so six remediations saying "gate this on an operator role" were placebos. "the session" and "the token" are hops; expand until it terminates. unverified is a legitimate answer that forces the dependent band provisional.

On the live analysis this immediately exposed that 20 of 38 beliefs are unverified — a quality signal the old model had no way to express.

Derived numbers. No count the tools compute may be typed into prose. --fix-numbers rewrites stale ones. Excludes tool-generated files and third-party audit documents (rewriting someone else's audit to match artifacts you changed afterwards would falsify their record).

Extended

  • DiscoveryGate — dynamic-dispatch-route [FORGOTTEN]. Fires with 4 hits on the exact file that hid the RPC bridges.
  • VerifyGate — upgrade/inventoryGap verdicts, a warning at zero upgrades, and an apply-reflection check: refutations must be tombstoned, downgrades and provisionals stamped.
  • Prioritize — absence claims must cite a line read or a sweep run.
  • RenderReport — §7c Corrections from corrections.json.
  • Start.md — Q1 is model selection as an author+attacker pairing, so the same model cannot land in both slots and fail stpa verify after the analysis is paid for. STRIDE removed from the intake and FullAnalysis — a same-model per-element pass is a circular cross-check; a real dual-method assessment belongs in its own skill.

Verification

Each gate tested positively and negatively — passes on the corrected artifacts, fails when its evidence is stripped:

controls   collisions 0 -> exit 0   | citation stripped -> exit 8
evidence   38/38 trust roots        | one removed        -> exit 9
verify     74/74 reviewed           | stamp removed      -> exit 7
prioritize absence evidence present | evidence removed   -> exit 1
discovery  dynamic-dispatch: 4 hits in sessions/routes/index.ts

Two false-positive classes were found by running the gates and fixed: "443 open" matching a port range, and "closes 12 findings" cluster-leverage counts read as totals. Both mattered — a gate that cries wolf is a gate that gets disabled.

Two bugs in my own gates, found the same way. The absence check first demanded file:line for an absence, which is incoherent; it now accepts a named negative search. And VerifyGate's apply check read 06-remediation.json (Prioritize's output, findings under waves[].items[]) instead of remediation.json, so every verdict looked unapplied.

The apply-reflection gate also caught a live instance of the bug it was written for, mid-implementation: re-merging reviews after the apply step had run left 11 verdicts filed but unstamped.

Notes

… review

A full-surface run of a ~3,600-file monorepo was authored by one model, adversarially
reviewed by a second (22 of 96 findings refuted, 14 downgraded — 48% corrected), then
audited by a third, which found two whole unmodelled control actions, a systemic
authority defect and a seeded super-admin account.

Every one of those errors was findable at analysis time. This makes them findable
mechanically instead of hoping the next reviewer is strong enough.

NEW GATES

  stpa controls   ControlInventory.ts — exit 8
    ControlStructureScan already emits a `guards` category. That run produced 2,508
    guard candidates and the analysis consumed ZERO. The author then wrote four
    findings asserting a control was absent, and all four were false — an image
    signature verifier, a PTY idle timeout, an auth middleware, and a config default
    that inverted the meaning of "not passed". For two of them the scan had already
    flagged a guard in the very file the finding called unguarded, including
    `WWW-Authenticate` INSIDE the middleware declared not to exist.
    This cross-checks every absence claim against those candidates. Nothing to fill in
    by hand — a gate that costs busywork gets bypassed. It cannot prove an absence; it
    proves you did not look, which is the actual failure mode.

  stpa evidence   EvidenceGate.ts — exit 9
    (a) Every process-model belief must name a trustRoot: database | iam | network |
        human | attacker-input | unverified. A run CREDITED "roles are re-read from the
        database on every request" as a working control; true for loginRequired, false
        for identityHasRole()/hasClaim(), which decrypt the role out of the caller's own
        token. Six remediations saying "gate this on an operator role" were placebos.
        "the session"/"the token" are hops, not roots.
    (b) No count the tools compute may be typed into prose. A scope document said "44
        tombstones" while the grid held 66. One checkable-and-wrong number discredits
        every number beside it. --fix-numbers rewrites them.

EXTENDED

  DiscoveryGate  +dynamic-dispatch-route modality [FORGOTTEN]. A route whose path
    segment NAMES the operation is a registry, not an endpoint, and collapses wrongly
    into one REST control action. This hid /api/orch/:command and /api/function/:command
    — generic unfiltered RPC bridges reaching .eval, arbitrary SQL and a self-update
    taking a caller-supplied image — through two passes AND an adversarial review. The
    new modality surfaces them at discovery time with 4 hits in one file.

  VerifyGate     +upgrade duty and +apply-reflection check (exit 7).
    0 of 96 verdicts were upgrades, because the reviewers were briefed only to refute.
    That silence is a property of the brief, not evidence of completeness — so `upgrade`
    and `inventoryGap` are now first-class verdicts and the gate warns at zero upgrades.
    Separately: a verdict the artifact does not reflect is worse than an unreviewed
    finding, because the scorecard claims it was checked. A reviewer caught the RPC
    bridges and the apply step reverted it. Re-merging reviews after the apply step has
    run is the specific mechanism, and it recurred while building this very gate.

  Prioritize     +absence-evidence check. A finding asserting a control is absent must
    cite a line it read or a negative search it ran. First version demanded file:line,
    which is incoherent for an absence — corrected to accept a named sweep.

  RenderReport   +§7c Corrections, rendered from corrections.json. Retractions
    accumulating as prose across five documents is how a corrected analysis becomes
    less readable than an uncorrected one.

  Start.md       Q1 is now model selection, asked as a PAIRING (author + attacker) so
    the same model cannot land in both slots and fail `stpa verify` AFTER the analysis
    is paid for. Enumerate the live runtime roster; prefer cross-vendor over cross-tier.
    STRIDE removed from the intake and from FullAnalysis: a same-model per-element pass
    yields a circular cross-check, and a real dual-method assessment belongs in its own
    skill.

  ComposeChains  dir-arg fix (also in #4, harmless if that merges first).

VERIFIED against the live analysis, positively and negatively: each gate passes on the
corrected artifacts and fails when its evidence is stripped. Two false-positive classes
were found and fixed by running them ("443 open" from a port range; "closes 12 findings"
cluster-leverage counts), because a gate that cries wolf is a gate that gets bypassed.
No dependencies added.
CI caught a real overreach in the absence-evidence check: it failed the shipped
worked example.

`Examples/ledgerline` is a Modality-B analysis — `scope.sources` is
`design-doc:DESIGN.md`, with no codebase at all. Its findings already cited evidence
properly (`location: "DESIGN.md — 'Invoice service — scoped by the orgId the gateway
attached'"`), but the check demanded `path.ext:line`, which a design document cannot
have. Impossible by construction, not a missing citation.

So the rule now keys off what the analysis was derived FROM: when no codebase appears
in `scope.sources`, a reference to the document IS the evidence. A codebase analysis
still requires a line that was read or a negative search that was run.

Verified all three ways:
  Examples/ledgerline (design-doc)          -> exit 0
  live codebase analysis, 74 findings       -> exit 0
  same analysis with the evidence stripped  -> exit 1

Also replicated the full smoke suite locally: grid arithmetic, the example at 100%,
scope + plan gates, and the self-contained report all pass.

This is the second false-positive class this check has produced (after "443 open"
matching a port range and "closes 12 findings" reading a leverage count as a total).
Each was found by running it rather than reasoning about it, which is the same lesson
the gate itself enforces on absence claims.
@MyAlterLego
MyAlterLego merged commit b8fad4c into main Jul 25, 2026
2 checks passed
@MyAlterLego
MyAlterLego deleted the feat/analysis-time-evidence-gates branch July 25, 2026 23:08
MyAlterLego pushed a commit that referenced this pull request Jul 25, 2026
The gates added in #6 each exist because a specific analysis shipped a specific
error. That reasoning lived only in a local SKILL.md, which is not in this repo and
sat on an ephemeral filesystem — so the checks would have outlived the explanation
of why they are worth their friction, and the first person to hit a red gate would
have had no way to judge whether it was telling them something true.

Records seven failure modes with the evidence: four false absence claims (one rated
band-1 and recorded as CONFIRMED); 2,508 guard candidates the analysis never
consumed; a belief credited as a working control because sourcedFrom named a hop
rather than an origin, invalidating six remediations; two generic RPC bridges that
survived two passes and a review; a one-directional review returning 0 upgrades in
96 verdicts; a verdict filed but reverted by the apply step; and a hand-typed count
that drifted from the grid.

Also records the three false positives these gates produced on first run, because a
gate that cries wolf gets disabled — that is a failure of the gate, not a cosmetic
issue.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants