Skip to content

Repository files navigation

grad-flow

A governed, end-to-end evaluation framework for a retrieval-augmented generation (RAG) assistant in a financial research context.

The system under evaluation is a single-pass RAG pipeline with explicit abstention: it retrieves supporting passages, generates an answer conditioned on them, and declines to answer when the retrieved context is insufficient. Control flow is fixed rather than model-directed — retrieval occurs once, before generation — which distinguishes it from an agentic system and determines what must be measured.

The system is built around two architectural constraints that apply to LLM deployments in regulated environments: all model traffic is routed through a single policy-enforcing gateway, and all changes to the pipeline are gated by an automated evaluation scorecard rather than subjective review.

The framework can run end-to-end without credentials or cost. When no provider API key is configured, the gateway returns deterministic mock responses, allowing the full control flow to be exercised offline. Configuring a provider key switches the same code path to live inference against Anthropic or Google Gemini models.

golden set ─▶ harness ─▶ pipeline ─▶ GATEWAY ─▶ model
                            │                      │
                            ▼                      ▼
                         metrics ◀──── answer + retrieved chunks
                            │
                            ▼
                    regression gate ─▶ baseline

Table of contents


Installation

make install

This creates a local virtual environment at .venv and installs the gateway and ragpipeline packages in editable mode, together with the development toolchain. All subsequent make targets resolve the virtual environment automatically; manual activation is not required.

Python 3.10 or later is required.


Usage

Quick verification

make demo

make demo starts the gateway, executes the evaluation suite, records a baseline, re-runs the suite to confirm the regression gate reports no change, and performs a multi-model comparison. The entire sequence runs offline at no cost.

Execution modes

The two entry points differ in process lifetime, and the distinction is significant in practice.

Mode Lifetime Intended use
make demo Ephemeral. The gateway is started in the background and terminated when the script exits. A single self-contained verification run that leaves no residual process. Suitable for CI or initial inspection.
make gateway Persistent. The gateway runs in the foreground until interrupted. Interactive model requests, or running evaluation stages against a live gateway.

Individual stages

make ingest          # build the vector index from data-platform/corpus/ (online, keyed)
make gateway         # terminal 1: gateway on port 8091 (override with PORT=9000)
make evals           # terminal 2: execute the evaluation suite, produce a scorecard
make baseline        # promote the current scorecard to the approved baseline
make evals           # re-run: the regression gate compares against the baseline
make compare-models  # multi-model comparison, as a model × metric matrix
make test            # unit test suite (offline)

The evaluation suite executes the golden set defined in data-platform/golden_set.jsonl, which currently contains five test cases.

make ingest is a one-off build step, run when the corpus changes rather than on every evaluation. It embeds the corpus (and the golden-set questions) into data-platform/index/, which is committed; the evaluation loop reads that index and never re-ingests. It requires a GEMINI_API_KEY to build the real semantic index — without one it builds a deterministic mock index for plumbing tests only.


Making a model request

With the gateway running, any registered model may be invoked through the single /v1/chat endpoint:

curl -sS -X POST http://localhost:8091/v1/chat \
  -H "Authorization: Bearer dev-research-key" \
  -H "Content-Type: application/json" \
  -d '{"model":"gemini-3.5-flash-lite",
       "messages":[{"role":"user","content":"Explain a 10-K in one sentence."}]}'

The response contains the generated text together with token usage and metered cost:

{"request_id":"…","model":"gemini-3.5-flash-lite",
 "text":"…answer…",
 "usage":{"input_tokens":9,"output_tokens":19,"cost_usd":0.0}}

Request parameters and constraints

Registered models. gemini-3.5-flash-lite (default), gemini-3.5-flash, gemini-3.6-flash, claude-sonnet-5, claude-haiku-4-5, and local-llama for chat, plus gemini-embedding-001 for retrieval embeddings (reached through /v1/embed, not /v1/chat). The authoritative definitions are in gateway/config.yaml. Provider model identifiers are subject to deprecation and tier restrictions; model-governance/registry.md documents how to enumerate the models available to a given API key.

Embedding requests. POST /v1/embed with {"model", "input": [...]} embeds one or more texts through the same policy-enforcing path as chat. This is how the corpus is ingested and how queries are embedded at retrieval time; there is no direct-to-provider path for embeddings either.

Live inference. Copy .env.example to .env and set ANTHROPIC_API_KEY and/or GEMINI_API_KEY, then restart the gateway. The same request path then reaches the live provider, and the evaluation suite scores against real model output.

Authorisation. Each API key maps to a team — the unit of authorisation, budgeting, and rate limiting, defined in gateway/config.yaml and explained in gateway/README.md. Each team carries an explicit model allowlist. dev-research-key may invoke every registered model; dev-evals-key is restricted to the external models and may not invoke local-llama. A request for a model outside the caller's allowlist returns HTTP 403.

Data classification. A request may declare "data_classification": "restricted". Restricted requests are rejected for any model marked external: true and may only be served by in-VPC models. This constraint is also enforced during failover.

Port configuration. The gateway listens on port 8091 by default, chosen to avoid the frequent conflicts on port 8080. Override with make gateway PORT=9000 and adjust the request URL accordingly.


Architecture

flowchart TD
    subgraph CICD["regression-gate/ — enforcement"]
        TRIGGER["PR / nightly trigger"]
        GATE["Regression gate<br/>vs approved baseline"]
        BASELINE[("baseline.json")]
    end
    subgraph DATA["data-platform/"]
        GOLD[("golden_set.jsonl<br/>analyst-verified cases")]
        FIX[("index/<br/>chunk vectors + query cache")]
        ANNOT["annotation<br/>(analyst verify/grade)"]
    end
    subgraph SANDBOX["sandbox/ — experiment env"]
        BAKE["compare_models.py<br/>model x metric matrix"]
    end
    subgraph HARNESS["evalharness/ — scoring machinery"]
        RUN["run_eval.py<br/>(lm-eval | native)"]
        METRICS["5 decomposed metrics<br/>utils.py"]
    end
    subgraph PIPE["ragpipeline/ — system under test"]
        RET["Retriever"]
        GEN["Generator (+abstain)"]
    end
    subgraph GW["gateway/ — policy choke point"]
        GATEWAY["authN/Z · data-class routing<br/>PII redact · cost · audit · failover"]
    end
    subgraph MODELS["models/ + providers"]
        ANT["Anthropic (Claude)"]
        GEM["Google (Gemini)"]
        LOCAL["self-hosted (in-VPC)"]
        MOCK["offline mock"]
    end
    subgraph OBS["observability/"]
        STORE[("results/<br/>scorecards + samples")]
        AUDIT[("audit/audit.jsonl")]
    end

    TRIGGER --> RUN
    GOLD --> RUN
    RUN --> PIPE
    RET --> FIX
    GEN --> GATEWAY
    RUN --> METRICS
    METRICS -->|"fuzzy scoring only"| GATEWAY
    GATEWAY --> ANT & GEM & LOCAL & MOCK
    METRICS --> STORE
    STORE --> GATE
    BASELINE --> GATE
    GATE -->|pass: promote| BASELINE
    GATE -->|fail: block| TRIGGER
    GATEWAY -.->|all calls logged| AUDIT
    BAKE --> RUN
    ANNOT --> GOLD
Loading

Request lifecycle

An evaluation run invokes evalharness/run_eval.py against the golden set. For each test case the harness executes the production ragpipeline modules (retrieval followed by generation), which reach the model exclusively through the gateway.

Every gateway request passes through the following sequence:

authenticate → rate limit → budget → authorize (model allowlist + data class)
  → redact PII → upstream call with failover → meter cost → append audit record

Results are written to observability/results/. The regression gate (regression-gate/compare_to_baseline.py) compares the run against the approved baseline and returns a non-zero exit status if any metric has regressed beyond tolerance.

Evaluation metrics

Scoring is defined in evalharness/tasks/utils.py. The metrics are deliberately decomposed so that a failing evaluation identifies which pipeline stage regressed.

Metric Method Failure mode detected
retrieval_recall Programmatic, using gold chunk identifiers Retrieval failures
numeric_exact Programmatic regex extraction with unit scaling; strict equality Magnitude errors, e.g. $4.2B reported as $4.2M
answer_correct LLM judge against a calibrated rubric Incorrect substantive conclusions
groundedness LLM judge, evaluated claim by claim Hallucination
abstention_correct Programmatic Fabrication when the source documents are silent

Invariants

Two properties are treated as non-negotiable and should be preserved in any future modification:

  1. All model traffic traverses the gateway, including judge invocations issued by the evaluation harness. No direct-to-provider path exists. Consequently, evaluation traffic is subject to the same authentication, redaction, routing, cost metering, and audit logging as production traffic, and the audit log provides evidence of this.
  2. Numeric facts are never evaluated by an LLM judge. numeric_exact performs regex extraction with billion/million/percent scaling and exact comparison. It is assigned zero regression tolerance, as is abstention_correct.

Design decisions

Component-oriented directory layout. Each top-level directory corresponds to a stage of the evaluation flow. The project is scoped to run on a single machine; production infrastructure (Redis, Presidio, PostgreSQL, a managed vector database) is documented as substitution points rather than implemented.

Separation of ragpipeline, sandbox, and regression-gate. These address distinct concerns and are intentionally not merged:

  • ragpipeline/ is the system under test: a retriever and a generator in a single configuration.
  • sandbox/ is the experiment environment. It executes the harness across a configuration sweep to compare multiple models cross-sectionally.
  • regression-gate/ is the regression gate. It compares the current run against an approved baseline longitudinally and blocks merges on regression.

sandbox and ci share the same metric implementations; scoring logic is not duplicated.

Multi-provider support from the outset. Anthropic and Google Gemini models are both reachable through the gateway. Adding a provider required changes only within the gateway, and switching the generator or judge between providers is a single-line configuration change.

Real transport with an offline default. The pipeline always issues a genuine HTTP request to the gateway. The gateway returns a deterministic mock response unless a provider key is present, so the policy and audit paths are exercised identically in both modes. A single environment variable switches between them.

Embeddings are governed like every other model call, and cached for determinism. Retrieval embeds the corpus (at ingest) and the query (at search) through the gateway's /v1/embed, so embedding traffic is subject to the same authentication, redaction, cost metering, and audit as generation — there is no direct-to-provider path. Two embedding operations are distinct: chunk embedding at ingest time, which is stored and searched, and query embedding at search time, which is what the search runs with. Because the golden-set questions are a fixed set, ingestion embeds them too and commits their vectors as a query-vector cache, so offline evaluation replays real embeddings deterministically with no key — the property the regression gate depends on. The retriever records the index manifest (model, dimension) and hard-fails on any mismatch rather than scoring retrieval on incomparable vectors. A managed vector store (pgvector) would replace the chunk-side storage and search, not the query cache; that substitution is planned for a later milestone.

lm-eval is the preferred harness engine, but optional. evalharness/run_eval.py uses lm_eval.simple_evaluate when the package is installed and otherwise falls back to a native runner that drives the same pipeline through the same scoring functions. This demonstrates that the metrics are engine-independent and would port unchanged to alternatives such as promptfoo or Braintrust. The native engine is the one exercised in environments where torch is unavailable, including Python 3.14; pip install -e '.[eval]' installs lm-eval where supported.


Repository structure

Each top-level module owns one responsibility in the evaluation flow. The boundaries are deliberate: they determine what can change independently, and which component a failing metric implicates.

Module Responsibility
gateway/ Enforces policy on every model request. It is the sole egress point to any provider, and the only component holding provider credentials. Decides whether a request may proceed (authentication, allowlists, data classification, budgets, rate limits), what it may reach (routing and failover), and records what occurred (cost metering, audit log).
ragpipeline/ The system under test. Implements the assistant itself — retrieval over the document index, then answer generation, including the decision to abstain when the retrieved context does not support an answer. This is production code, executed unmodified by the harness. Retrieval is vector search over an ingested corpus (data-platform/index/), with the query embedded through the gateway's /v1/embed; a committed query-vector cache keeps offline runs deterministic.
evalharness/ Defines what "correct" means and measures it. Owns the five decomposed metrics, the task definition, and the runner that executes the pipeline against the golden set to produce a scorecard. It is the measuring instrument, deliberately separate from the system it measures.
sandbox/ Compares alternatives. Executes the harness across a sweep of configurations to evaluate competing models or settings against one another under identical conditions. Answers "which option performs best?"
regression-gate/ Prevents regression. Compares a run against the approved baseline and fails the build when a metric degrades beyond its tolerance. Answers "did this change make the system worse?"
data-platform/ Owns the evaluation assets and the data that feeds them: the analyst-verified golden set, the retrieval corpus and index, and the annotation process by which production failures become new test cases. The golden set is the durable asset; it outlasts any framework choice.
models/ Governs model selection rather than serving inference. Documents which models are approved for which purpose, and holds the judge calibration procedure that establishes whether judge-scored metrics can be trusted.
observability/ Records evidence. Retains per-run scorecards and per-sample logs for debugging and for reproducing any published result, and the gateway audit trail demonstrating that all traffic passed through the governed path.

sandbox/ and regression-gate/ apply the same metric implementations along different axes: sandbox/ compares options against each other at a point in time, while regression-gate/ compares the current state against a previously approved one. Neither reimplements scoring.


Testing and code quality

The development toolchain is installed by make install via the [dev] extra. None of the commands in this section perform network requests or incur cost.

Commands are shown with the uv run prefix, which resolves the project environment without requiring manual activation. Equivalent alternatives are the make targets, or activating .venv once and omitting the prefix.

Dependency groups

Dependencies are separated into three groups in pyproject.toml. uv run installs only the primary group by default, which is why --extra dev is required to expose the test and lint tooling.

Group Contents Purpose
dependencies fastapi, uvicorn, httpx, pydantic, pyyaml Runtime requirements
[dev] extra pytest, pytest-cov, ruff, mypy Testing and static analysis
[eval] extra lm-eval Optional harness engine

lm-eval is isolated because it is heavy — roughly 45 transitive packages including pandas, scipy, scikit-learn and pyarrow — not because it is difficult to install. It requires no torch for this use (torch is an optional backend for running local models) and installs cleanly on Python 3.14. The native engine exists so the base install stays small and the flow runs anywhere, not as a workaround.

Unit tests

make test                                              # full suite
uv run --extra dev pytest -v                           # verbose output
uv run --extra dev pytest tests/test_policy.py         # single module
uv run --extra dev pytest -k "restricted or numeric"   # select by name
uv run --extra dev pytest -x --lf                      # halt on first failure, then rerun failures

Coverage

uv run --extra dev pytest \
  --cov=gateway --cov=ragpipeline --cov=harness --cov=sandbox \
  --cov-report=term-missing

uv run --extra dev pytest \
  --cov=gateway --cov-report=html        # report written to htmlcov/index.html

Aggregate coverage is approximately 38%. The distribution is more informative than the aggregate figure:

Module Coverage Notes
gateway/policy.py ~98% Policy enforcement; covered to near-completion by design
evalharness/tasks/utils.py ~94% Metric implementations; covered to near-completion by design
gateway/main.py ~59% Retry logic covered; the request handler is not
gateway/providers.py ~41% Mock and retry parsing covered; live HTTP adapters are not
ragpipeline/retrieval.py ~100% Ranking and the manifest hard-fails; covered by design
ragpipeline/generation.py 0% Abstention branch is a known gap
run_eval.py, compare_models.py 0% I/O orchestration; low unit-test value

Coverage is concentrated in the logic that carries correctness and compliance risk. The uncovered modules are predominantly I/O orchestration, where unit tests provide limited assurance relative to their cost. ragpipeline/retrieval.py — cosine ranking and the manifest hard-fails — is covered by tests/test_retrieval.py. The remaining gap that merits attention is abstention detection in generation.py, since abstention is a scored metric and an undetected regression there would invalidate a scorecard.

Static analysis

Ruff is used for linting and formatting. It consolidates the functionality of flake8, isort, pyupgrade, a substantial portion of pylint, and part of bandit. Ruff does not perform type checking, so mypy is used alongside it.

uv run --extra dev ruff check .              # lint using the project configuration
uv run --extra dev ruff check --statistics . # per-rule counts
uv run --extra dev ruff check --fix .        # apply safe automatic fixes
uv run --extra dev ruff format .             # format
uv run --extra dev ruff check --show-files . # list files subject to linting

uv run --extra dev mypy gateway ragpipeline     # type checking

The rule selection is defined in [tool.ruff.lint] in pyproject.toml: E, F, W (pycodestyle and pyflakes), B (bugbear), UP (pyupgrade), S (bandit), SIM (simplify), and I (import ordering). Alternative rule sets may be applied for a single invocation without modifying the committed configuration:

uv run --extra dev ruff check --select ALL --statistics .
uv run --extra dev ruff check --select B --output-format concise .

Two exclusions are deliberate and documented in pyproject.toml:

  • claude-discussions/ is excluded. It is an archived design artifact rather than maintained source, and linting it produced no actionable output.
  • S101 is suppressed under tests/, and S603 globally. Test code is expected to use bare assert statements, which accounted for 68 of 88 initial findings. S603 reports the subprocess invocations in sandbox/compare_models.py, which pass an argument list without a shell and are therefore not a security concern.

Testing strategy

The repository contains two complementary forms of testing.

Unit tests (tests/) verify the implementation. They are fast, offline, and concentrated on behaviour that must not regress silently: the restricted-data routing rule, team allowlists, budgets, rate limits, PII redaction, strict numeric matching, abstention branches, and retry behaviour. Two tests guard against configuration drift by asserting that every allowlist entry and every failover target resolves to a defined model.

The evaluation harness verifies the behaviour of the assistant itself, through the golden set, the five decomposed metrics, and the regression gate.

The two are not substitutes. Unit tests cannot establish answer quality, and evaluations cannot establish that the policy layer is correctly enforced.


Security and data handling

This repository is intended for public distribution and has been prepared accordingly.

  • No credentials are committed. Provider keys are read from .env, which is excluded from version control. The gateway configuration references keys by environment variable name only. Run artefacts (observability/results/, audit/*.jsonl, regression-gate/baseline.json) and .venv/ are likewise excluded.
  • All data is synthetic. The golden set refers to a fictional entity and contains no real company data or material non-public information.
  • Two files are committed by design. gateway/config.yaml contains illustrative budgets and pricing but no secrets, and data-platform/golden_set.jsonl is the evaluation asset. Both carry explanatory notes. As documented in data-platform/README.md, if real corpus data or material non-public information is introduced into the golden set, the repository must be made private.

The gateway audit log records request metadata only — team, requested and served model, data classification, token counts, cost, latency, and status — and does not record prompt or response content.


Troubleshooting

When diagnosing a failed request, use curl -sS rather than curl -s. The -s flag suppresses transport errors and returns an empty response, concealing the underlying cause.

Symptom Probable cause Resolution
Empty response curl -s is suppressing a connection error Re-run with curl -sS and refer to the rows below
curl: (7) Failed to connect The gateway is not running, or the port is incorrect Start the gateway with make gateway and confirm the port (8091 by default)
Plain-text 404 page not found Another service is bound to that port The gateway returns JSON exclusively. Select a free port with make gateway PORT=9000 or stop the conflicting service
uvicorn: command not found Commands are being run outside the project virtual environment Run make install; make targets and scripts then resolve .venv automatically
Gateway unreachable after make demo make demo terminates the gateway on exit Use make gateway for a persistent instance
Response contains mock placeholder text No provider key is configured Set ANTHROPIC_API_KEY and/or GEMINI_API_KEY in .env and restart the gateway
HTTP 502 during a multi-model run Provider rate limit exhausted The gateway honours the provider's retry interval automatically. Free-tier quotas are low; reduce the number of models under comparison or the golden set size

Roadmap

M1 — End-to-end framework (complete). The full evaluation loop executes offline, gateway policy is enforced and audited, and multi-model comparison is operational.

M2 — Retrieval implementation (complete). data-platform/ingest/ingest.py chunks and embeds the synthetic corpus in data-platform/corpus/ into a local vector index, and ragpipeline/retrieval.py performs vector search over that index behind the established interface. Embedding traffic traverses the gateway's /v1/embed endpoint. The golden-set questions are embedded at ingest time into a committed query-vector cache, so offline evaluation is deterministic and requires no key. Retrieval measures substantive quality, and the multi-model comparison discriminates between models.

Ingestion generates its own chunk_id values, so the order is fixed: build the index, re-annotate every gold_chunk_ids in golden_set.jsonl against the produced identifiers, then rely on the retriever. tests/test_golden_set.py enforces referential integrity. Expanding the golden set to validate retrieval quality is deferred to M3, where it is paired with judge calibration.

M3 — Golden set expansion and judge calibration. Expand the golden set to 100–300 analyst-verified cases and run model-governance/judge/calibrate_judge.py to establish at least 90% agreement with analyst grading before judge-scored metrics are relied upon. Add a reporting view over observability/. Objective: establish metric trustworthiness.

On completion, update:

  • ragpipeline/llm.py — resolve the TODO(M3): judge and generator currently default to the same model (self-evaluation bias)
  • Record the calibration evidence in model-governance/judge/ as required by its README
  • Revisit the multi-model comparison: five golden cases cannot discriminate between models; a larger set can

M4 — Continuous integration. Execute the regression gate on merge requests affecting ragpipeline/, prompts, or configuration; block merges on regression and promote baselines on approval. Objective: automate enforcement.

On completion, update:

  • Decide how the baseline is shared. regression-gate/baseline.json is currently git-ignored, so a fresh runner has none — and run_evals.sh exits 0 when the baseline is absent, meaning CI would report success while gating nothing. Either commit an approved baseline or make its absence a hard failure in CI.
  • Revisit the regression-gate/ name (TODO(M4) in CLAUDE.md). It was renamed from ci/ because that described its consumer rather than its function. Note .gitlab-ci.yml must sit at the repository root regardless.
  • Ensure the pipeline runs bash regression-gate/run_evals.sh rather than reimplementing the sequence.

M5 — Production feedback loop. Deploy the assistant, sample groundedness in production, collect user feedback, monitor drift, and route observed failures back into the golden set through data-platform/annotation/. Objective: ensure the evaluation suite reflects production conditions.

References

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages