A governed, end-to-end evaluation framework for a retrieval-augmented generation (RAG) assistant in a financial research context.
The system under evaluation is a single-pass RAG pipeline with explicit abstention: it retrieves supporting passages, generates an answer conditioned on them, and declines to answer when the retrieved context is insufficient. Control flow is fixed rather than model-directed — retrieval occurs once, before generation — which distinguishes it from an agentic system and determines what must be measured.
The system is built around two architectural constraints that apply to LLM deployments in regulated environments: all model traffic is routed through a single policy-enforcing gateway, and all changes to the pipeline are gated by an automated evaluation scorecard rather than subjective review.
The framework can run end-to-end without credentials or cost. When no provider API key is configured, the gateway returns deterministic mock responses, allowing the full control flow to be exercised offline. Configuring a provider key switches the same code path to live inference against Anthropic or Google Gemini models.
golden set ─▶ harness ─▶ pipeline ─▶ GATEWAY ─▶ model
│ │
▼ ▼
metrics ◀──── answer + retrieved chunks
│
▼
regression gate ─▶ baseline
- Installation
- Usage
- Making a model request
- Architecture
- Design decisions
- Repository structure
- Testing and code quality
- Security and data handling
- Troubleshooting
- Roadmap
make installThis creates a local virtual environment at .venv and installs the gateway
and ragpipeline packages in editable mode, together with the development
toolchain. All subsequent make targets resolve the virtual environment
automatically; manual activation is not required.
Python 3.10 or later is required.
make demomake demo starts the gateway, executes the evaluation suite, records a
baseline, re-runs the suite to confirm the regression gate reports no change,
and performs a multi-model comparison. The entire sequence runs offline at no
cost.
The two entry points differ in process lifetime, and the distinction is significant in practice.
| Mode | Lifetime | Intended use |
|---|---|---|
make demo |
Ephemeral. The gateway is started in the background and terminated when the script exits. | A single self-contained verification run that leaves no residual process. Suitable for CI or initial inspection. |
make gateway |
Persistent. The gateway runs in the foreground until interrupted. | Interactive model requests, or running evaluation stages against a live gateway. |
make ingest # build the vector index from data-platform/corpus/ (online, keyed)
make gateway # terminal 1: gateway on port 8091 (override with PORT=9000)
make evals # terminal 2: execute the evaluation suite, produce a scorecard
make baseline # promote the current scorecard to the approved baseline
make evals # re-run: the regression gate compares against the baseline
make compare-models # multi-model comparison, as a model × metric matrix
make test # unit test suite (offline)The evaluation suite executes the golden set defined in
data-platform/golden_set.jsonl, which currently contains five test cases.
make ingest is a one-off build step, run when the corpus changes rather than on
every evaluation. It embeds the corpus (and the golden-set questions) into
data-platform/index/, which is committed; the evaluation loop reads that index
and never re-ingests. It requires a GEMINI_API_KEY to build the real semantic
index — without one it builds a deterministic mock index for plumbing tests only.
With the gateway running, any registered model may be invoked through the single
/v1/chat endpoint:
curl -sS -X POST http://localhost:8091/v1/chat \
-H "Authorization: Bearer dev-research-key" \
-H "Content-Type: application/json" \
-d '{"model":"gemini-3.5-flash-lite",
"messages":[{"role":"user","content":"Explain a 10-K in one sentence."}]}'The response contains the generated text together with token usage and metered cost:
Registered models. gemini-3.5-flash-lite (default), gemini-3.5-flash,
gemini-3.6-flash, claude-sonnet-5, claude-haiku-4-5, and local-llama for
chat, plus gemini-embedding-001 for retrieval embeddings (reached through
/v1/embed, not /v1/chat). The authoritative definitions are in
gateway/config.yaml. Provider model identifiers are subject to deprecation and
tier restrictions; model-governance/registry.md documents how to enumerate the
models available to a given API key.
Embedding requests. POST /v1/embed with {"model", "input": [...]} embeds
one or more texts through the same policy-enforcing path as chat. This is how the
corpus is ingested and how queries are embedded at retrieval time; there is no
direct-to-provider path for embeddings either.
Live inference. Copy .env.example to .env and set ANTHROPIC_API_KEY
and/or GEMINI_API_KEY, then restart the gateway. The same request path then
reaches the live provider, and the evaluation suite scores against real model
output.
Authorisation. Each API key maps to a team — the unit of authorisation,
budgeting, and rate limiting, defined in gateway/config.yaml and explained in
gateway/README.md. Each team carries an explicit model allowlist.
dev-research-key may invoke every registered model; dev-evals-key is
restricted to the external models and may not invoke local-llama. A request for
a model outside the caller's allowlist returns HTTP 403.
Data classification. A request may declare
"data_classification": "restricted". Restricted requests are rejected for any
model marked external: true and may only be served by in-VPC models. This
constraint is also enforced during failover.
Port configuration. The gateway listens on port 8091 by default, chosen to
avoid the frequent conflicts on port 8080. Override with make gateway PORT=9000
and adjust the request URL accordingly.
flowchart TD
subgraph CICD["regression-gate/ — enforcement"]
TRIGGER["PR / nightly trigger"]
GATE["Regression gate<br/>vs approved baseline"]
BASELINE[("baseline.json")]
end
subgraph DATA["data-platform/"]
GOLD[("golden_set.jsonl<br/>analyst-verified cases")]
FIX[("index/<br/>chunk vectors + query cache")]
ANNOT["annotation<br/>(analyst verify/grade)"]
end
subgraph SANDBOX["sandbox/ — experiment env"]
BAKE["compare_models.py<br/>model x metric matrix"]
end
subgraph HARNESS["evalharness/ — scoring machinery"]
RUN["run_eval.py<br/>(lm-eval | native)"]
METRICS["5 decomposed metrics<br/>utils.py"]
end
subgraph PIPE["ragpipeline/ — system under test"]
RET["Retriever"]
GEN["Generator (+abstain)"]
end
subgraph GW["gateway/ — policy choke point"]
GATEWAY["authN/Z · data-class routing<br/>PII redact · cost · audit · failover"]
end
subgraph MODELS["models/ + providers"]
ANT["Anthropic (Claude)"]
GEM["Google (Gemini)"]
LOCAL["self-hosted (in-VPC)"]
MOCK["offline mock"]
end
subgraph OBS["observability/"]
STORE[("results/<br/>scorecards + samples")]
AUDIT[("audit/audit.jsonl")]
end
TRIGGER --> RUN
GOLD --> RUN
RUN --> PIPE
RET --> FIX
GEN --> GATEWAY
RUN --> METRICS
METRICS -->|"fuzzy scoring only"| GATEWAY
GATEWAY --> ANT & GEM & LOCAL & MOCK
METRICS --> STORE
STORE --> GATE
BASELINE --> GATE
GATE -->|pass: promote| BASELINE
GATE -->|fail: block| TRIGGER
GATEWAY -.->|all calls logged| AUDIT
BAKE --> RUN
ANNOT --> GOLD
An evaluation run invokes evalharness/run_eval.py against the golden set. For each
test case the harness executes the production ragpipeline modules (retrieval
followed by generation), which reach the model exclusively through the gateway.
Every gateway request passes through the following sequence:
authenticate → rate limit → budget → authorize (model allowlist + data class)
→ redact PII → upstream call with failover → meter cost → append audit record
Results are written to observability/results/. The regression gate
(regression-gate/compare_to_baseline.py) compares the run against the approved baseline and
returns a non-zero exit status if any metric has regressed beyond tolerance.
Scoring is defined in evalharness/tasks/utils.py. The metrics are deliberately
decomposed so that a failing evaluation identifies which pipeline stage
regressed.
| Metric | Method | Failure mode detected |
|---|---|---|
retrieval_recall |
Programmatic, using gold chunk identifiers | Retrieval failures |
numeric_exact |
Programmatic regex extraction with unit scaling; strict equality | Magnitude errors, e.g. $4.2B reported as $4.2M |
answer_correct |
LLM judge against a calibrated rubric | Incorrect substantive conclusions |
groundedness |
LLM judge, evaluated claim by claim | Hallucination |
abstention_correct |
Programmatic | Fabrication when the source documents are silent |
Two properties are treated as non-negotiable and should be preserved in any future modification:
- All model traffic traverses the gateway, including judge invocations issued by the evaluation harness. No direct-to-provider path exists. Consequently, evaluation traffic is subject to the same authentication, redaction, routing, cost metering, and audit logging as production traffic, and the audit log provides evidence of this.
- Numeric facts are never evaluated by an LLM judge.
numeric_exactperforms regex extraction with billion/million/percent scaling and exact comparison. It is assigned zero regression tolerance, as isabstention_correct.
Component-oriented directory layout. Each top-level directory corresponds to a stage of the evaluation flow. The project is scoped to run on a single machine; production infrastructure (Redis, Presidio, PostgreSQL, a managed vector database) is documented as substitution points rather than implemented.
Separation of ragpipeline, sandbox, and regression-gate. These address distinct
concerns and are intentionally not merged:
ragpipeline/is the system under test: a retriever and a generator in a single configuration.sandbox/is the experiment environment. It executes the harness across a configuration sweep to compare multiple models cross-sectionally.regression-gate/is the regression gate. It compares the current run against an approved baseline longitudinally and blocks merges on regression.
sandbox and ci share the same metric implementations; scoring logic is not
duplicated.
Multi-provider support from the outset. Anthropic and Google Gemini models are both reachable through the gateway. Adding a provider required changes only within the gateway, and switching the generator or judge between providers is a single-line configuration change.
Real transport with an offline default. The pipeline always issues a genuine HTTP request to the gateway. The gateway returns a deterministic mock response unless a provider key is present, so the policy and audit paths are exercised identically in both modes. A single environment variable switches between them.
Embeddings are governed like every other model call, and cached for
determinism. Retrieval embeds the corpus (at ingest) and the query (at search)
through the gateway's /v1/embed, so embedding traffic is subject to the same
authentication, redaction, cost metering, and audit as generation — there is no
direct-to-provider path. Two embedding operations are distinct: chunk embedding
at ingest time, which is stored and searched, and query embedding at search
time, which is what the search runs with. Because the golden-set questions are a
fixed set, ingestion embeds them too and commits their vectors as a query-vector
cache, so offline evaluation replays real embeddings deterministically with no key
— the property the regression gate depends on. The retriever records the index
manifest (model, dimension) and hard-fails on any mismatch rather than scoring
retrieval on incomparable vectors. A managed vector store (pgvector) would replace
the chunk-side storage and search, not the query cache; that substitution is
planned for a later milestone.
lm-eval is the preferred harness engine, but optional. evalharness/run_eval.py
uses lm_eval.simple_evaluate when the package is installed and otherwise falls
back to a native runner that drives the same pipeline through the same scoring
functions. This demonstrates that the metrics are engine-independent and would
port unchanged to alternatives such as promptfoo or Braintrust. The native engine
is the one exercised in environments where torch is unavailable, including
Python 3.14; pip install -e '.[eval]' installs lm-eval where supported.
Each top-level module owns one responsibility in the evaluation flow. The boundaries are deliberate: they determine what can change independently, and which component a failing metric implicates.
| Module | Responsibility |
|---|---|
gateway/ |
Enforces policy on every model request. It is the sole egress point to any provider, and the only component holding provider credentials. Decides whether a request may proceed (authentication, allowlists, data classification, budgets, rate limits), what it may reach (routing and failover), and records what occurred (cost metering, audit log). |
ragpipeline/ |
The system under test. Implements the assistant itself — retrieval over the document index, then answer generation, including the decision to abstain when the retrieved context does not support an answer. This is production code, executed unmodified by the harness. Retrieval is vector search over an ingested corpus (data-platform/index/), with the query embedded through the gateway's /v1/embed; a committed query-vector cache keeps offline runs deterministic. |
evalharness/ |
Defines what "correct" means and measures it. Owns the five decomposed metrics, the task definition, and the runner that executes the pipeline against the golden set to produce a scorecard. It is the measuring instrument, deliberately separate from the system it measures. |
sandbox/ |
Compares alternatives. Executes the harness across a sweep of configurations to evaluate competing models or settings against one another under identical conditions. Answers "which option performs best?" |
regression-gate/ |
Prevents regression. Compares a run against the approved baseline and fails the build when a metric degrades beyond its tolerance. Answers "did this change make the system worse?" |
data-platform/ |
Owns the evaluation assets and the data that feeds them: the analyst-verified golden set, the retrieval corpus and index, and the annotation process by which production failures become new test cases. The golden set is the durable asset; it outlasts any framework choice. |
models/ |
Governs model selection rather than serving inference. Documents which models are approved for which purpose, and holds the judge calibration procedure that establishes whether judge-scored metrics can be trusted. |
observability/ |
Records evidence. Retains per-run scorecards and per-sample logs for debugging and for reproducing any published result, and the gateway audit trail demonstrating that all traffic passed through the governed path. |
sandbox/ and regression-gate/ apply the same metric implementations along different axes:
sandbox/ compares options against each other at a point in time, while regression-gate/
compares the current state against a previously approved one. Neither reimplements
scoring.
The development toolchain is installed by make install via the [dev] extra.
None of the commands in this section perform network requests or incur cost.
Commands are shown with the uv run prefix, which resolves the project
environment without requiring manual activation. Equivalent alternatives are the
make targets, or activating .venv once and omitting the prefix.
Dependencies are separated into three groups in pyproject.toml. uv run
installs only the primary group by default, which is why --extra dev is
required to expose the test and lint tooling.
| Group | Contents | Purpose |
|---|---|---|
dependencies |
fastapi, uvicorn, httpx, pydantic, pyyaml | Runtime requirements |
[dev] extra |
pytest, pytest-cov, ruff, mypy | Testing and static analysis |
[eval] extra |
lm-eval | Optional harness engine |
lm-eval is isolated because it is heavy — roughly 45 transitive packages
including pandas, scipy, scikit-learn and pyarrow — not because it is difficult
to install. It requires no torch for this use (torch is an optional backend
for running local models) and installs cleanly on Python 3.14. The native engine
exists so the base install stays small and the flow runs anywhere, not as a
workaround.
make test # full suite
uv run --extra dev pytest -v # verbose output
uv run --extra dev pytest tests/test_policy.py # single module
uv run --extra dev pytest -k "restricted or numeric" # select by name
uv run --extra dev pytest -x --lf # halt on first failure, then rerun failuresuv run --extra dev pytest \
--cov=gateway --cov=ragpipeline --cov=harness --cov=sandbox \
--cov-report=term-missing
uv run --extra dev pytest \
--cov=gateway --cov-report=html # report written to htmlcov/index.htmlAggregate coverage is approximately 38%. The distribution is more informative than the aggregate figure:
| Module | Coverage | Notes |
|---|---|---|
gateway/policy.py |
~98% | Policy enforcement; covered to near-completion by design |
evalharness/tasks/utils.py |
~94% | Metric implementations; covered to near-completion by design |
gateway/main.py |
~59% | Retry logic covered; the request handler is not |
gateway/providers.py |
~41% | Mock and retry parsing covered; live HTTP adapters are not |
ragpipeline/retrieval.py |
~100% | Ranking and the manifest hard-fails; covered by design |
ragpipeline/generation.py |
0% | Abstention branch is a known gap |
run_eval.py, compare_models.py |
0% | I/O orchestration; low unit-test value |
Coverage is concentrated in the logic that carries correctness and compliance
risk. The uncovered modules are predominantly I/O orchestration, where unit tests
provide limited assurance relative to their cost. ragpipeline/retrieval.py —
cosine ranking and the manifest hard-fails — is covered by
tests/test_retrieval.py. The remaining gap that merits attention is abstention
detection in generation.py, since abstention is a scored metric and an
undetected regression there would invalidate a scorecard.
Ruff is used for linting and formatting. It consolidates the functionality of flake8, isort, pyupgrade, a substantial portion of pylint, and part of bandit. Ruff does not perform type checking, so mypy is used alongside it.
uv run --extra dev ruff check . # lint using the project configuration
uv run --extra dev ruff check --statistics . # per-rule counts
uv run --extra dev ruff check --fix . # apply safe automatic fixes
uv run --extra dev ruff format . # format
uv run --extra dev ruff check --show-files . # list files subject to linting
uv run --extra dev mypy gateway ragpipeline # type checkingThe rule selection is defined in [tool.ruff.lint] in pyproject.toml: E,
F, W (pycodestyle and pyflakes), B (bugbear), UP (pyupgrade), S
(bandit), SIM (simplify), and I (import ordering). Alternative rule sets may
be applied for a single invocation without modifying the committed
configuration:
uv run --extra dev ruff check --select ALL --statistics .
uv run --extra dev ruff check --select B --output-format concise .Two exclusions are deliberate and documented in pyproject.toml:
claude-discussions/is excluded. It is an archived design artifact rather than maintained source, and linting it produced no actionable output.S101is suppressed undertests/, andS603globally. Test code is expected to use bareassertstatements, which accounted for 68 of 88 initial findings.S603reports thesubprocessinvocations insandbox/compare_models.py, which pass an argument list without a shell and are therefore not a security concern.
The repository contains two complementary forms of testing.
Unit tests (tests/) verify the implementation. They are fast, offline, and
concentrated on behaviour that must not regress silently: the restricted-data
routing rule, team allowlists, budgets, rate limits, PII redaction, strict
numeric matching, abstention branches, and retry behaviour. Two tests guard
against configuration drift by asserting that every allowlist entry and every
failover target resolves to a defined model.
The evaluation harness verifies the behaviour of the assistant itself, through the golden set, the five decomposed metrics, and the regression gate.
The two are not substitutes. Unit tests cannot establish answer quality, and evaluations cannot establish that the policy layer is correctly enforced.
This repository is intended for public distribution and has been prepared accordingly.
- No credentials are committed. Provider keys are read from
.env, which is excluded from version control. The gateway configuration references keys by environment variable name only. Run artefacts (observability/results/,audit/*.jsonl,regression-gate/baseline.json) and.venv/are likewise excluded. - All data is synthetic. The golden set refers to a fictional entity and contains no real company data or material non-public information.
- Two files are committed by design.
gateway/config.yamlcontains illustrative budgets and pricing but no secrets, anddata-platform/golden_set.jsonlis the evaluation asset. Both carry explanatory notes. As documented indata-platform/README.md, if real corpus data or material non-public information is introduced into the golden set, the repository must be made private.
The gateway audit log records request metadata only — team, requested and served model, data classification, token counts, cost, latency, and status — and does not record prompt or response content.
When diagnosing a failed request, use curl -sS rather than curl -s. The -s
flag suppresses transport errors and returns an empty response, concealing the
underlying cause.
| Symptom | Probable cause | Resolution |
|---|---|---|
| Empty response | curl -s is suppressing a connection error |
Re-run with curl -sS and refer to the rows below |
curl: (7) Failed to connect |
The gateway is not running, or the port is incorrect | Start the gateway with make gateway and confirm the port (8091 by default) |
Plain-text 404 page not found |
Another service is bound to that port | The gateway returns JSON exclusively. Select a free port with make gateway PORT=9000 or stop the conflicting service |
uvicorn: command not found |
Commands are being run outside the project virtual environment | Run make install; make targets and scripts then resolve .venv automatically |
Gateway unreachable after make demo |
make demo terminates the gateway on exit |
Use make gateway for a persistent instance |
| Response contains mock placeholder text | No provider key is configured | Set ANTHROPIC_API_KEY and/or GEMINI_API_KEY in .env and restart the gateway |
| HTTP 502 during a multi-model run | Provider rate limit exhausted | The gateway honours the provider's retry interval automatically. Free-tier quotas are low; reduce the number of models under comparison or the golden set size |
M1 — End-to-end framework (complete). The full evaluation loop executes offline, gateway policy is enforced and audited, and multi-model comparison is operational.
M2 — Retrieval implementation (complete). data-platform/ingest/ingest.py
chunks and embeds the synthetic corpus in data-platform/corpus/ into a local
vector index, and ragpipeline/retrieval.py performs vector search over that
index behind the established interface. Embedding traffic traverses the gateway's
/v1/embed endpoint. The golden-set questions are embedded at ingest time into a
committed query-vector cache, so offline evaluation is deterministic and requires
no key. Retrieval measures substantive quality, and the multi-model comparison
discriminates between models.
Ingestion generates its own chunk_id values, so the order is fixed: build the
index, re-annotate every gold_chunk_ids in golden_set.jsonl against the
produced identifiers, then rely on the retriever. tests/test_golden_set.py
enforces referential integrity. Expanding the golden set to validate retrieval
quality is deferred to M3, where it is paired with judge calibration.
M3 — Golden set expansion and judge calibration. Expand the golden set to
100–300 analyst-verified cases and run model-governance/judge/calibrate_judge.py to
establish at least 90% agreement with analyst grading before judge-scored metrics
are relied upon. Add a reporting view over observability/. Objective: establish
metric trustworthiness.
On completion, update:
-
ragpipeline/llm.py— resolve theTODO(M3): judge and generator currently default to the same model (self-evaluation bias) - Record the calibration evidence in
model-governance/judge/as required by its README - Revisit the multi-model comparison: five golden cases cannot discriminate between models; a larger set can
M4 — Continuous integration. Execute the regression gate on merge requests
affecting ragpipeline/, prompts, or configuration; block merges on regression and
promote baselines on approval. Objective: automate enforcement.
On completion, update:
- Decide how the baseline is shared.
regression-gate/baseline.jsonis currently git-ignored, so a fresh runner has none — andrun_evals.shexits 0 when the baseline is absent, meaning CI would report success while gating nothing. Either commit an approved baseline or make its absence a hard failure in CI. - Revisit the
regression-gate/name (TODO(M4)inCLAUDE.md). It was renamed fromci/because that described its consumer rather than its function. Note.gitlab-ci.ymlmust sit at the repository root regardless. - Ensure the pipeline runs
bash regression-gate/run_evals.shrather than reimplementing the sequence.
M5 — Production feedback loop. Deploy the assistant, sample groundedness in
production, collect user feedback, monitor drift, and route observed failures
back into the golden set through data-platform/annotation/. Objective: ensure
the evaluation suite reflects production conditions.
{"request_id":"…","model":"gemini-3.5-flash-lite", "text":"…answer…", "usage":{"input_tokens":9,"output_tokens":19,"cost_usd":0.0}}