Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Changelog

## 0.17.3 — Unreleased

- Add bounded handoff chunks and host-visible completeness checks with budgeted recovery.
- Exclude incomplete/unknown live inputs before optional grading and model comparisons.
- Record ordinary-work feedback without extra model calls, automatic training or policy changes.

## 0.17.2 — Unreleased

- Make Dynamic routing the ordinary automatic-selection wizard path on Codex and Claude.
Expand Down
2 changes: 1 addition & 1 deletion docs/contributing-agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,7 @@ workers or claim synthetic savings as observed results. See the

## Distribution and versioning

The current package version is 0.17.2; increment all three manifests for further
The current package version is 0.17.3; increment all three manifests for further
installer-visible changes. Existing native Claude agents/commands/hooks remain
preserved. Never rewrite historical plans to claim newer evidence. No repository
change implicitly installs personally, publishes a release or edits the separate
Expand Down
2 changes: 1 addition & 1 deletion docs/installation.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Guildhall installation and host support

Version **0.17.2 is a release candidate**, available through the `main` commands
Version **0.17.3 is a release candidate**, available through the `main` commands
below after merge. It includes all-role routing eligibility, usage/evidence/study tools, independent
`guildhall-quest` and `guildhall-routing-setup` skills and native Codex metadata. Claude's `/guildhall:quest`, nineteen agent definitions
and hooks remain available. Choose one route per quest to avoid duplicate entry
Expand Down
67 changes: 67 additions & 0 deletions docs/reviews/2026-09-29-dynamic-routing-verification.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
# Practical dynamic routing verification

The four-PR stack implements task-level Jev routing on Codex, native Claude and
standalone Claude. Final package version is 0.17.3. Installation remains off;
users can explicitly approve Dynamic routing without a qualification study.
Advanced adaptive retains its benchmark requirements. No personal settings,
marketplaces or saved study results were changed during this implementation.

## Stack and behavior

1. [PR 47](https://github.com/GrillerGeek/guildhall/pull/47): v5 dynamic contracts,
reviewed controls, versioned approvals, fallback and router-identity handling.
2. [PR 48](https://github.com/GrillerGeek/guildhall/pull/48): pinned host catalogs,
controlled task briefs, meaningful choice criteria and privacy boundaries.
3. [PR 49](https://github.com/GrillerGeek/guildhall/pull/49): wizard and all three
host routes, explicit role locks versus Claude/default fallbacks, migration.
4. `routing/delivery-integrity`: bounded handoffs, study delivery guards and local
feedback from ordinary work. This branch targets PR 49's branch for review.

Review and merge in order. The first three PRs have passing GitHub validation and
pinned/latest installer checks. The final PR should ship with the feature, while
benchmark qualification remains optional. These intermediate package versions
are a review stack, not a request to publish or install each step personally.

## Delivery and feedback

Required references become bounded identified chunks. Host-visible content is
checked for omissions, changed bytes, explicit/recognized truncation, duplicate
attempts and cumulative read budgets. Recovery appends only missing authorized
chunks without resetting time/attempt counts or replaying a worker. All three
host routes share the checker; instructions specify actual host output controls.

New live studies cannot claim workers without a frozen expected-input manifest.
Invalid prior live delivery stops further claims before spending on later trials.
Incomplete/unknown live inputs are excluded from blind grading; grading refuses
them, and export withholds the whole model comparison rather than cherry-picking
valid-looking outputs. Original outcomes, grades and consumed usage remain.
Legacy live delivery without evidence stays unknown; historical synthetic
arithmetic remains explicitly synthetic.

Ordinary-work feedback reuses existing test/review outcomes, requested/trusted
observed settings and normalized usage. It correlates task, host, policy and
worker scope, keeps cache/reasoning accounting, leaves unknown subscription
allowance/cost unknown, and proposes reviewed catalog updates only. Known forced
substitution suspends choices without replay; unknown served identity does not.
No helper launches models, adds graders, trains or automatically edits policy.

## Validation and limits

Offline suite: 178 tests on Python 3.12 and 3.14, including all-role/all-host
routing, distinct same-role tasks, approval/scope/lock/fallback behavior, installed
wizard operations, privacy, truncation/recovery, invalid study exclusion and
normal-work feedback. Native validator: 19 agents, no errors/warnings. Portable
validator, generated drift checks, skill validation and whitespace checks pass.

Isolated native Codex and Claude installs passed exact bytes/modes and removal of
the source fixture. The pinned skills 1.5.25 installer passed both skills
independently, for Codex and Claude, from explicit bundle paths and repository-root
discovery, with source removal. Installer receipts are local temporary artifacts;
no personal profile or credentials were copied into those fixtures.

Fake providers and supplied synthetic host records establish wiring and failure
handling, not live recommendation quality or savings. No paid workers, Jev calls,
qualification studies or regrading runs were started. Delivery checks validate
supplied host-visible evidence; they cannot authenticate a fabricated source
label or prove comprehension. Hosts lacking visibility report unknown. Codex
catalog setup uses the actual exposed roster; no universal model list is shipped.
2 changes: 1 addition & 1 deletion plugin/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "guildhall",
"version": "0.17.2",
"version": "0.17.3",
"description": "The Guildhall \u2014 a gathering place for adventurers. A TDD-ordered coding agent harness for Claude Code, tuned for Opus-tier orchestration (Opus 5 recommended seat). The /quest slash command runs Mordain the Guildmaster, who writes a durable plan file, then dispatches 18 specialist adventurers across three tiers: Opus (architecture-reviewer, security-reviewer, reliability-reviewer, migration-safety-reviewer), Sonnet (test-author, feature-implementer, ui-test-author, docs-writer, pr-author, prototype-builder, debug-investigator, observability-reviewer, performance-reviewer, ops-readiness-reviewer, accessibility-reviewer), and Haiku (refactorer, plugin-validator, fog-cartographer). Post-green reviews fan out in parallel \u2014 two always-on (security, docs) plus six gated production-readiness reviewers (observability, reliability, performance, ops-readiness, migration-safety, accessibility) that fire only when their trigger matches the diff. Rook (pr-author) closes the quest with a platform-agnostic PR draft and folds the runbook into the body. Integrates with IDD-framework specs.",
"author": {
"name": "GrillerGeek"
Expand Down
2 changes: 1 addition & 1 deletion plugin/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "guildhall",
"version": "0.17.2",
"version": "0.17.3",
"author": {
"name": "GrillerGeek"
},
Expand Down
2 changes: 1 addition & 1 deletion plugin/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ Full character sheets in [`CHARACTERS.md`](CHARACTERS.md).

## Installation

Version **0.17.2 includes the installed routing tools**. The [installation guide](https://github.com/GrillerGeek/guildhall/blob/main/docs/installation.md)
Version **0.17.3 includes the installed routing tools**. The [installation guide](https://github.com/GrillerGeek/guildhall/blob/main/docs/installation.md)
covers this repository's Codex, Claude and standalone routes, updates and removal.
The `main` route receives 0.16.0 after merge; the separate marketplace below is
not updated by this change.
Expand Down
6 changes: 6 additions & 0 deletions plugin/commands/quest.md
Original file line number Diff line number Diff line change
Expand Up @@ -405,3 +405,9 @@ The tone is a Guildmaster's fireside account, not a machine's log. Keep it truth
For schema-v3/v4 routing, read `${CLAUDE_PLUGIN_ROOT}/skills/guildhall-quest/references/host-evidence.md`. Run preflight before paid studies; compare reviewed observations after each worker and suspend on drift without replay.

Schema v4 supports all 18 specialist roles under explicit allowlists and scoped qualification. Follow the [role matrix and migration guide](../skills/guildhall-quest/references/role-eligibility.md); upgrading never enables a role automatically.


Use [bounded input delivery and feedback](../skills/guildhall-quest/references/input-delivery.md)
for required worker references on every path. Record host-visible completeness
or unknown, repair only missing authorized chunks within the existing budget,
and retain actual outcomes/usage in the plan without extra grading workers.
2 changes: 1 addition & 1 deletion plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "guildhall",
"version": "0.17.2",
"version": "0.17.3",
"author": {
"name": "GrillerGeek"
},
Expand Down
14 changes: 8 additions & 6 deletions plugin/portable/references/global-routing.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,8 +63,8 @@ Use an ordinary setup conversation outside quest execution. For example:
> Configure Guildhall specialist model routing globally for this host. Read the
> installed global and model-routing guides. Discover supported model/effort
> settings without paid probes. Prepare an off global policy, keep unknown metrics
> and qualifications unknown, and show the proposed shadow policy, all-project
> scope, objective, candidates, budgets, outbound fields and host evidence.
> and qualifications unknown, and show the proposed Dynamic policy, all-project
> scope, objective, candidates, budgets, outbound fields and worker controls.
> Activate only after I explicitly approve that exact proposal. Retain any
> existing project policy and explain which source currently wins.

Expand All @@ -83,16 +83,17 @@ different source do not copy consent. Credentials stay in the host environment,
using the policy's `key_env`; never store the secret value in either file.

Each approval binds the effective canonical policy hash, source and host route,
current host configuration fingerprint, and reviewed evidence hashes. Changing
current host configuration fingerprint, and either v5 reviewed control facts or
legacy/qualified reviewed evidence hashes. Changing
another host's global entry or JSON formatting does not invalidate it. Changing
the selected policy or host configuration requires renewed activation. Expiry,
revocation, missing reviewed host evidence or corrupt approval state prevents
revocation, missing required control/evidence review or corrupt approval state prevents
reuse. Approval expiry may be omitted; qualification still has its own expiry.
When reviewed host evidence has an expiry, do not approve beyond that expiry.

The helper fingerprints the supplied host request fields (sorting supported
settings), excluding the evidence hash itself; that hash must separately appear
in approved evidence. Supply fresh, truthful host metadata each time, not a stale
in approved evidence. V5 uses the control fingerprint described below instead. Supply fresh, truthful host metadata each time, not a stale
snapshot to keep approval working. Host/role qualification checks still happen
in the routing engine. `ready` means reusable consent, not adaptive qualification
or proof of which model executed.
Expand Down Expand Up @@ -165,7 +166,8 @@ recognize a valid opt-out or valid off policy, and skip both Python helpers.
Check global defaults when the project file is absent. If selected configuration
cannot be validated, stop configuration resolution; never guess that it is off.

For enabled routing, call `status` with current host evidence. Use its `policy`
For enabled routing, call `status` with current host controls/evidence shaped to
the selected request version. V5 dynamic requires no execution-evidence hashes. Use its `policy`
and `activation` verbatim in the existing routing request, setting request schema
version to the selected policy's schema version. Add a summary hash only after
the exact summary's approval. Non-ready results do not authorize external calls;
Expand Down
6 changes: 6 additions & 0 deletions plugin/portable/references/hosts.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,3 +105,9 @@ worker and bundled role body without assuming native registration or hooks.
Codex uses actual exposed model/effort arguments and a fresh independent context;
missing served identity alone is not a reason to demand a study. Never copy effort
names between hosts. Reuse the same control approval across role fallback choices.


For required reference material, use [bounded input delivery](input-delivery.md)
and record ordinary-work feedback from existing tests/reviews/usage only. Repair
missing chunks within the existing budget; never treat source hashes or worker
assertions as delivery proof or run extra workers to collect feedback.
104 changes: 104 additions & 0 deletions plugin/portable/references/input-delivery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
# Bounded input delivery and ordinary-work feedback

A correct source file and its hash do not prove that a worker received the whole
file. Treat delivery as a separate fact from model identity and output quality.
Use this procedure for required role/reference material on all three hosts and
for optional study inputs. It does not authorize another worker or paid grader.

## Freeze and read bounded material

1. Inventory only the role contract and references authorized by the handoff.
Identify required sections with stable material IDs. Test-author inventory
comes only from its Spec/API/test handoff, never implementation context.
Preserve the source revision and permitted read scope. Do not load unrelated
documents just to create an inventory.
2. `scripts/routing_delivery.py` accepts JSON stdin operations `plan`, `emit` and
`assess`. Plan takes `materials: [{id, text}]`, `max_read_attempts` and
`max_read_ms`; derive these limits from the task's remaining authorized budget,
not a new retry allowance. It returns chunk IDs, byte lengths and SHA256 values.
Store a large manifest in task-owned temporary storage; do not print it through
an undersized tool output. The helper reads no files and launches no models.
3. Emit takes the same authorized `materials` and one `chunk_id`. It returns only
that identified chunk: at most 512 Unicode characters / 2048 UTF-8 bytes. Read
one chunk at a time with enough actual tool output budget, including the JSON
wrapper. A prose request for a larger output limit does not set the limit.
4. On Codex, set the actual shell tool's `max_output_tokens` and the orchestration
tool's own output cap. When functions.exec is present, use a literal first-line
pragma, e.g. `// @exec: {"max_output_tokens": 3000}`, plus the nested command's
`max_output_tokens: 3000` for a single chunk. Do not combine a whole reference
bundle into that one capped call. On Claude native/skill routes, use the actual
Read offset/limit or bounded shell output supported by that worker tool; names
and controls vary, so do not copy Codex parameters into Claude.
5. Inspect the visible result for omissions, truncation markers and the complete
identified content. Source-side hashes alone and a worker saying “read it all”
are insufficient. If trusted host output is unavailable, record `unknown`.
Do not relabel source-file bytes as the observed tool output.

## Assess, recover and preserve limits

Assess takes `packet: {schema_version: 1, host, worker_id, manifest, observations}`.
Each observation has `chunk_id`, zero-based `attempt`, the actual visible `text`,
`truncated` (boolean or null), `source` (`host_tool_output`, `worker_assertion`,
`unavailable`), local evidence reference (or null), and `elapsed_ms` (or null).
Use `host_tool_output` only for host-visible evidence you actually inspected.
These are supplied records: the helper cannot authenticate a fabricated source
label or prove worker comprehension. Inspect the underlying evidence honestly.

The helper checks visible content/length/hash, duplicate attempts, truncation
flags/markers, omissions and cumulative read budgets. It returns complete,
incomplete, unknown or budget_exceeded, missing IDs and the recoverable subset.
Append only missing chunk reads within the remaining budget; retain all earlier
attempts. Unknown timing does not authorize budgeted recovery. A complete later
read can repair a truncated chunk, but does not erase consumed attempts/time.
A duplicate attempt is invalid input. Never automatically restart the worker.

For ordinary work, repair required missing inputs within scope or report the
input-delivery failure. Unknown completeness must be reported as unknown; it
is not proof of complete delivery, and does not itself reinstate a study gate
for Dynamic routing. Existing tests/review still determine the work's outcome.
Keep these receipts in the permitted quest plan (or final fast-lane report).

## Optional study integration

New study state uses version 2; existing manifests/state remain readable and
original outcomes/grades stay unchanged. New live trials cannot be claimed without a frozen delivery manifest, and an
invalid prior live trial stops further claims before more model usage. Before claiming a trial, freeze the
expected manifest with `study_runner.py delivery_plan --directory … --run …
--input manifest.json`. All trials for one fixture must use the same manifest.
After actual reads, append the cumulative assess packet with the `delivery`
operation. It must match that manifest, host and worker; earlier observations
cannot be removed. Record the ordinary outcome and actual consumption normally.

`blind` excludes incomplete/unknown live delivery before paying for grading.
`grade` refuses those trials. `export` returns `invalid_input_delivery` and no
comparison dataset if any trial is invalid, even if an old grade exists; it does
not cherry-pick a favorable subset or manufacture quality failures. Missing
legacy live delivery is unknown. Historical synthetic fixtures keep explicitly
synthetic arithmetic compatibility; this is not live delivery evidence.
Keep prior exported reports unchanged and annotate their limits separately.
No new study or regrading run is required for ordinary Dynamic activation.

## Feedback from work already performed

`scripts/routing_feedback.py` accepts the v5 request, its decision, a worker
outcome, observation and optional existing usage packet. Outcome fields are
worker_id, status (completed/failed/interrupted), tests
(passed/failed/unknown/not_applicable), review
(accepted/rejected/unknown/not_applicable), retries, local evidence references
and delivery status. Observation fields are source
(host_metadata/worker_assertion/unknown), evidence, requested_resolved and
observed model/effort pairs (nullable), and configuration_supported (nullable).

Reuse trusted host metadata, actual tests/reviews/retries and complete usage from
that work. Compare resolved identities only when known; requested aliases are not
served names. Worker assertions never become observed identity. Known forced
substitution/unsupported configuration suspends future choices without resetting
counters or replaying work. Unknown identity alone does not suspend Dynamic mode.

Optional usage input uses the existing routing_usage schema, correlated to this
worker and host/role/category. Cache/reasoning semantics and incomplete counts
remain intact. Raw tokens are a subscription proxy; account-wide percentages do
not establish per-task allowance or monetary cost. The result leaves subscription
allowance percent null, requests zero extra runs and proposes catalog review only.
It never trains, mutates a policy/catalog, gives qualification or writes files.
Do not schedule duplicate runs, independent grading or extra retries for feedback.
6 changes: 6 additions & 0 deletions plugin/portable/references/qualification-study.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Run a bounded routing study

This advanced workflow is optional. Dynamic routing does not require a study.
Before a paid trial, follow [bounded delivery](input-delivery.md), freeze its
expected input manifest and verify actual visible chunks. Exclude incomplete or
unknown live delivery before grading or comparing model quality/efficiency.
A failed study does not authorize another run or prevent Dynamic activation.

Guildhall supplies a development headroom analyzer and a study controller. They
run no models themselves. The controller prepares isolated Git worktrees and
explicit dispatch packets for the actual host, then imports measured outcomes
Expand Down
Loading
Loading