Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Changelog

## 0.17.1 — Unreleased

- Supply Jev with controlled task briefs and distinct reviewed profile preferences.
- Pin host-scoped catalogs; preserve unknown measurements and exclude unsupported profiles.
- Keep model identities, arbitrary catalog prose and provenance local.

## 0.17.0 — Unreleased

- Add schema-v5 dynamic routing with control-based approval and no study requirement.
Expand Down
2 changes: 1 addition & 1 deletion docs/contributing-agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,7 @@ workers or claim synthetic savings as observed results. See the

## Distribution and versioning

The current package version is 0.17.0; increment all three manifests for further
The current package version is 0.17.1; increment all three manifests for further
installer-visible changes. Existing native Claude agents/commands/hooks remain
preserved. Never rewrite historical plans to claim newer evidence. No repository
change implicitly installs personally, publishes a release or edits the separate
Expand Down
16 changes: 16 additions & 0 deletions docs/reviews/2026-09-27-task-routing-inputs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Task-level routing inputs

Jev receives a controlled assignment brief and distinct profile preferences,
with provenance and exact model identities retained locally. Catalog snapshots
are bound to policy approval and compiled against current worker controls. Null
measurements stay valid; partial observations never become complete metrics.
Unknown context cannot meet an enforced context requirement. Claude starter
aliases use documented qualitative descriptions; Codex discovers its roster from
the active host, not a universal hard-coded model list.

Offline fake-provider tests inspect the actual payload and demonstrate different
same-role decisions for routine/intensive tasks on all three hosts. They verify
privacy, injection rejection, catalog drift, unsupported settings, unknown
capacity and measurement completeness. These tests establish plumbing, not live
recommendation quality or token savings. Native and portable validation pass;
no live models or personal configuration were used.
2 changes: 1 addition & 1 deletion plugin/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "guildhall",
"version": "0.17.0",
"version": "0.17.1",
"description": "The Guildhall \u2014 a gathering place for adventurers. A TDD-ordered coding agent harness for Claude Code, tuned for Opus-tier orchestration (Opus 5 recommended seat). The /quest slash command runs Mordain the Guildmaster, who writes a durable plan file, then dispatches 18 specialist adventurers across three tiers: Opus (architecture-reviewer, security-reviewer, reliability-reviewer, migration-safety-reviewer), Sonnet (test-author, feature-implementer, ui-test-author, docs-writer, pr-author, prototype-builder, debug-investigator, observability-reviewer, performance-reviewer, ops-readiness-reviewer, accessibility-reviewer), and Haiku (refactorer, plugin-validator, fog-cartographer). Post-green reviews fan out in parallel \u2014 two always-on (security, docs) plus six gated production-readiness reviewers (observability, reliability, performance, ops-readiness, migration-safety, accessibility) that fire only when their trigger matches the diff. Rook (pr-author) closes the quest with a platform-agnostic PR draft and folds the runbook into the body. Integrates with IDD-framework specs.",
"author": {
"name": "GrillerGeek"
Expand Down
2 changes: 1 addition & 1 deletion plugin/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "guildhall",
"version": "0.17.0",
"version": "0.17.1",
"author": {
"name": "GrillerGeek"
},
Expand Down
2 changes: 1 addition & 1 deletion plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "guildhall",
"version": "0.17.0",
"version": "0.17.1",
"author": {
"name": "GrillerGeek"
},
Expand Down
71 changes: 71 additions & 0 deletions plugin/portable/references/task-routing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Task-level routing inputs (schema v5)

Dynamic routing is approved selection, not benchmark-proven improvement. Use
`routing_catalog.py` to compile a reviewed catalog and actual host controls on
JSON stdin: `{"catalog": <catalog-v1>, "host": <request-v5 host>}`. It returns
supported candidates, excluded IDs and their `catalog_revision`; it does not
write files, activate routing or call models. Set the policy's candidates and
revision to this snapshot. Changes require a new policy preview and approval.

## Catalogs and provenance

The [catalog schema](../resources/schemas/catalog-v1.schema.json) uses complete
host-scoped candidate records. Start with the matching resource:
[Codex](../resources/catalogs/codex-skill.json),
[native Claude](../resources/catalogs/claude-native.json), or
[standalone Claude](../resources/catalogs/claude-skill.json).
Codex's starter is deliberately empty: populate exact model/effort pairs and
permitted values from the currently callable worker tool and current host metadata.
Do not copy the setup author's personal model roster into another installation.
No paid discovery or another installed CLI is necessary or authoritative.

Claude's aliases are descriptive starting points. Their routine/extended/intensive
and efficiency descriptors are qualitative priors based on the official
[model configuration documentation](https://code.claude.com/docs/en/model-config),
reviewed 2026-09-27. Verify actual supported aliases, forced settings and provider
mappings on this host. These descriptions are not measured cost, capacity or
quality. Context and capabilities remain unknown until supported by host metadata
or documentation. Native frontmatter defaults stay unchanged. Standalone Claude
must discover its own tool controls. Keep Fable excluded.

Each `routing_profile` has controlled work types, reasoning depth, complexity,
risk and efficiency preferences plus local source/revision. Use `documented`
only when the linked source supports the description; use `user_preference` for
reviewed judgments. No universal model ranking ships. Ask for the missing preference
when metadata gives no meaningful distinction; do not label all profiles the same
and promise useful routing. Presets do not authorize any models or roles.

`facts_source` records where capacity/capabilities came from. Unknown facts remain
null/empty and cannot satisfy an enforced context/capability constraint. A task
with no established hard context minimum can use `context_bucket: unknown`;
never change a known requirement to unknown just to make a candidate eligible.
Unsupported settings are excluded. Recompile and obtain review when actual
controls or the catalog changes. Catalog updates do not silently widen policy.

Measurements remain optional and role/category-scoped. Schema v5 adds source,
sample_count, completeness and revision; partial observations cannot satisfy
numeric ceilings or become provider metrics. Unknown cost/quota remains unknown.
Raw subscription tokens do not establish weighted allowance or monetary savings.

## Task brief and privacy

From the worker's permitted handoff, set role/category, ambiguity, risk, context
bucket, reasoning depth, change breadth, expected output and verification needs.
Each field has a controlled vocabulary in the
[request schema](../resources/schemas/request-v5.schema.json); use `unknown` where
needed. Test-author facts come only from its allowed Spec/API/test inputs. Do not
read an implementation to classify that assignment. No extra classifier model
call is needed. The same specialist may receive very different briefs.

The `categories-v2` outbound contract sends these fields, objective, capability
count, known capacity, complete scoped measurements and each profile's controlled
preferences/basis. Each candidate gets different descriptive criteria when its
reviewed preferences differ. Only request-local `p0`, `p1`, … labels leave the
host; model names, catalog IDs, provenance URLs, paths, revisions and evidence
hashes stay local. The router's own model selector is necessarily sent to Jev.
Optional summary mode still needs approval of the exact text for each task.

The TypeSafe [choice API](https://docs.typesafe.ai/introduction/quickstart) supports
text state and per-choice criteria (reviewed 2026-09-27). Returned confidence is
recorded, not treated as calibrated coding success. Receipts retain local task
facts and catalog revision; do not invent a provider rationale or savings claim.
183 changes: 183 additions & 0 deletions plugin/portable/resources/catalogs/claude-native.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,183 @@
{
"schema_version": 1,
"host": "claude-native",
"profiles": [
{
"id": "haiku",
"host": "claude-native",
"model": "haiku",
"effort": null,
"roles": [
"accessibility-reviewer",
"architecture-reviewer",
"debug-investigator",
"docs-writer",
"feature-implementer",
"fog-cartographer",
"migration-safety-reviewer",
"observability-reviewer",
"ops-readiness-reviewer",
"performance-reviewer",
"plugin-validator",
"pr-author",
"prototype-builder",
"refactorer",
"reliability-reviewer",
"security-reviewer",
"test-author",
"ui-test-author"
],
"categories": [
"docs",
"pr",
"implementation",
"tests",
"debug",
"review",
"prototype",
"refactor"
],
"capabilities": [],
"context_tokens": null,
"quality": null,
"latency_ms": null,
"cost_usd": null,
"usage_tokens": null,
"profile_revision": "starter-2026-09-27",
"qualification": null,
"measurements": [],
"facts_source": {
"kind": "unknown",
"reference": null
},
"routing_profile": {
"work_types": [],
"reasoning_depth": "routine",
"complexity": [],
"risk": [],
"efficiency": "low_overhead",
"basis": "documented",
"source": "https://code.claude.com/docs/en/model-config",
"revision": "2026-09-27"
}
},
{
"id": "sonnet",
"host": "claude-native",
"model": "sonnet",
"effort": null,
"roles": [
"accessibility-reviewer",
"architecture-reviewer",
"debug-investigator",
"docs-writer",
"feature-implementer",
"fog-cartographer",
"migration-safety-reviewer",
"observability-reviewer",
"ops-readiness-reviewer",
"performance-reviewer",
"plugin-validator",
"pr-author",
"prototype-builder",
"refactorer",
"reliability-reviewer",
"security-reviewer",
"test-author",
"ui-test-author"
],
"categories": [
"docs",
"pr",
"implementation",
"tests",
"debug",
"review",
"prototype",
"refactor"
],
"capabilities": [],
"context_tokens": null,
"quality": null,
"latency_ms": null,
"cost_usd": null,
"usage_tokens": null,
"profile_revision": "starter-2026-09-27",
"qualification": null,
"measurements": [],
"facts_source": {
"kind": "unknown",
"reference": null
},
"routing_profile": {
"work_types": [],
"reasoning_depth": "extended",
"complexity": [],
"risk": [],
"efficiency": "balanced",
"basis": "documented",
"source": "https://code.claude.com/docs/en/model-config",
"revision": "2026-09-27"
}
},
{
"id": "opus",
"host": "claude-native",
"model": "opus",
"effort": null,
"roles": [
"accessibility-reviewer",
"architecture-reviewer",
"debug-investigator",
"docs-writer",
"feature-implementer",
"fog-cartographer",
"migration-safety-reviewer",
"observability-reviewer",
"ops-readiness-reviewer",
"performance-reviewer",
"plugin-validator",
"pr-author",
"prototype-builder",
"refactorer",
"reliability-reviewer",
"security-reviewer",
"test-author",
"ui-test-author"
],
"categories": [
"docs",
"pr",
"implementation",
"tests",
"debug",
"review",
"prototype",
"refactor"
],
"capabilities": [],
"context_tokens": null,
"quality": null,
"latency_ms": null,
"cost_usd": null,
"usage_tokens": null,
"profile_revision": "starter-2026-09-27",
"qualification": null,
"measurements": [],
"facts_source": {
"kind": "unknown",
"reference": null
},
"routing_profile": {
"work_types": [],
"reasoning_depth": "intensive",
"complexity": [],
"risk": [],
"efficiency": "thorough",
"basis": "documented",
"source": "https://code.claude.com/docs/en/model-config",
"revision": "2026-09-27"
}
}
]
}
Loading
Loading