You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Hidden skills-expert bot: build skills and benchmark whether each one adds value #1504
As discussed in the DM: skills are getting a global library, per-turn selection (#1502), and candidate capture from completed tasks (#1365), but nothing measures whether a skill actually helps. As libraries grow, users will accumulate skills that cost tokens every turn without improving outcomes, and today there is no way to tell a useful skill from a neutral or harmful one. The agreed direction was a hidden skills expert bot, like the Claude skill building skill, that benchmarks whether a skill actually adds value.
Proposal
A hidden skills-expert bot inspired by Anthropic's skill-creator (github.com/anthropics/skills/tree/main/skills/skill-creator), adapted to OMB's architecture:
Measure loop: every test prompt runs twice, once with the skill and once without it (or against a snapshot of the old version when improving). Where Claude needs subagents, OMB dispatches the runs to real bots in a room, which is the native primitive. Assertions are graded per run, script-first for anything programmatically checkable.
Alpha report: pass rate, time, and tokens with mean plus standard deviation and the with-skill versus baseline delta, per assertion. The analyzer flags non-discriminating assertions (pass with and without the skill) and high-variance cases, so a skill that wins on pass rate but doubles token cost does not look free. Verdict per skill: keep, revise, or drop.
Human review before revision: an in-app panel shows outputs and grades side by side with the previous iteration, per-case feedback, empty feedback counts as approved. This fits the simpler skills UI direction already on the priorities list.
Trigger evals: 20 realistic queries split into should-trigger and near-miss should-not-trigger sets with a train and held-out split. This doubles as the validation suite for the per-turn router in Route only the relevant skills into each turn using embeddings (per-turn skill selection) #1502: the router should pick the right skill on the positive set and stay quiet on the near-misses.
The bot stays hidden from the roster, invoked from the skills UI.
Sequencing: happy to hold implementation until the code-health stack (#1402 to #1498) merges, same as #1502. The grader and mock-provider pieces can be shared with the behavior evals proposal in #1503. Related: #1502, #1365.
Alternatives considered
No benchmarking: the status quo, libraries rot as they grow and every skill costs tokens forever.
External eval tooling: rejected, OMB is local-first and this needs to run inside the app against the user's own bots.
Judge-only scoring without baselines: gives opinions, not an alpha number, and cannot detect a skill that changes nothing.
Area
Harness server / API (server/)
Problem
As discussed in the DM: skills are getting a global library, per-turn selection (#1502), and candidate capture from completed tasks (#1365), but nothing measures whether a skill actually helps. As libraries grow, users will accumulate skills that cost tokens every turn without improving outcomes, and today there is no way to tell a useful skill from a neutral or harmful one. The agreed direction was a hidden skills expert bot, like the Claude skill building skill, that benchmarks whether a skill actually adds value.
Proposal
A hidden skills-expert bot inspired by Anthropic's skill-creator (github.com/anthropics/skills/tree/main/skills/skill-creator), adapted to OMB's architecture:
Sequencing: happy to hold implementation until the code-health stack (#1402 to #1498) merges, same as #1502. The grader and mock-provider pieces can be shared with the behavior evals proposal in #1503. Related: #1502, #1365.
Alternatives considered