Skip to content

Hidden skills-expert bot: build skills and benchmark whether each one adds value #1504

Description

@bradhallett

Area

Harness server / API (server/)

Problem

As discussed in the DM: skills are getting a global library, per-turn selection (#1502), and candidate capture from completed tasks (#1365), but nothing measures whether a skill actually helps. As libraries grow, users will accumulate skills that cost tokens every turn without improving outcomes, and today there is no way to tell a useful skill from a neutral or harmful one. The agreed direction was a hidden skills expert bot, like the Claude skill building skill, that benchmarks whether a skill actually adds value.

Proposal

A hidden skills-expert bot inspired by Anthropic's skill-creator (github.com/anthropics/skills/tree/main/skills/skill-creator), adapted to OMB's architecture:

  1. Build loop: capture intent from the conversation (including turning a completed workflow into a skill, ties into feat(skills): reflect a judged-complete board task into a candidate skill for review (phase 4 part 3) #1365), interview for edge cases, draft the SKILL.md with progressive disclosure (metadata always visible, body on trigger, resources on demand).
  2. Measure loop: every test prompt runs twice, once with the skill and once without it (or against a snapshot of the old version when improving). Where Claude needs subagents, OMB dispatches the runs to real bots in a room, which is the native primitive. Assertions are graded per run, script-first for anything programmatically checkable.
  3. Alpha report: pass rate, time, and tokens with mean plus standard deviation and the with-skill versus baseline delta, per assertion. The analyzer flags non-discriminating assertions (pass with and without the skill) and high-variance cases, so a skill that wins on pass rate but doubles token cost does not look free. Verdict per skill: keep, revise, or drop.
  4. Human review before revision: an in-app panel shows outputs and grades side by side with the previous iteration, per-case feedback, empty feedback counts as approved. This fits the simpler skills UI direction already on the priorities list.
  5. Trigger evals: 20 realistic queries split into should-trigger and near-miss should-not-trigger sets with a train and held-out split. This doubles as the validation suite for the per-turn router in Route only the relevant skills into each turn using embeddings (per-turn skill selection) #1502: the router should pick the right skill on the positive set and stay quiet on the near-misses.
  6. The bot stays hidden from the roster, invoked from the skills UI.

Sequencing: happy to hold implementation until the code-health stack (#1402 to #1498) merges, same as #1502. The grader and mock-provider pieces can be shared with the behavior evals proposal in #1503. Related: #1502, #1365.

Alternatives considered

  • No benchmarking: the status quo, libraries rot as they grow and every skill costs tokens forever.
  • External eval tooling: rejected, OMB is local-first and this needs to run inside the app against the user's own bots.
  • Judge-only scoring without baselines: gives opinions, not an alpha number, and cannot detect a skill that changes nothing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions