Eval-Driven Skill Improvement Loop
Improve a skill against preserved tasks, baselines and observable assertions.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
Shorten a review skill while preserving triggers/locations. Teaching complete, missing-evidence and ambiguous cases compare old/candidate behavior rather than one satisfying output.
Mechanism
Snapshot old skill/resources/cases/criteria and fix harness/budget. Run matching designs, inspect outputs/traces and repair failure mechanisms minimally instead of memorized case answers. Recheck affected/regression cases and retain versions/repetitions/errors/gaps for new-case confirmation. Without execution this remains a plan.
Bad example
Rewrite elegantly, try one normal request and claim improvement without baseline/failures.
Good example
Freeze old version/three case types and check finding trigger/location/evidence. Compare matched runs, inspect missing-field failures and revise, preserving raw outputs. Report brevity and quality separately; model calls follow authorized budgets.
Why the change matters
Preserved cases give common reference and failures identify mechanisms. Process/artifact inspection distinguishes nonactivation, tool failure and answer errors.
Observable expectation
Teaching candidate missing locations regresses even if shorter. Repair then verify same cases/new paraphrases; missing evidence must not become invented findings. This design is not measured efficacy.
Limits
Limited cases overfit; independent confirmation/grade calibration matter. Model/harness changes invalidate results; count failures/judge costs. Single success/tool engagement proves no universal benefit, and frozen parallel instructions authorize no current delegation.
Sources and evidence
- anthropics/skills · Step 1: Spawn all runs
File at this version8a1541c4a3ff - ComposioHQ/awesome-claude-skills · Review report / Low Accuracy
File at this versionbe2a406907db - ComposioHQ/awesome-claude-skills · Evaluation prompt
File at this versionbe2a406907db - affaan-m/ECC · Philosophy / Eval Types / Grader Types
File at this versionef648e01899b - anthropics/skills · Same-task baseline and objective assertions
File at this version8a1541c4a3ff - anthropics/skills · Generalization and repeated-helper evidence
File at this version8a1541c4a3ff