P119 · Skill authoring

Eval-Driven Skill Improvement Loop

Improve a skill against preserved tasks, baselines and observable assertions.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Shorten a review skill while preserving triggers/locations. Teaching complete, missing-evidence and ambiguous cases compare old/candidate behavior rather than one satisfying output.

Mechanism

Snapshot old skill/resources/cases/criteria and fix harness/budget. Run matching designs, inspect outputs/traces and repair failure mechanisms minimally instead of memorized case answers. Recheck affected/regression cases and retain versions/repetitions/errors/gaps for new-case confirmation. Without execution this remains a plan.

Bad example

Rewrite elegantly, try one normal request and claim improvement without baseline/failures.

Good example

Freeze old version/three case types and check finding trigger/location/evidence. Compare matched runs, inspect missing-field failures and revise, preserving raw outputs. Report brevity and quality separately; model calls follow authorized budgets.

Why the change matters

Preserved cases give common reference and failures identify mechanisms. Process/artifact inspection distinguishes nonactivation, tool failure and answer errors.

Observable expectation

Teaching candidate missing locations regresses even if shorter. Repair then verify same cases/new paraphrases; missing evidence must not become invented findings. This design is not measured efficacy.

Limits

Limited cases overfit; independent confirmation/grade calibration matter. Model/harness changes invalidate results; count failures/judge costs. Single success/tool engagement proves no universal benefit, and frozen parallel instructions authorize no current delegation.

Sources and evidence

Read the editorial criteria

Related methods