No-Op Test (Behavior-vs-Default Pruning)
Remove instructions only after testing their effect against the model's baseline behavior.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
Prune the instruction “cite evidence for every factual conclusion” because the model may already do it. First test its effect on current citation behavior rather than asking the model whether it matters.
Mechanism
Freeze model, harness, tasks, citation criteria and tolerated regression. Compare full prompt, only-this-sentence-removed and no-skill control with repeated matched runs. Track engagement separately from citation quality and recompute missing-citation counts from raw rows. Check exact-text dependencies in external parsers. Retain the rule with insufficient evidence; restore minimally after regression.
Bad example
The rule sounds obvious, so delete it. Ask the model if it needs it and treat no as a successful experiment.
Good example
Compare three prompt arms on the same evidence tasks with model version, repetitions, engagement and missing citations recorded. Remove only after meeting predefined quality gates for this sample and checking external dependencies. Retain under uncertainty; without actual runs do not claim no effect.
Why the change matters
Ablation tests behavioral contribution; the default control separates whole-skill from single-sentence effects. Observed answers rather than self-report make pruning reviewable. Separating engagement from benefit prevents loading markers from replacing quality.
Observable expectation
A teaching table can contain case, arm, rep, cited and engaged. Full and removed arms both citing 3/3 shows no difference in those observations, not proof of no effect. Missing citations after removal warrant attribution checks and minimal restoration. Numbers illustrate a hypothetical table only.
Limits
No-op conclusions depend on model, task and statistical detection power; small samples miss rare regression. Costs depend on more than prompt length, so fewer words do not establish savings. Paid experiments need authorization; revisit after model/harness changes.
Sources and evidence
- mattpocock/skills · Steps and completion criteria
File at this versiond81f3a183412 - mattpocock/skills · Hunt no-ops sentence by sentence
File at this versiond81f3a183412 - affaan-m/ECC · Internal behavior-change check
File at this versionef648e01899b - mattpocock/skills · No-Op Test (Behavior-vs-Default Pruning)
File at this versiond81f3a183412 - anthropics/skills · Behavioral removal hypothesis
File at this version8a1541c4a3ff - addyosmani/agent-skills · Separate activation indicators from outcome scores
File at this version9d0c60d406b4 - anthropics/skills · Disable/enable lever and recompute raw rows
File at this version8a1541c4a3ff - anthropics/skills · Engagement versus outcome
File at this version8a1541c4a3ff