P192 · Evaluation & feedback

No-Op Test (Behavior-vs-Default Pruning)

Remove instructions only after testing their effect against the model's baseline behavior.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Prune the instruction “cite evidence for every factual conclusion” because the model may already do it. First test its effect on current citation behavior rather than asking the model whether it matters.

Mechanism

Freeze model, harness, tasks, citation criteria and tolerated regression. Compare full prompt, only-this-sentence-removed and no-skill control with repeated matched runs. Track engagement separately from citation quality and recompute missing-citation counts from raw rows. Check exact-text dependencies in external parsers. Retain the rule with insufficient evidence; restore minimally after regression.

Bad example

The rule sounds obvious, so delete it. Ask the model if it needs it and treat no as a successful experiment.

Good example

Compare three prompt arms on the same evidence tasks with model version, repetitions, engagement and missing citations recorded. Remove only after meeting predefined quality gates for this sample and checking external dependencies. Retain under uncertainty; without actual runs do not claim no effect.

Why the change matters

Ablation tests behavioral contribution; the default control separates whole-skill from single-sentence effects. Observed answers rather than self-report make pruning reviewable. Separating engagement from benefit prevents loading markers from replacing quality.

Observable expectation

A teaching table can contain case, arm, rep, cited and engaged. Full and removed arms both citing 3/3 shows no difference in those observations, not proof of no effect. Missing citations after removal warrant attribution checks and minimal restoration. Numbers illustrate a hypothetical table only.

Limits

No-op conclusions depend on model, task and statistical detection power; small samples miss rare regression. Costs depend on more than prompt length, so fewer words do not establish savings. Paid experiments need authorization; revisit after model/harness changes.

Sources and evidence

Read the editorial criteria

Related methods