Keep selection feedback out of revision prompts
Keep withheld-case results out of the prompt that proposes revisions, and name the resulting selection set accurately.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
A routing description is revised from training failures and reserved queries select winners each round. Even if the reviser never reads queries, repeated selection makes them validation/selection, not an independent final test.
Mechanism
Fix train, selection and final-confirmation identities and feedback permissions. Give revisers training material only, removing withheld history fields and inspecting full prompts/referenced files for leaks. Name selection honestly; hold final cases until selection ends. Once final feedback informs revision, it is no longer untouched. Label stored thresholds/scores historical and unreviewed gates provisional.
Bad example
Feed withheld failures into every rewrite while calling the best score an untouched test result.
Good example
Expose only training failures and filtered history to the reviser. Select candidates using a separate validation score and confirm on new cases unused in selection. Record set identities and feedback access; stored thresholds are not fresh evaluations.
Why the change matters
Isolation reduces direct leakage, but adaptive selection still consumes selection information. Correct naming prevents validation wins from becoming independent generalization claims.
Observable expectation
Teaching prompt snapshots exclude selection cases/score fields, and selection/final tables have IDs. Using final failures for another revision turns that set into development feedback requiring new confirmation. Inspect references/history, not just one filtered field.
Limits
Small samples may justify directional results rather than fabricated three-way splits. Repetitions are not automatically independent cases, and deterministic curated corpora do not yield population intervals by assertion. Frozen sources sometimes call selection test; explain actual use.
Sources and evidence
- anthropics/skills · Blind history and selection
File at this version8a1541c4a3ff - anthropics/skills · Generalize intent instead of enumerating cases
File at this version8a1541c4a3ff - anthropics/skills · Two-way selection versus optional untouched final test
File at this version8a1541c4a3ff - anthropics/skills · Selection is not holdout
File at this version8a1541c4a3ff - nextlevelbuilder/ui-ux-pro-max-skill · Explicit calibration/held_out case labels; partial illustration of channel separation
File at this version09170eec67ee - nextlevelbuilder/ui-ux-pro-max-skill · Declared tuning exclusion for held_out and provisional comparison boundary
File at this version09170eec67ee - nextlevelbuilder/ui-ux-pro-max-skill · Stored held_out summary, not a fresh evaluation result
File at this version09170eec67ee