P267 · Evaluation & feedback

Freeze comparison references across rounds

Compare iterative candidates with the same saved reference artifacts and retain older references when adding a harder comparison.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Iterate a report prompt using per-case blind win rates. Regenerating opponents each round changes what v1/v2 face. Teaching reference R0 should be saved once and fixed.

Mechanism

Save case-specific R0 outputs/hashes and version cases, rubric and environment. Compare every candidate blindly with the same R0, retaining labels/evidence. If saturated, add stronger R1 with a new metric column while preserving R0. Do not rewrite historical denominators. Fresh no-change system controls may address drift separately, but must not replace this fixed reference.

Bad example

Generate a new baseline each round yet compare win rates directly. Replace the original column when introducing a harder reference.

Good example

Freeze R0 per case and compare v1/v2 to those same files. Add R1 as a separate column when needed, retaining R0 identity/metric. Record judge/case revisions and ties; drift controls do not replace the fixed opponent.

Why the change matters

A fixed opponent preserves the meaning of across-round win rates. Adding rather than replacing references retains history and provides a more discriminating future comparator.

Observable expectation

Teaching v1/v2 compare against the same three R0 cases. Changed R0 hashes invalidate or require rebuilding comparisons. Adding R1 yields vs-R0 and vs-R1 columns rather than calling new-reference scores growth in the old metric. Keep item mappings/artifacts accessible.

Limits

References become stale and judges/environments drift; fixed comparisons do not establish broad superiority. Item grading can miss cross-case style collapse, needing set-level checks. Fixed references and rerun unchanged systems have different purposes that must be stated.

Sources and evidence

Read the editorial criteria