Regrade all variants after a criterion change
When a metric or grader changes, apply it consistently to stored outputs from every compared variant before recommending a winner.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
A report benchmark graded structure but omitted factual fidelity. Teaching baseline B and candidate C have saved outputs. Under the new criterion, compare regraded results for both, not C’s new score against B’s old score.
Mechanism
Version rubric, grader and formulas while freezing each variant’s output identity. Apply the same new criterion to baseline and all candidates; preserve old grades with their revision. Rebuild item/aggregate tables, checking unavailable evidence and grading failures, and revise recommendations if rankings change. New generation is a separate experiment, not regrading stored artifacts.
Bad example
Grade only latest C for facts and compare with B’s old structure score to declare a winner.
Good example
Regrade saved B, C and other outputs with grader-v2, recording hashes, old/new criteria and evidence. Recommend afterward. Explain a ranking reversal as a different signal; variants without originals are unregradable, not assigned invented scores.
Why the change matters
Fixed artifacts and shared new measurement separate grading changes from generation changes. Recommendations use comparable evidence and historical corrections remain traceable.
Observable expectation
Teaching structure scores B=0.7/C=0.9 become fact scores B=0.8/C=0.6, potentially reversing recommendation. New scores compare; old C0.9 versus new B0.8 does not. These are hypothetical values with criterion columns retained.
Limits
Stored outputs may lack newly needed evidence, and model grading adds variance/cost. Regrading does not undo selection/tuning on the wrong signal; new validation or search may be needed. Experiments and budgets remain within authorization.