P269 · Evaluation & feedback

Record the origin of evaluation references

Record whether reference answers are independently verified or generated by a system under comparison.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

When comparing summarizers, the answer key comes from the incumbent. Teaching source says unknown cause, while its summary invents database failure. Exact matching would reward the error.

Mechanism

Record reference origin, generating system/version, source input and independent-verification status. Recheck facts against source rather than calling compared-model outputs gold. Use rubrics/blind comparison for multiple valid phrasings, retaining incumbent outputs as attributed references. Test graders with good/bad controls and preserve annotation disagreement.

Bad example

Mark different wording wrong and trust the incumbent-generated database cause without verification.

Good example

Label the old summary incumbent-generated and not independently verified. Check both against source unknown cause rather than golding the database assertion; allow equivalent language. Retain the old artifact for comparison and independently inspect factual criteria.

Why the change matters

Origin determines what matching measures. Model-derived references can reward imitation; independent facts separate style from correctness.

Observable expectation

The teaching incumbent’s invented database cause fails fidelity; a new unknown-cause summary can pass despite low text match. Reference records expose origin/version/evidence. Pending verification is not a fabricated human-reviewed label.

Limits

Human verification can fail too. Correct references do not establish a well-calibrated grader, and saved model outputs are not independent answers. Use/retention of production materials follows authorization and data contracts.

Sources and evidence

Read the editorial criteria