P237 · Evaluation & feedback

Bind grader results to the full expectation set

Validate grade identities and reconstruct totals from the declared expectation set.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A fix case declares five expectations, but a grader returns four rows while claiming all five pass. A fluent summary cannot replace checking the canonical expectation set and item results.

Mechanism

Freeze case identity, expectation revision and IDs1–5. Parse grades and require each ID exactly once, in range, with boolean passed and evidence. Restore canonical text by ID instead of accepting rewritten requirements. Reject missing/duplicate rows and inconsistent integer summaries. Derive pass/fail/total/rate from valid rows, then independently assess evidence truth.

Bad example

Four rows look good; summary says 5/5, so report 100% without noticing absent ID5.

Good example

Validate one boolean result for each ID1–5 first. Missing ID5 invalidates complete grading. Recompute summaries from rows, reject integer conflicts and use derived rate. Bind the report to canonical expectations and check evidence rather than letting the grader omit difficult items.

Why the change matters

Canonical sets define denominators, unique identity defines coverage and recomputation defines arithmetic. Together they prevent invented coverage in polished summaries; evidence semantics remain another gate.

Observable expectation

Teaching IDs1–4 passing and 5 failing yield passed=4, failed=1, total=5, rate=0.8. Missing 5, duplicate 4 or total=4 is invalid. A rate-only mistake of 0.9 may be recomputed to 0.8 under the contract, but cannot repair absent rows.

Limits

Identity/arithmetic validation neither establishes grade truth nor calibrates a model. Empty expectation sets or new verdict states need contracts. Grader failure differs from model failure, and multi-case/repetition aggregation needs the correct statistical unit.

Sources and evidence

Read the editorial criteria