P52 · Evaluation & feedback

Rubric-Based Model Evaluation

Use a model evaluator with a fixed rubric and checkable evidence.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Compare incident summaries for factual fidelity. Teaching input says a test failed at 12:04, with production impact and cause unknown. A preserves uncertainty; B claims a database failure affected every customer. Fluency cannot justify new facts.

Mechanism

Freeze the same input, sample and rubric before grading: time matches, impact is supported, cause is not invented. Return pass/fail/unknown and evidence sentences per criterion. Verify objective fields with code/source before accepting semantic model judgment. Hide irrelevant model labels, allow uncertainty/ties, calibrate on human-labeled examples and inspect extra claims beyond the rubric.

Bad example

Pick the more polished summary and give model B a high score without checking the incident source.

Good example

Evaluate A and B under the same three criteria, quoting time, impact and cause statements. Source says production and cause are unknown, so B’s database/every-customer additions fail. Return item reasons and unresolved checks rather than treating prose quality as factual accuracy.

Why the change matters

A fixed rubric makes criteria visible and evidence sentences contestable. Objective checks combined with semantic review reduce rewarding unsupported facts because they sound fluent.

Observable expectation

Teaching A can pass these criteria if it preserves all three facts; B fails impact and cause. Changing A’s date causes a time failure; missing source makes relevant checks unknown. Compare grades with human labels and record disagreement. This is not an actual model ranking.

Limits

Graders may favor their own style, longer outputs or labels; calibration and counterexamples matter. Fixed samples do not represent all tasks, and rubric success does not exclude other errors. The frozen source’s PASS/FAIL protocol does not require dropping uncertainty in this example.

Sources and evidence

Read the editorial criteria

Related methods