P108 · Evaluation & feedback

Blind A/B Comparison

Compare outputs on matched tasks while hiding version identity from the evaluator.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Compare two incident-summary prompts using the same teaching input: time 12:04 and unknown cause. One output retains uncertainty and another invents a cause. The evaluator must not know which is newer.

Mechanism

Freeze cases, rubric, run conditions and raw outputs. Generate both versions and use randomized anonymous labels with balanced order, keeping the mapping hidden. Grade factual fidelity and task completion, allowing ties or inability to judge. Save decisions before decoding. Repeat cases and report item results, failed runs and variation without selecting only wins.

Bad example

Tell the evaluator A is improved and ask whether it is better. Retain only judgments supporting A.

Good example

Compare both outputs on the same frozen incident with hidden version identity and randomized order. Cite time/cause evidence under one rubric, allow ties and save decisions before decoding. Record failures and repeated variation; newer is not quality evidence.

Why the change matters

Matched inputs reduce task differences; hidden labels reduce expectations that a revision must improve. Saving before decoding prevents identity-driven edits, and repeats expose chance wins.

Observable expectation

A teaching evaluator should prefer preserved uncertainty without referencing revision identity. Swapping order should retain evidence-based reasoning; indistinguishable outputs tie. Retain cases, outputs, decisions and mapping. This expectation is not a measured win rate.

Limits

Blinding does not remove style, position or grader bias. Small samples do not establish general benefit. Paid runs and external sharing still require authorization; the source’s preference for rare ties should not force invented differences.

Sources and evidence

Read the editorial criteria

Related methods