P357 · Evaluation & feedback

Compare Candidate and Baseline on the Same Trial Panel

Use the same declared trials for candidate and baseline and retain each paired result before applying an aggregate gate.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Trials 1/2 give candidate 0.9/0.5 and baseline 0.8/0.6. Best-versus-worst cherry-picking creates 0.3 while paired mean difference is 0.

Mechanism

Predeclare unique trial IDs, revisions, inputs, environment, grader and aggregate gates. Save both results/evidence per ID under matching tasks/data, preserving errors/gaps. Compute paired means, wins/ties and count from the complete panel, checking duplicate IDs and score ranges. Apply fixed gates with uncertainty/pilot scope.

Bad example

Compare 0.9 with 0.6, claim 0.3 over two trials and hide the other pair.

Good example

Retain 1:0.9 versus 0.8 and 2:0.5 versus 0.6. Both means are 0.7, delta 0 and wins 1/2, failing a predeclared positive-gain gate. Report missing/errors and raw evidence instead of mixing trials.

Why the change matters

Pairing reduces case-selection confounding, and whole-panel retention prevents favorable-row selection. Fixed aggregation stops post-hoc extremes changing meaning.

Observable expectation

Teaching deltas+0.1/−0.1 average 0 with win rate 0.5. Duplicate ID1 or an unpaired second result is incomplete; finite/range checks apply. Hypothetical values establish neither significance nor production benefit.

Limits

Same seeds do not ensure identical conditions and may not provide determinism. Two trials are small; pilots are not population statistics. Define failure handling/denominators before seeing results rather than deleting unfavorable calls.

Sources and evidence

Read the editorial criteria