P244 · Evaluation & feedback

Noise-aware intervention comparison

Compare one change under equivalent measurement conditions and require signal beyond normal variance.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A parser changes from one 100 ms run to 97 ms while ordinary runs span 95–105 ms. The teaching difference lies within variation, not established improvement.

Mechanism

Define correctness gate, metric, inputs, cache, hardware and search budget first. Change one hypothesis and repeat both versions under matched conditions/budgets, retaining raw values; interleave order to reduce drift where appropriate. Compare effect and uncertainty using a predefined method. Reject correctness failures; claim benefit only beyond explainable noise. Evaluate separate maintenance benefits under a separate contract.

Bad example

Measure the baseline cold once and the candidate warm once, call 3% success and omit failed cases.

Good example

Repeat old/new parsers with identical inputs and cache conditions, retaining correctness checks and per-run time. Against teaching 95–105 ms variation,100→97 alone is not attributable. Bound attempts and report no detected benefit when indistinguishable; evaluate maintenance benefit separately if intended.

Why the change matters

Cache and environment create differences unrelated to code. Matched conditions, single changes and repeats constrain confounding, while correctness prevents speed from dropping required work.

Observable expectation

Teaching 100→97 is only a candidate signal in 95–105 background. Repeated old≈100/new≈80 under matched conditions warrants analysis with sample/method recorded. Skipping validation fails correctness even if faster. Numbers illustrate criteria, not measured performance.

Limits

Finite samples cannot exclude all drift, and ranges are not automatically confidence intervals. Define statistical units/methods. Nonperformance benefit must not be called speed benefit. Apply keep/revert policies to approved objectives without discarding user changes autonomously.

Sources and evidence

Read the editorial criteria