P257 · Evaluation & feedback

Separate build variance from scoring variance

When a prompt builds an artifact that is later evaluated, measure build variability separately from repeated scoring of one artifact.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A prompt builds a retrieval index that queries later score. Stable repeated scores for one index do not establish stable generation; randomness occurs at both build and scoring levels.

Mechanism

Independently rebuild unchanged baselines, storing build_id, prompt, inputs, model and generation traces. Score each with the same queries/repetitions, separating between-build and within-build variation. Use matching design for candidates and retain grouped raw values. Many scores on one build are not many independent builds. Analyze uncertainty/budget at actual random levels.

Bad example

Build one baseline and rescore it dozens of times, using a narrow interval to claim a better index-building prompt.

Good example

Build independent B1/B2/B3 indexes from the same prompt and score each under identical query design; rebuild candidates too. Retain build identity and within-group scores and report both variation levels. More scores on one index cannot establish stable construction.

Why the change matters

Rescoring reveals fixed-artifact grading noise only. Independent rebuilding exposes generation variation, preventing a lucky build from masquerading as a prompt gain.

Observable expectation

Teaching B1 scores 0.79/0.80, B2 scores 0.65/0.66 and B3 scores 0.78/0.79: small within-group but large between-build variation. Do not use B1 alone as the benefit threshold. Label values hypothetical and verify distinct builds rather than reused caches.

Limits

Shared inputs/caches/random state can couple builds. Replication/statistical design depends on hierarchy and cost; source counts guarantee nothing. If generation is deterministic, verify that premise rather than mechanically requiring repeated model generation.

Sources and evidence

Read the editorial criteria