P361 · Evaluation & feedback

Retry Success vs Repeated Reliability

Distinguish succeeding once with retries from succeeding on every repeated trial.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Three runs of one teaching routing fixture yield pass, fail, fail. That satisfies at-least-one-within-three but not all-three success; operational questions differ.

Mechanism

Fix repetitions, retry rules, task and budget and retain each outcome/failure class. Compute any-success, all-success and observed success fraction separately, noting independent trials versus stateful sequential retries. Count failed resources too and use the preselected operational metric without deleting failures.

Bad example

Call one success in three three-run reliable and count only successful-call cost.

Good example

Record[pass,fail,fail]: within-three=true, all-three=false and 1/3 observed successes with all resources/failures. This teaching sequence is not a population probability; apply declared reliability gates.

Why the change matters

Any-success asks whether repeated opportunities can produce success; all-success asks stability. Separating them prevents retry wins from becoming every-time reliability claims.

Observable expectation

Teaching all-pass makes both true; all-fail makes both false; the original has one true/one false. Recompute from complete arrays and classify infrastructure failures under the contract.

Limits

Small/correlated trials do not establish universal reliability. pass@k/pass^k can have statistical definitions elsewhere; declare these as observed events. Source 90% targets are not universal.

Sources and evidence

Read the editorial criteria