P256 · Evaluation & feedback

Keep attempts, grades and traces coherent

Distinguish a scorable result from an attempt failure and keep each result’s grade, trace and usage bound to the same invocation.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Attempt a1 for case-c1 repetition 1 times out; after recovery a2 returns a scorable answer. Treating a1 as a negative result blocks resume, while pairing a2’s grade with a1’s trace mixes evidence.

Mechanism

Use (case,rep) for result slots and attempt_id for invocations. Durably commit scorable rows with same-call output, grade, trace, model and usage. Append infrastructure failures without scorable output to attempts while leaving result gaps. Define retries, refusal, truncation and grader-failure semantics and account all usage. Resume missing slots without replacing valid results with another call. A post-row trace failure must not duplicate billing records.

Bad example

Store a1 timeout as false in the result slot; attach a1 logs to successful a2 and count only successful-call cost.

Good example

Record a1 timeout class and available usage without inventing an answer. If a2 is allowed by retry policy, bind its result, grade, trace and usage to a2, preserving a1 history. Resume idempotently at case-c1/rep1 and report failure rates alongside conditional grades.

Why the change matters

Infrastructure no-answer differs from a model’s negative answer. Attempt identity traces retries/costs; result-slot semantics ensure resume skips only genuinely completed scoring objects.

Observable expectation

The teaching ledger has a1 timeout and a2’s valid result, with one result row bound to a2. If trace writing fails after a2’s row commits, record the artifact gap without duplicating usage as another failed attempt. Restart skips a2 and resumes remaining gaps.

Limits

Local timeout does not establish remote cancellation; retries can repeat effects and need isolation/idempotency. Report exclusions, truncations, refusals and errors beside conditional scores, not best-retry selection. Atomic commits/concurrency need implementation verification.

Sources and evidence

Read the editorial criteria