Keep attempts, grades and traces coherent
Distinguish a scorable result from an attempt failure and keep each result’s grade, trace and usage bound to the same invocation.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
Attempt a1 for case-c1 repetition 1 times out; after recovery a2 returns a scorable answer. Treating a1 as a negative result blocks resume, while pairing a2’s grade with a1’s trace mixes evidence.
Mechanism
Use (case,rep) for result slots and attempt_id for invocations. Durably commit scorable rows with same-call output, grade, trace, model and usage. Append infrastructure failures without scorable output to attempts while leaving result gaps. Define retries, refusal, truncation and grader-failure semantics and account all usage. Resume missing slots without replacing valid results with another call. A post-row trace failure must not duplicate billing records.
Bad example
Store a1 timeout as false in the result slot; attach a1 logs to successful a2 and count only successful-call cost.
Good example
Record a1 timeout class and available usage without inventing an answer. If a2 is allowed by retry policy, bind its result, grade, trace and usage to a2, preserving a1 history. Resume idempotently at case-c1/rep1 and report failure rates alongside conditional grades.
Why the change matters
Infrastructure no-answer differs from a model’s negative answer. Attempt identity traces retries/costs; result-slot semantics ensure resume skips only genuinely completed scoring objects.
Observable expectation
The teaching ledger has a1 timeout and a2’s valid result, with one result row bound to a2. If trace writing fails after a2’s row commits, record the artifact gap without duplicating usage as another failed attempt. Restart skips a2 and resumes remaining gaps.
Limits
Local timeout does not establish remote cancellation; retries can repeat effects and need isolation/idempotency. Report exclusions, truncations, refusals and errors beside conditional scores, not best-retry selection. Atomic commits/concurrency need implementation verification.
Sources and evidence
- anthropics/skills · Resume and error sidecar
File at this version8a1541c4a3ff - anthropics/skills · No answer is not negative answer
File at this version8a1541c4a3ff - anthropics/skills · Resolved run scope, canary and complete result record
File at this version8a1541c4a3ff - anthropics/skills · Result-before-companion completion flaw
File at this version8a1541c4a3ff