P274 · Evaluation & feedback

Classify the failing layer when optimization stalls

When repeated edits fail to move an evaluation, classify remaining failures by the layer that could explain them before adding more instructions.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Repeated stronger use-search instructions leave retrieval scores flat. Teaching traces show search was never registered; more prompt text cannot fix missing harness tools. Wrong grading or variance may also explain failure.

Mechanism

Read representative full traces, outputs and configuration. Provisionally classify evidenced content gaps, grader disagreement, infrastructure, unreachable structure and variance, allowing multiple layers/unknowns. Record cases/counts/locations and probe the best explanation, such as registration-only repair on the same case. Grader changes require consistent regrading; indistinguishable movement calls for uncertainty rather than unlimited instructions.

Bad example

After every flat round add must always search without checking availability or grading.

Good example

Inspect failures first. Missing search registration targets harness wiring and a same-case probe; correct citations rejected by grading target rubric checks/regrading. Classify with concrete evidence and preserve unknowns, adding content only for genuine content gaps.

Why the change matters

The tuned artifact need not be the bottleneck. Layer diagnosis directs intervention toward observations instead of instructions the agent cannot execute or reach.

Observable expectation

Teaching missing registration followed by observed calls supports a wiring explanation but still requires outcome checks. Successful calls with unchanged grades warrant grading inspection. Identical-config flips suggest variance. Every bucket needs case evidence/probes rather than inferred cause from low scores alone.

Limits

Taxonomies aid diagnosis rather than prove it, and cases can span layers. Environment/grader drift remains possible; regrading/repetition follows approved budgets. Unclassified failures stay unknown rather than forced content gaps.

Sources and evidence

Read the editorial criteria