evaluation

Evaluation & feedback

P27

Scoring Rubrics

Evaluation & feedback

Define anchored criteria and action thresholds before assigning a score.

P28

Self-Critique

Evaluation & feedback

Check a draft against named failure modes before presenting it.

P47

Evidence-First Review

Evaluation & feedback

Demonstrate a defect through its trigger, path and consequence.

P48

Rule-Catalog Review (YAML)

Evaluation & feedbackSource unconfirmed

Store review rules as versioned data with explicit evaluation and failure actions.

P108

Blind A/B Comparison

Evaluation & feedback

Compare outputs on matched tasks while hiding version identity from the evaluator.

P174

Reachability-Scoped Delta

Evaluation & feedbackSource unconfirmed

Scope upgrade advice to behavior the current code can actually reach.

P227

Grade the declared artifact kind

Evaluation & feedback

Choose the grading evidence according to whether the promised artifact is execution or conversation.

P242

Measured baseline ratchet

Evaluation & feedback

Adopt a measured nonregression direction when an arbitrary target would fail the inherited baseline.

P245

Owned negative routing controls

Evaluation & feedback

A negative routing example must route to its intended owner, not merely avoid one wrong skill.

P255

Keep answer keys outside agent reach

Evaluation & feedback

Make evaluation answers structurally unavailable to the system being tested, including history and earlier-trial artifacts.

P265

Use fresh no-change controls after drift

Evaluation & feedback

When execution conditions change, compare a candidate against a contemporaneous unchanged control rather than only an old baseline number.

P275

Evaluate the real application entry point

Evaluation & feedback

Exercise the application path that assembles context, routes tools and handles retries, rather than recreating its final model request.

P344

Choose Metrics from Failure Costs

Evaluation & feedback

Choose evaluation metrics according to the decision and the cost of its failure modes before optimizing a score.

P354

Related-Input Invariant Oracle

Evaluation & feedback

Check defined relationships between transformed inputs and outputs when exact fixtures alone leave broad behavior unchecked.

P384

Independent completion witness

Evaluation & feedback

Separate externally witnessed behavior from completion claims made by the code being evaluated.

P387

Persona-grounded peer benchmarking

Evaluation & feedback

Define the persona and journey boundaries before comparing steps and evidence; retain estimates, unknowns and pending target choices.