Evidence Chain / Proof-of-Work
Link conclusions to observed evidence and record missing coverage.
Link conclusions to observed evidence and record missing coverage.
Define anchored criteria and action thresholds before assigning a score.
Check a draft against named failure modes before presenting it.
Develop alternatives independently before cross-reviewing their trade-offs.
Demonstrate a defect through its trigger, path and consequence.
Store review rules as versioned data with explicit evaluation and failure actions.
Use a model evaluator with a fixed rubric and checkable evidence.
Test whether a fresh reader can understand the document without the author's history.
Compare outputs on matched tasks while hiding version identity from the evaluator.
Render the artifact and inspect its visible output before delivery.
Report an implemented requirement that is itself defective without hiding the conflict.
Start review from the diff and expand only for a named cross-cutting risk.
Reuse valid test evidence and run focused checks when a concrete doubt remains.
Scope upgrade advice to behavior the current code can actually reach.
Distinguish tool presence, feasible execution and observed execution in reports.
Compare baseline and candidate behavior on identical inputs, including failures.
Remove instructions only after testing their effect against the model's baseline behavior.
Evaluate specification fit and engineering standards as separate verdicts.
Report an untestable behavior boundary as an architectural finding.
Use static tracing for unsafe execution and state the runtime evidence it cannot supply.
Measure the produced artifact against its delivery constraints and apply targeted corrections.
Freeze extraction fields and validate a small sample before scaling.
Derive negative scenarios from each trust boundary before selecting controls.
Choose the grading evidence according to whether the promised artifact is execution or conversation.
A local lenient reader passing is not proof that the real consumer accepts the same artifact.
Validate grade identities and reconstruct totals from the declared expectation set.
Review the validator as well as the implementation so green cannot be bought by weakening the contract.
Adopt a measured nonregression direction when an arbitrary target would fail the inherited baseline.
Compare one change under equivalent measurement conditions and require signal beyond normal variance.
A negative routing example must route to its intended owner, not merely avoid one wrong skill.
Treat reviewer output as hypotheses and classify against the actual artifact before acting.
Require the requested behavior and the standing project readiness bar separately.
Repeat a successful check only when relevant code or baseline assumptions have changed.
Check the telemetry path by generating a known signal and finding it through the intended channel.
Make evaluation answers structurally unavailable to the system being tested, including history and earlier-trial artifacts.
Distinguish a scorable result from an attempt failure and keep each result’s grade, trace and usage bound to the same invocation.
When a prompt builds an artifact that is later evaluated, measure build variability separately from repeated scoring of one artifact.
Attach optional verification questions to actual claims and assumptions, and omit them when the user already requested that verification.
When a metric or grader changes, apply it consistently to stored outputs from every compared variant before recommending a winner.
When execution conditions change, compare a candidate against a contemporaneous unchanged control rather than only an old baseline number.
For an agent that changes an environment, score the resulting state and use transcripts for explicit process constraints.
Compare iterative candidates with the same saved reference artifacts and retain older references when adding a harder comparison.
Record whether reference answers are independently verified or generated by a system under comparison.
Keep withheld-case results out of the prompt that proposes revisions, and name the resulting selection set accurately.
After selecting interacting settings in separate stages, evaluate the final combination on fresh runs against criteria fixed before confirmation.
When repeated edits fail to move an evaluation, classify remaining failures by the layer that could explain them before adding more instructions.
Exercise the application path that assembles context, routes tools and handles retries, rather than recreating its final model request.
Walk backward from a failing effect through callers and values to identify where the bad input originated.
Determine whether metrics are increments or cumulative observations before aggregating them.
Derive expected values independently and check that plausible wrong behavior would fail the test.
Validate a visual sequence as a causal user journey, not just a screen count.
Keep an assessment from leaking its answer through presentation differences.
Verify a guardrail rejects an intentional violation, then returns to clean success.
Distinguish incomplete valid fixtures from intentionally invalid input.
Turn syntactic mistakes into deterministic checks and reserve prose rules for judgment.
Compare design alternatives inside the real host context and hold data constant.
Distinguish historical evidence, unverified snapshots, current observations and unknown state.
Apply declared non-negotiable blockers before deriving a readiness decision from an aggregate quality score.
Label a judgment that was given the measured answer as conditioned description rather than independent corroboration.
Choose evaluation metrics according to the decision and the cost of its failure modes before optimizing a score.
State what a measurement actually counts before using it as evidence for a different resource or behavior.
Check defined relationships between transformed inputs and outputs when exact fixtures alone leave broad behavior unchecked.
Use the same declared trials for candidate and baseline and retain each paired result before applying an aggregate gate.
Revoke downstream passes that depend on a prerequisite which failed its own validation.
Distinguish succeeding once with retries from succeeding on every repeated trial.
Evaluate important task slices beside aggregate metrics using gates declared before inspecting the result.
Report whether the requested behavior works separately from whether an added instruction improves it.
Compare result trends only when the covered checks and measurement conditions match.
Preserve a finding's evidence limits in every summary and proposed regression assertion.
Separate externally witnessed behavior from completion claims made by the code being evaluated.
Test a shared-state invariant across competing asynchronous operations, rather than reviewing each operation alone.
Define the persona and journey boundaries before comparing steps and evidence; retain estimates, unknowns and pending target choices.
Fix a stochastic evaluation panel before seeing outcomes and retain every planned attempt.