P160 · Evaluation & feedback

Reviewer Does Not Re-Run the Implementer's Tests

Reuse valid test evidence and run focused checks when a concrete doubt remains.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A reviewer receives complete test logs bound to current code, covering normal cache hits. Reading reveals no assertion for misses. Reuse valid coverage without treating untested paths as passed.

Mechanism

Verify command, code/test identity, environment, configuration, complete output and scope, locating raw reports. Reuse matching evidence with attribution. For a named doubt unanswered by existing runs, perform an allowed focused check or specify the recommended command. Changes to code/shared dependencies invalidate affected evidence. Record gaps and environment; unreadable original logs are not permission to fabricate replacements.

Bad example

Always rerun everything, or never test. Treat a success statement as validation of every path.

Good example

Match existing logs to the same code/configuration and complete command, then reuse normal-path results. Propose a focused regression for the uncovered cache miss; record execution if possible, otherwise unrun. Revalidate changed paths rather than treating the source’s economy policy as a permanent test ban.

Why the change matters

Identity and scope determine reuse, avoiding redundant work without removing justified doubt. Focused tests have a specific purpose instead of allowing a green suite to conceal missing behavior.

Observable expectation

Teaching matrix: same-version complete logs can be reused; truncated evidence calls for reading the original; code/environment differences require reassessment; an unasserted miss remains a gap. Distinguish reused, newly run and recommended-but-unrun checks.

Limits

Reuse does not guarantee trustworthy logs; consequential checks may require independence. Cross-cutting risks can justify broad tests under current authorization. Frozen reviewer budgets are workflow policy. Unseen evidence differs from absent evidence, and warnings need impact-based assessment.

Sources and evidence

Read the editorial criteria

Related methods