P380 · Evaluation & feedback

Comparable-coverage trend reporting

Compare result trends only when the covered checks and measurement conditions match.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Yesterday’s score 6 covered test+lint; today’s9 covers test only. Apparent improvement may come from omitting failing lint.

Mechanism

Record run/unrun/unavailable checks with tools, fixtures, thresholds, weights and environment. Compare composites only for matching coverage/definitions; display unmatched history as incomparable and compare valid individual metrics. N/A has no trend and unknown is not 0. Verify sources instead of treating local scores as universal quality.

Bad example

Say 6→9 improved 3 points without mentioning skipped lint or changed thresholds.

Good example

Compare check sets and definitions first. Missing lint means coverage_changed, not composite gain. Show comparable test results and why lint was unrun. Resume trends after conditions match, retaining actual scope per run.

Why the change matters

Removing difficult categories changes denominator/weights and can create attractive totals. Coverage/condition checks preserve trend meaning and expose unmeasured work.

Observable expectation

Teaching{test,lint}→{test} prevents 6→9 composite gain. Test changing from 10 to 2 cases also changes scope despite the same name. Compute applicable deltas only under matching sets/samples/rules, preserving historical gaps.

Limits

Names alone miss versions/workloads, and local weights lack automatic validity. Time/environment also affect comparability. Incomparable does not mean no improvement; these totals simply cannot demonstrate it.

Sources and evidence

Read the editorial criteria