Choose Metrics from Failure Costs
Choose evaluation metrics according to the decision and the cost of its failure modes before optimizing a score.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
An alert dataset has 100 cases:98 routine and 2 urgent. Missing urgency is costly, yet always-routine gets 98% accuracy. That value alone cannot justify deployment.
Mechanism
Define the decision, false-positive/negative/latency costs and population before choosing metrics/denominators/gates. Retain confusion counts, urgent recall and unnecessary alerts, inspecting costly errors directly. Compare simple baselines and diagnose signal/labels/thresholds before complexity. Show unknown policies and representativeness.
Bad example
Approve always-routine for 98% accuracy while ignoring both urgent misses.
Good example
Report 98/100 correct, urgent recall 0/2 and unnecessary routine alerts 0/98 on the same cases. Apply a predeclared urgency policy, inspect missed cases/labels and state false-alarm trade-offs rather than hiding failure under overall accuracy.
Why the change matters
Majority classes dominate accuracy while decision harm may concentrate in a minority. Failure-cost metrics keep those mistakes visible and support a specific choice.
Observable expectation
Teaching always-routine has recall 0 and 2 misses despite 98% accuracy. Always-urgent gets 2/2 recall but 98/98 unnecessary alerts. Reconcile denominators with raw cases; these are hypothetical arithmetic examples.
Limits
Two urgent cases cannot establish reliability; representative data, uncertainty and labels matter. Costs/thresholds are product-specific and shift with distribution/feedback. No real business or safe-deployment claim is supplied.
Sources and evidence
- affaan-m/ECC · Problem framing, mistake budget, metrics and error analysis
File at this versionef648e01899b