Question-to-signal observability
Choose instrumentation from the questions an operator must answer.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
Operators need failed runs, latency locations and duplicate retries for an import. Teaching fields are run_id/stage/error class/duration; full payloads may be private. Logging everything can cost more without answering those questions.
Mechanism
Define questions/decisions first, choose minimal events for individual failures, bounded metrics for trends and correlated spans for paths. Specify fields/sampling/scope and verify answerability rather than infer root cause from a mnemonic or dump payloads.
Bad example
Store all data/secrets and arbitrary counts, calling observability complete despite no run correlation.
Good example
Map teaching failures→run/stage/error events, timing→stage durations/correlated paths, duplicates→intent/attempt identity. Keep metric labels bounded without run_id per-series explosion. Exclude secrets/unrelated payload and test actual answerability rather than count-only diagnosis.
Why the change matters
Questions determine fields/correlation; volume does not establish completeness. Signals connect cases/trends/paths with bounded disclosure.
Observable expectation
A failed fixture traces to stage/class, slow work to observed spans and retries to intent. Missing/incomplete telemetry stays unknown.
Inspect purpose/bounded labels/no raw payload retention; no real monitoring connects.
Limits
Logs need evidence for causes, metrics establish no root cause and sampling omits cases. Source counts/mnemonics are heuristics; validate task coverage.