P176 · Evaluation & feedback

Differential Baseline-Equivalence Oracle

Compare baseline and candidate behavior on identical inputs, including failures.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Upgrade a serializer while preserving public behavior. Teaching inputs include an empty list, null field and invalid date. Normal candidate tests may pass despite changed errors or ordering, so define equivalence explicitly.

Mechanism

Freeze baseline, candidate, identical fixtures and environment. Choose byte comparison or documented semantic normalization without erasing contractual ordering; compare error categories and effects too. Save item outputs/deltas, separating shared failures from regressions and adding boundaries. Explain approved differences before reporting tested equivalence scope.

Bad example

Run only candidate tests and declare complete equivalence when green; omit errors and empty inputs.

Good example

Run old and new serializers on identical empty-list, null and invalid-date fixtures. Compare public outputs and errors under the contract, retain array order and record every delta/shared failure. Resolve unexplained differences before claiming tested equivalence, not universal equivalence.

Why the change matters

Candidate-only tests may miss the old contract. Differential outputs provide a common-input reference, while explicit comparison rules prevent normalization from hiding consumer-relevant changes.

Observable expectation

Both versions returning [] agree for that case. Omitting a required null field is a delta, as is InvalidDate becoming TypeError. A failure shared by both remains a known problem rather than a candidate fix.

Limits

The baseline can be wrong; equivalence is not correctness. Claims depend on fixtures, environment and definition. Nondeterministic outputs need controlled conditions or semantic predicates. Frozen document/contract comparisons support the mechanism, not observed serializer tests.

Sources and evidence

Read the editorial criteria

Related methods