P367 · Evaluation & feedback

Aggregate and Slice Evaluation Gates

Evaluate important task slices beside aggregate metrics using gates declared before inspecting the result.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A bilingual assistant passes 95/100 while all five Chinese cases fail.95% overall does not establish the promised language capability.

Mechanism

Before inspecting candidates, fix aggregate/important slices, dataset IDs, gates and tolerances. Retain pass/fail/total/gaps and inspect concentration/insufficient coverage; overlapping groups do not sum twice. Required-slice violations prevent promotion under project policy. Missing thresholds are policy gaps, not post-hoc convenient gates.

Bad example

Approve bilingual capability at 95%, hiding Chinese failures or removing its gate after seeing them.

Good example

Report 95/100 overall,0/5 Chinese and 95/95 English, retaining not-ready under predeclared bilingual rules. Inspect failures/representativeness without lowering gates and disclose missing/overlapping/unmeasured slices.

Why the change matters

Averages reflect dominant groups; slices expose failures within promised capabilities. Predeclared policy prevents selection of groups/thresholds merely to pass.

Observable expectation

Teaching 95 English+5 Chinese=100 reconciles totals. Two missing Chinese rows make coverage incomplete rather than silently 0/3. New cross-device groups overlap and cannot be added again to total.

Limits

Small slices have uncertainty and do not represent every language/device. Many slices invite selective comparisons; choose from needs. Offline gates do not guarantee production and source thresholds/tools are task-specific.

Sources and evidence

Read the editorial criteria