Aggregate and Slice Evaluation Gates
Evaluate important task slices beside aggregate metrics using gates declared before inspecting the result.
These examples and illustrative results are independently authored teaching materials, not measured model results.
Use case
A bilingual assistant passes 95/100 while all five Chinese cases fail.95% overall does not establish the promised language capability.
Mechanism
Before inspecting candidates, fix aggregate/important slices, dataset IDs, gates and tolerances. Retain pass/fail/total/gaps and inspect concentration/insufficient coverage; overlapping groups do not sum twice. Required-slice violations prevent promotion under project policy. Missing thresholds are policy gaps, not post-hoc convenient gates.
Bad example
Approve bilingual capability at 95%, hiding Chinese failures or removing its gate after seeing them.
Good example
Report 95/100 overall,0/5 Chinese and 95/95 English, retaining not-ready under predeclared bilingual rules. Inspect failures/representativeness without lowering gates and disclose missing/overlapping/unmeasured slices.
Why the change matters
Averages reflect dominant groups; slices expose failures within promised capabilities. Predeclared policy prevents selection of groups/thresholds merely to pass.
Observable expectation
Teaching 95 English+5 Chinese=100 reconciles totals. Two missing Chinese rows make coverage incomplete rather than silently 0/3. New cross-device groups overlap and cannot be added again to total.
Limits
Small slices have uncertainty and do not represent every language/device. Many slices invite selective comparisons; choose from needs. Offline gates do not guarantee production and source thresholds/tools are task-specific.
Sources and evidence
- affaan-m/ECC · Evaluate before promotion
File at this versionef648e01899b - affaan-m/ECC · Versioned representative baseline with task-specific thresholds, slices and regression deltas
File at this versionef648e01899b