P35 · Agent orchestration

Cost-Optimized Model Routing

Select a model using measured task requirements and an explicit escalation path.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Route mechanical rename and concurrency review under budget. Teaching candidates A/B are unmeasured on these task classes. Lower per-token prices do not establish lower completed-task cost; stronger labels do not establish relevant quality.

Mechanism

Follow host/user selection rules, define representative cases, quality gates and an approved measurement budget. Compare task-complete cost/quality including failed retries/checks. Configure routes/escalation only where allowed. Unmeasured candidates stay proposals, and changed model/prompt/workloads need new evidence.

Bad example

Use the cheapest for everything and trust confidence, or use the strongest for every rename while ignoring retry cost.

Good example

Define separate rename/concurrency fixtures. Under approved budget compare A/B quality and completed-task cost using actual usage/current applicable rates, counting failed retries and verification. Choose qualifying candidates only within host policy, with explicit escalation. Without measurements report unverified routing proposals, not fixed savings or universally best models.

Why the change matters

Task distribution and hard failures can reverse cost ranking. Quality plus complete costs grounds selection in this workload rather than cheap/powerful role recipes.

Observable expectation

An illustrative ledger records cases, configurations, results, usage/rates and reruns; unmeasured A/B totals/quality stay unknown. Inspect matched conditions and complete failure costs instead of single-case generalization.

No paid models or host settings are changed here.

Limits

Frozen model lists, ratios, effort thresholds/prices are not current promises. Evaluation has cost and live shadow effects need budget/safe replay. Lower cost that loses required quality is not successful optimization.

Sources and evidence

Read the editorial criteria

Related methods