P381 · Context management

Budget context by when it loads

Measure always-loaded metadata separately from the full content eagerly loaded for one invocation.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

A pack shrinks an entrypoint from 20 KiB to 2 KiB but requires another 18-KiB reference on every invocation. Teaching always-loaded name/description totals 0.4 KiB; full front matter is inventoried separately but included in the body ledger. Compare consumed content rather than equating file shrinkage with token savings.

Mechanism

Build separate discovery-metadata, full-front-matter and invocation body/required-reference ledgers. Explain actual host loading and duplication. Compare the same pack using one tokenizer, or label byte figures as estimates. List conditional references by trigger rather than treating all packaged files as always loaded.

Bad example

The entry fell from 20 to 2 KiB, so claim a 90% invocation-context reduction while omitting its mandatory 18-KiB reference and discovery overhead.

Good example

Inspect before/after loading for this pack. Both discovery lists contain 0.4-KiB name/description; full front matter is noted but included in body accounting. Old invocation reads 20 KiB; new reads 2+18 KiB. Report distinct ledgers/methods without double-counting input. Entry bytes alone do not establish 90% invocation savings; keep token figures estimated when unmeasured.

Why the change matters

Moving text into an eagerly required reference changes location, not consumption. Timing-aware complete ledgers show discovery versus invocation costs and avoid omitted or double-counted references.

Observable expectation

The illustrative invocation total stays 20 KiB and discovery stays 0.4 KiB. Without tokenizer records, exact token totals remain unknown.

Check mandatory/conditional loading, front-matter duplication and matched comparison methods. Entry shrinkage alone is insufficient evidence of reduced cost.

Limits

Bytes, tokens, cache billing and context occupancy differ by host/payload. Source budgets/ratios are not universal capacity or pricing promises; smaller context does not establish equal quality. No tokenizer API is called here.

Sources and evidence

Read the editorial criteria