P338 · Tool use

Deterministic Parse with Uncertain-Case Escalation

Run a deterministic parser first and send only structurally uncertain items for model review.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Convert regularly formatted quiz text into structured data. Teaching Q1 contains a question, A/B/C choices and answer B; Q2 has the same fields but no answer. Preserve determinate extraction without asking the model to invent an answer from general knowledge.

Mechanism

Define parsing grammar and validation for ID, question, choices and answer. Parse deterministically first and retain raw text with locations. Escalate missing fields, malformed structures and unassigned text so entire items cannot disappear silently. Give the model only relevant source and uncertainty reasons, separating extraction from solving. Apply the same schema and source checks afterward; unsupported answers remain missing.

Bad example

Send Q1 and answerless Q2 to the model for complete JSON. Require an answer for each and accept its additions.

Good example

Parse Q1 and Q2 under the known grammar while retaining source. Extract B for Q1. Escalate Q2 with missing_answer so the model checks for an overlooked source answer rather than solving the quiz. Revalidate repairs; if the source has no answer, preserve missing and request source completion.

Why the change matters

Stable structures suit reproducible parsers, while models handle bounded ambiguity. Shared acceptance gates prevent the model path from weakening requirements. Retained source also exposes items the parser did not recognize.

Observable expectation

Teaching output accepts Q1 with answer=B and keeps Q2 in the uncertain queue with no answer. Add Q3 with a malformed ID delimiter: report its unparsed text rather than dropping it. Reconcile input count with accepted and pending counts and check every accepted field against the source.

Limits

Heuristic scores are not calibrated correctness probabilities. Format changes require parser updates, and unrestricted prose may not fit a fixed grammar. Models and schemas can both accept wrong meanings, requiring source or human checks. Frozen success-rate and cost claims do not establish performance here.

Sources and evidence

Read the editorial criteria