P275 · Evaluation & feedback

Evaluate the real application entry point

Exercise the application path that assembles context, routes tools and handles retries, rather than recreating its final model request.

Editorially reviewed

These examples and illustrative results are independently authored teaching materials, not measured model results.

Use case

Compare support prompts in an application that assembles customer context, retrieves, retries and routes tools. Copying the prompt into a direct model request evaluates a replacement path.

Mechanism

Locate the real handler/production configuration and design isolated test accounts or a thin wrapper. Substitute unauthorized external effects only, retaining context/retrieval/retry/tool wiring and recording differences. Inspect one complete saved pilot row for served model, usage, trace and result; fix missing runner fields before scaling. Match entry, parameters and budget to the plan.

Bad example

Copy system prompt/model name into a new script, bypass the handler and claim improved support application quality.

Good example

Send a teaching query through the real test entry, preserving retrieval/routes and stubbing outgoing messages. Inspect a complete pilot and actual configuration before scaling. State production/substitute differences; bare-model grades are not whole-application effects.

Why the change matters

Application quality includes paths before/after the model. Real entry points include them, while thin substitutes limit effects without silently changing the object measured.

Observable expectation

The teaching pilot records context assembly, retrieval, allowed tools and retries, with no real outgoing messages. Missing usage/traces prevent full evaluation. A bare-prompt experiment must be labeled that way.

Limits

Test accounts do not themselves isolate production effects. Substitutes/configuration/fallbacks can differ, and reported model identity may not reveal every hidden substitution. This article calls no model or messaging service.

Sources and evidence

Read the editorial criteria