A focused model may be suitable for a clear reference and quantity in a short document, while a contradictory attachment needs more careful review. Routing those cases differently is worth testing. It is not enough for the first model to say that it is confident and the second model to agree.
This article proposes a small-versus-larger-model evaluation, using GPT-6 Luna as the focused candidate and a current approved comparison model for review. All sample cases are fictional. No accuracy, cost saving or production benefit has been measured, and no model calls are performed by this article.
Check the current catalog before choosing candidates
The September changelog records GPT-6 Luna's 22 September 2026 release. The fresh 6 October snapshot also records the 25 September image-encoding fix and recommends rerunning affected visual evaluations. A pilot using image inputs should identify the tested model and configuration rather than treating earlier results as automatically current. Source: OpenAI changelog
The official model catalog checked on 6 October lists GPT-6 Luna for cost-sensitive, high-volume workloads, GPT-6.1 Sol for balancing intelligence and cost, and GPT-6 Astra for complex reasoning and coding. These descriptions support candidate selection, not a ranking of legal or business extraction accuracy. Source: OpenAI model catalog
The current Luna reference lists text and image inputs and structured-output support. Confirm the actual interface, model identifier and available features in the intended environment. A supported format does not establish that the workflow's particular documents are easy or that a review model will resolve every ambiguity correctly. Source: GPT-6 Luna model reference
Define easy and ambiguous from evidence rules
Use the workflow owner's approved extraction requirements. An easy candidate might have one explicit account reference, one quantity and a readable source. Missing required fields, duplicated references, inconsistent units and contradictory attachments should trigger review through application rules, regardless of the model's self-reported certainty.
Keep model-reported uncertainty as additional evidence, not the sole routing control. A wrong answer can look complete. Review a selected sample of accepted results against source ground truth to detect false certainty and missed escalation.
The reviewer model produces another proposal, not an authoritative business decision. Apply the same source and validation checks, with professional or domain review where required. Two models can repeat the same unsupported interpretation.
Small-versus-large model evaluation protocol
Use this complete proposed protocol for a fictional service-request extraction task. The domain owner approves expected fields and dispositions before any run.
Paths: compare a single approved review-model path with a hybrid path that begins with Luna and escalates under the rules below. Use the same source set, output requirements and acceptance criteria for both. Record model identifiers, configuration and source versions.
| Fixture | Expected extraction or disposition | Hybrid routing requirement | Review evidence |
|---|---|---|---|
| MODEL-A clear request | One source-supported test account reference and quantity | Candidate acceptance after deterministic validation | Verify exact values and source location |
| MODEL-B missing reference | Missing-field hold; no guessed account | Application escalates or holds independently of model confidence | No invented identifier or downstream action |
| MODEL-C two conflicting references | Explicit unresolved conflict | Escalate with both source references | Review model must not choose a target without approved evidence |
| MODEL-D unreadable quantity | Unreadable-source status | Hold or route to approved human source review | No plausible replacement number |
| MODEL-E conditional wording | Preserve the condition and uncertainty | Escalate when the approved task cannot resolve the condition deterministically | Domain reviewer checks interpretation and limitations |
| MODEL-F duplicate copies | One verified source-derived item with collection provenance | Apply approved identity rule; hold conflicting versions | No double count or false independent confirmation |
For every run, store fixture and task version, expected fields, observed fields, source citations, routing reason, application checks, model-reported uncertainty, escalation outcome, accepted or repaired status, reviewer corrections, elapsed completed-task time and observable usage. Keep the ground-truth record independent of the generated answer.
Cost record: include all first-stage calls, escalated calls, retries, discarded outputs, tool or extraction charges and professional review effort. Divide by accepted completed tasks under the approved accounting rule, with unresolved cases shown separately. A low first-stage token price is not a total-workflow saving.
Error review: classify wrong extraction, unsupported certainty, unnecessary escalation, missed escalation, reviewer-model error and source-coverage failure separately. An answer with correct field shape but wrong source value is an extraction failure. A difficult case sent to the reviewer is not a failure merely because it costs more.
Decision rule: the owner compares observed quality, missed ambiguity, review burden and total cost with the one-model baseline. Keep the tested task scope explicit and retain an approved fallback. No fixed pass percentage or savings target is asserted here; those criteria require the business owner's actual risk and cost decisions.
The protocol is ready when expected cases, routing rules, handling arrangements and review owners are approved. It produces an evaluation plan, not evidence that either path has passed.
Use structured output without trusting its meaning
OpenAI's Structured Outputs guide describes schema-constrained responses and refusal handling. A schema can require reference, quantity, source and uncertainty, but does not guarantee factual accuracy. Treat refusals and unusable outputs as explicit outcomes rather than accepting empty fields as a successful extraction. Source: OpenAI structured outputs
Validate required values against source evidence. Do not accept an output solely because it uses the permitted enum or correct data type. A valid account-reference format can still identify the wrong fixture account.
Also inspect context passed to the reviewer. Escalation needs the relevant permitted sources and first-stage result, with its limitations. An unsupported summary of the source can hide the original contradiction from the second model.
Work through normal, missing and duplicate cases
In a hypothetical normal case, MODEL-A contains a clear test reference and quantity. Luna proposes the expected values, the application validates them and a source check confirms their support. The pilot records the observed outcome and cost without extrapolating to all live documents.
In a missing-information case, MODEL-B lacks its required reference. The deterministic rule holds or escalates it even if the model returns a confident complete-looking answer. The reviewer must preserve the missing evidence rather than guess an identity from a display name.
In a duplicate case, two copies share the same verified document identity. The approved rule prevents repeated items while retaining provenance. If their values differ, the case becomes an explicit conflict; the second model is not authorised to silently choose the more convenient value.
FAQ about routing extraction between models
Is Luna automatically accurate enough for simple documents?
No. Its documented capabilities make it a candidate to evaluate. Define the task and compare source-supported results on the approved sample set before choosing a production scope.
Can model confidence alone decide escalation?
It should not be the only control. Use deterministic checks for missing or conflicting evidence and review accepted outputs for false certainty. A confidently wrong answer can bypass a confidence-only route.
Does a larger reviewer model replace human review?
Not where the approved process requires domain or professional judgement. It provides another proposal that still needs source checks and the appropriate decision owner. Measure remaining review effort rather than assuming it disappears.
If your business needs help defining this process, explore Custom AI agents, the wider AI automation services, and our custom-agent workflow guide. The agents and automation comparison and custom AI agent glossary explain the terms. To discuss your records and approval rules, get in touch.

