What should we retest after OpenAI's September image-input fix if we process scans?

Rerun previously misread scans after OpenAI’s September image-input fix. Compare field evidence, review exceptions and decide whether to change model routing.

AI Automation
6 October 2026Updated 06 Oct 202610 min readBukhosi Moyo

Quick Answer

Retest the exact scans that previously failed, plus previously correct controls, using unchanged prompts, schemas and image preparation. Compare each extracted field with human-checked evidence, not just the old output. Review missing values, ambiguous text, duplicate pages and unsupported guesses separately. Keep document-model routing unchanged until the affected route passes your proposed acceptance rules and a human owner approves any limited change.

Key Takeaways

  • Start with previously misread scans and previously correct controls.
  • Freeze image preparation, prompts and schemas for a meaningful comparison.
  • Score evidence, missing values and duplicates separately from formatting.
  • Change routing only for document groups supported by reviewed results.

Want the full breakdown? Scroll below.

Person planning a workflow on a whiteboard
On this pageJump to a section
  1. 11. Establish whether your scan route was affected
  2. 22. Build a fixed set of failures and controls
  3. 33. Write the reference answers before reviewing new output
  4. 44. Freeze the run configuration and retain evidence
  5. 55. Score changes at field and document level
  6. 66. Keep exception review separate from downstream action
  7. 77. Decide routing changes by document group
  8. 8Reusable regression-review checklist
  9. 9Hypothetical walkthrough: one correction and two exceptions
  10. 10FAQs about rerunning scan extraction
  11. 11Turn the review into a bounded next step
  12. 12Sources

Share this article

Bukhosi Moyo

Growth Partner

Need help growing your company?

We build SEO-first websites and growth systems for South African businesses.

Get Started

Retest previously misread scans, previously correct controls and exception handling before changing document-model routing. Use the same files and extraction settings, then compare field values against human-checked page evidence. A vendor fix is a reason to evaluate again, not a reason to send every scan through a different model.

OpenAI’s Source: 25 September 2026 changelog reports an image-encoding fix affecting GPT-6 Sol and GPT-6 Luna and recommends rerunning image-input evaluations. The procedure below is a proposed regression review for a South African document team, not a claim of measured improvements.

1. Establish whether your scan route was affected

Start with request records that show which model received image inputs. Do not treat every document failure as part of this fix.

Trace a failed document from upload to extraction. Did your application send rendered page images to GPT-6 Sol or GPT-6 Luna? Did it send only text produced by an earlier OCR step? Did another processor handle the scan entirely? Record the answer for each route before selecting cases.

The September announcement concerns those named models’ image understanding. It does not establish that an OCR-only route improved, that every scan error came from encoding, or that a later model is the right replacement.

For the proposed review, classify routes as directly relevant, indirect or outside scope. Include indirect routes only when the affected model used images to check or interpret earlier extraction.

Confirm endpoint access and supported settings in your own project. Have the responsible security or privacy person approve the evaluation environment and document handling. Do not infer regional processing or retention suitability from the fix announcement.

2. Build a fixed set of failures and controls

Choose a frozen document set that contains known failures and enough previously correct examples to expose regressions. Do not select only scans that now look promising.

Group failures by visible problem: faint text, rotated pages, cropped edges, stamps over values, handwritten additions, table alignment and multi-page confusion. Add missing-page and duplicate-page cases, even when the original complaint concerned a wrong digit. A route must handle absence and repetition as well as readable text.

Use actual authorised documents where appropriate, or approved redacted copies. If redaction changes the pixels, keep that version fixed throughout the comparison and note that it differs from the original failure input.

A hypothetical starting set might contain 24 previously failed documents and 12 controls. Those counts are illustrative, not a statistically justified sample size. Select coverage according to your document mix and consequences.

Keep a separate later holdout set. The first set can guide diagnosis; the holdout should check whether a proposed route change works beyond the examples used to design it. Symaxx’s document processing service is the relevant service area for this extraction-and-evidence boundary.

3. Write the reference answers before reviewing new output

Create human-checked expected values and evidence locations before scoring the rerun. Old model output is a historical record, not ground truth.

For each required field, record the document identifier, page, visible label, raw text, expected normalised value and acceptable state. Proposed states are present, missing, ambiguous and unreadable. Keep missing distinct from unreadable: a blank field is not the same as a field hidden by a poor scan.

Define normalisation narrowly. You might propose accepting spacing differences in a supplier name while requiring exact invoice digits. Preserve the raw date before converting it. If the document and agreed business context do not resolve the date format, mark it ambiguous rather than silently choosing one.

Have a second reviewer resolve disagreements on consequential fields. Record the reason for the reference decision and keep unresolved cases outside the ordinary accuracy denominator.

For invoices, a finance reviewer should decide what requires escalation. Extracting a VAT amount or account number does not validate tax treatment, authorise payment or establish that the document is genuine.

4. Freeze the run configuration and retain evidence

Change as little as possible during the first rerun so that differences remain interpretable. Keep the prompt, schema, renderer, image size, page order and preprocessing configuration unchanged where available.

Store a file checksum, configuration versions, requested model identifier, run date, request identifier and raw response. Record returned model metadata when available. Keep any historical pre-fix output beside the new output without overwriting either.

You may not be able to recreate the pre-fix service behaviour. In that case, compare the new run with archived output and the reference answers, and describe it as a historical comparison rather than a controlled replay of the old model.

If settings or files have changed since the failure, flag the case as confounded. Test a preprocessing improvement separately rather than crediting it to the image fix.

OpenAI’s Source: Structured Outputs guide describes schema-constrained responses. Use that capability only after confirming support for your chosen configuration. Treat a valid object as a formatting result; separately check whether each value matches the scan.

5. Score changes at field and document level

Compare correctness, evidence and exception behaviour separately. A document with more populated fields is not necessarily a better extraction.

Use these proposed comparison labels: corrected, unchanged correct, unchanged wrong, newly wrong and unresolved. Add operational outcomes such as refused, incomplete, timed out or failed. Never drop failed requests from the report merely because they contain no extracted fields.

Count exact matches among fields whose reference values are settled. Separately count unsupported populated values where the reference says missing or unreadable. Check page coverage, row association and duplicate handling at document level. Correct amounts attached to the wrong line items should fail the association check.

Report results by failure group and document type, not just as one average. A hypothetical improvement on clear invoices could conceal new errors on stamped delivery notes.

For repeatability, propose several runs on the same selected difficult cases and retain every result. Agree the repetition count before running. This can reveal unstable outputs, but does not establish reliability across the full production workload. Measure review effort directly if it matters; do not assume better-looking extraction saves time.

6. Keep exception review separate from downstream action

Send missing, ambiguous and duplicate cases to a reviewer with the source page visible. The model should not resolve these cases by inventing a value or triggering a business action.

Propose an exception queue containing the candidate value, reference evidence, reason for referral and assigned owner. Let reviewers accept a supported extraction, correct it with evidence, request a clearer scan or leave it unresolved. Preserve both the initial output and the human correction.

During evaluation, disable production writes, notifications and payment-related actions. OpenAI’s Source: function-calling guide explains that the application executes model-requested functions. Use that application boundary to validate inputs and enforce review gates rather than treating a tool request as permission.

The AI agents versus automation comparison can help frame whether fixed rules are sufficient. For this review, prefer an explicit sequence over an agent that chooses new routes during the test. If you use a custom AI agent, keep its permissions and choices inside the evaluation boundary.

7. Decide routing changes by document group

Keep the existing route unless reviewed evidence supports a narrowly defined change. Do not promote a model across all document types because one failure group improves.

Before the rerun, propose acceptance rules with the process owner. One possible rule is no newly wrong critical fields in the selected set, no unsupported values for missing critical fields, and mandatory referral of unresolved critical fields. These are proposed gates, not universal safety guarantees or proof that unseen documents will pass.

Use three decisions: retain the current route, investigate further, or approve a limited routing trial. Assign an owner, eligible document group, monitoring measures and rollback condition to any trial. Continue human review for consequential outputs.

If you already use another processor, compare it on the same inputs and reference answers. Google’s Source: Document AI overview describes OCR, extraction, classification and splitting processors. That establishes possible comparison categories, not a winner for your scans. Check specific processor availability, version and regional suitability before proposing a switch.

Reusable regression-review checklist

Use this proposed checklist as the review record. Complete it before approving a routing trial.

  • Scope: record affected image-input routes, model identifiers and excluded text-only routes.
  • Inputs: freeze authorised failure scans, correct controls and a separate holdout; retain checksums and page counts.
  • Reference: record each field’s expected value, state, page evidence and reviewer; settle consequential disagreements.
  • Configuration: record prompt, schema, renderer, preprocessing, endpoint and run settings; flag changed inputs.
  • Baseline: preserve historical outputs; mark unavailable baselines rather than reconstructing them.
  • Rerun: retain every request outcome and raw response; keep production actions disabled.
  • Comparison: label fields corrected, unchanged correct, unchanged wrong, newly wrong or unresolved.
  • Exceptions: check missing values, unreadable text, ambiguous dates, row association and duplicate pages.
  • Decision: apply pre-agreed proposed gates per document group; assign retain, investigate or limited trial.
  • Ownership: name the approving owner, exception reviewer, trial boundary and rollback condition.
Finding Proposed response
Corrected failures, controls preserved Consider a limited trial after holdout review
Any new critical error or unsupported critical value Hold the routing change and investigate
Missing baseline or changed preprocessing Report current quality; avoid attributing improvement to the fix
Unresolved source text Request evidence or a clearer scan; retain human review

Hypothetical walkthrough: one correction and two exceptions

A normal case can support a limited trial, while an ambiguous or duplicate case still needs human handling. All records and values in this walkthrough are hypothetical.

Suppose a supplier invoice clearly shows reference INV-084 and total R2,760.00 on its first page. The archived extraction returned INV-034. The rerun returns INV-084 with the correct total and page reference. A reviewer confirms both against the scan. Label the reference corrected and the total unchanged correct. This is one corrected case, not proof of better invoice processing overall.

Now suppose another invoice’s date reads 04/05/26, with no agreed evidence establishing the format, and its bank-detail box is cropped away. The expected output is an ambiguous date and an unreadable bank-detail field, not a guessed date and account number. If the rerun supplies confident values, record unsupported extraction and hold any proposed critical-field routing change. A finance reviewer requests a complete scan or other authorised evidence.

Finally, suppose a PDF contains the same invoice page twice. The proposed expected handling is one candidate invoice record linked to both page occurrences, with a duplicate flag. A reviewer checks whether the pages are identical or represent a revision before resolving the flag. Neither the correction nor the duplicate review authorises payment.

FAQs about rerunning scan extraction

What if we did not save output from before 25 September?

Evaluate current output against human-checked references, but do not report a measured before-and-after improvement. Reconstruct the failure category from tickets only where evidence exists. Keep those cases labelled separately from cases with archived responses. Your current evaluation can still inform routing, but the September fix cannot be isolated as the cause. Start retaining inputs, settings and outputs for future comparisons within approved access and retention rules.

Should we improve image quality during this rerun?

Keep the original image preparation for the first comparison. If you resize, deskew or render pages differently, you introduce another change. Run a separate proposed comparison for preprocessing, with the same reference answers and clearly labelled variants. A reviewer should check that cropping or cleaning has not removed stamps, notes or other evidence. Use the resulting evidence to choose a configuration, not to claim that the vendor fix caused every gain.

Should we retry scans that previously reached a human reviewer?

Yes, where authorised, but keep the rerun separate from completed business records. Compare against the preserved original scan and reviewed answer without overwriting the human decision. Avoid re-running downstream actions. If new output contradicts a consequential historical record, refer it to the responsible person rather than changing that record automatically. Agree retention, access and escalation arrangements with the relevant security, finance, legal or process owner.

Turn the review into a bounded next step

Choose a fixed evaluation and a named decision owner before commissioning a broader workflow change. The custom AI agents workflow resource can help frame orchestration if your review spans several tools, but tool orchestration should not replace evidence checks.

If your business needs help defining this regression set and its exception gates, get in touch with Symaxx about AI automation. Bring authorised sample scans, archived errors and the current route configuration. The useful starting point is a scoped evaluation plan, not a promise that a model update will resolve every extraction problem.

Sources

Share this article

Bukhosi Moyo

Written by

Bukhosi Moyo

CEO & Founder

Bukhosi is the founder and lead SEO strategist at Symaxx. He architects search-first digital systems for South African businesses, combining technical engineering with commercial strategy to build long-term organic assets.

Feedback

Was this helpful?

Tell us how this article felt in one click.

Back to Insights

Need help executing this strategy?

Our team turns these insights into revenue-generating search architectures for your business.