How do we test document automation with scans that are cropped, rotated or unreadable?

Build a poor-input test pack for cropped, rotated and unreadable scans. Check evidence, rejection, resubmission and human review before accepting extracted data

AI Automation
6 October 2026Updated 06 Oct 202610 min readBukhosi Moyo

Quick Answer

Test damaged scans by creating controlled copies of readable documents, defining the expected state for each copy, and checking both extracted fields and downstream actions. A plausible value is not a pass if its supporting text is missing or unreadable. Include recoverable rotation, cropped required fields, blur, missing pages and duplicates. Verify that each case reaches preparation, resubmission, review or rejection with a clear reason and an accountable human owner.

Key Takeaways

  • Compare damaged copies with readable originals, but keep original answers out of the extraction workflow.
  • Require visible evidence for accepted fields, not just valid formatting or high confidence.
  • Test queue routing, resubmission messages and blocked actions alongside extraction.
  • Separate unreadable evidence, ambiguous values, technical failures and duplicate submissions.

Want the full breakdown? Scroll below.

Person planning a workflow on a whiteboard
On this pageJump to a section
  1. 11. Define what acceptance means before damaging files
  2. 22. Create paired originals and controlled damage variants
  3. 33. Separate recoverable orientation from lost evidence
  4. 44. Require field evidence as well as structured output
  5. 55. Test exception routing and resubmission messages
  6. 66. Run a reusable poor-input acceptance pack
  7. 77. Evaluate failures by consequence, then hand over
  8. 8Worked walkthrough: readable, cropped and duplicate invoices
  9. 9FAQs about damaged-scan testing
  10. 10Sources

Share this article

Bukhosi Moyo

Growth Partner

Need help growing your company?

We build SEO-first websites and growth systems for South African businesses.

Get Started

Test document automation with deliberately damaged copies of readable scans, then check whether the workflow handles missing evidence correctly. Cropping, rotation and blur should not turn uncertain text into accepted facts. The test must cover extracted fields, exception states, operator instructions and blocked downstream actions, not just whether the system returns data.

The procedure below is a proposed acceptance approach for document processing. Its rules and examples are hypothetical, not measured results. The aim is to decide whether a workflow is ready for controlled use and where human judgement remains necessary.

1. Define what acceptance means before damaging files

Accept a document for preparation only when its required fields have readable supporting evidence and its completeness checks pass. This is a proposed technical acceptance state, not approval of an invoice, identity, employment record or legal document.

Start with one document type and one downstream task. For a hypothetical supplier-invoice workflow, that task might be preparing a record for accounts-payable review. Agree which fields are required for that preparation and which can remain absent without blocking it.

Define four business-facing states:

  • Ready for preparation: required evidence is readable and checks pass.
  • Resubmission required: a better source is needed to recover missing evidence.
  • Human review: evidence exists, but interpretation or document identity is uncertain.
  • Rejected input: the submission cannot enter this supported workflow, with a reason and correction path where possible.

Keep technical failure separate. A service timeout is not evidence that the scan is unreadable.

Have the process owner approve these definitions before testing. For finance, tax, legal, employment or security-sensitive documents, the relevant specialist must decide what evidence is sufficient and which actions remain restricted.

2. Create paired originals and controlled damage variants

Build each test around a readable original and labelled damaged copies so that the expected answer is known independently of the extraction system.

Use synthetic documents or suitably authorised, sanitised samples. Have the security or privacy owner decide how files may be stored, accessed and deleted. Do not assume that removing a name removes every sensitive detail.

For each original, record the expected fields and their page locations. Then create variants that isolate specific failures. Hypothetical examples include:

  • Rotate a complete page by 90 degrees without removing content.
  • Crop a required total from the bottom of the page.
  • Crop an optional footer while retaining required fields.
  • Blur a reference until two characters cannot be distinguished.
  • Reduce contrast until a required section becomes unreadable.
  • Remove a page containing required detail.
  • Submit an identical copy under another filename.

Only after testing individual faults should you combine them, such as rotation plus cropping. Record the transformation settings and resulting file identity so another tester can reproduce the case.

Keep the original answer sheet outside the extraction context. Otherwise, a correct-looking answer may come from leaked test information rather than the damaged scan.

3. Separate recoverable orientation from lost evidence

Treat rotation as potentially recoverable, but treat absent pixels as absent evidence. Preprocessing should not be credited with recovering information that the submitted file does not contain.

Google’s Source: Document AI overview describes OCR, deskewing and readability analysis, with image-quality analysis enabled for the relevant OCR processor. These capabilities do not establish how your chosen configuration will handle your test files. Check processor suitability and deployment availability before relying on them.

For the rotated case, inspect the original submission, any transformed image and the extracted result. Confirm that page content remains intact and that the evidence location can still be traced. If your workflow proposes an orientation correction, test that correction as part of the pipeline rather than manually repairing files before submission.

For a crop, inspect whether required content actually survives. A readable heading above a missing total does not support a total value. Sharpening an image is also not permission to infer indistinct digits.

Record where each decision occurs: input validation, preprocessing, extraction or business checks. This makes a failure actionable instead of leaving the team with a vague instruction to improve OCR.

4. Require field evidence as well as structured output

Check every required field against the submitted image, even when the response is well formatted or reports strong confidence.

Microsoft’s Source: Azure Document Intelligence overview describes extraction of text, tables and document structure, and typed field values. Typed output is useful for integration, but a currency-shaped value does not prove that its amount is visible in a damaged scan.

For this proposed workflow, store each field’s candidate value, raw extracted text, page reference, evidence location where available, and evidence state. Suggested evidence states are readable, missing, ambiguous and unavailable. Distinguish a genuinely blank optional field from a required field that has been cropped away.

If an additional language-model step formats the record, allow unknown values explicitly. OpenAI’s Source: Structured Outputs guide explains schema-constrained responses. Use that capability for response shape, not as a substitute for checking source evidence.

Test whether an unsupported candidate remains visible as a flagged candidate rather than silently entering the accepted record. Where a provider does not return usable evidence locations, define and test an alternative operator inspection method before acceptance.

5. Test exception routing and resubmission messages

Pass an exception only when it reaches the right queue with an understandable reason and no prohibited downstream action.

A cropped required field should produce a specific proposed message: “The bottom of the page is missing, including the total. Please upload the complete page.” An unreadable document needs a different request: “Please provide a clearer scan with all page edges visible.” A timeout should trigger a technical recovery path, not a request to rescan.

The operator view should show the submitted image, affected fields, reason, allowed next action and case owner. Avoid a generic failed label that forces staff to investigate from scratch.

Test the response to a replacement file. Link it to the earlier submission, retain the earlier result for the authorised review trail, and evaluate the replacement on its own evidence. Do not carry unsupported values forward automatically.

When planning queue orchestration, the n8n AI agents workflow resource can frame implementation discussions. The acceptance requirement remains independent of the tool: a review label must actually block preparation or export when the agreed rules require it.

6. Run a reusable poor-input acceptance pack

Use the same checklist and expected-state table for each release, adapting required fields to your document type. The pack below is proposed; all example identifiers and transformations are hypothetical.

Poor-input acceptance test pack

Setup checklist

  • Name the document type, required fields, process owner and reviewer.
  • Store an authorised readable original and independent answer sheet.
  • Record each variant’s file ID, parent ID and damage settings.
  • Freeze extraction, preprocessing and routing configuration for the run.
  • Keep answer sheets and original values out of extraction inputs.
Hypothetical case Proposed expected state Evidence and handling check
P01: readable original Ready for preparation Required values match visible text; no business approval occurs.
P02: complete rotated page Ready if recovery succeeds; otherwise review Correct orientation, intact content and traceable field evidence.
P03: required total cropped Resubmission required Total remains unknown; request the complete page.
P04: optional footer cropped Ready if required checks pass Missing optional content does not invent a value.
P05: blurred reference Human review Show ambiguity; reviewer resolves from evidence or requests replacement.
P06: unreadable required section Resubmission required Identify the affected section; block preparation.
P07: required page missing Resubmission required Request the missing page; do not accept an incomplete bundle.
P08: duplicate submission Human review Link the suspected duplicate; do not create another prepared record.
P09: unreadable file format Rejected input Explain the supported correction path; create no prepared record.
P10: processing timeout Technical hold Preserve the case and recover without duplicate output.

Result record: case ID; run ID; configuration version; expected state; actual state; field evidence; downstream events; operator action; pass/fail; defect owner; retest result.

Completion check: every case has evidence, routing and action results. Resolve unexpected ready states and unsupported accepted values before sign-off.

7. Evaluate failures by consequence, then hand over

Judge false acceptance separately from unnecessary review, because the two failures need different remedies.

Report results by damage type and required field, not just one overall extraction score. A workflow can read most text while still accepting an unsupported total. Count state mismatches, unsupported accepted fields, readable cases sent unnecessarily to resubmission, duplicate outputs and unclear operator instructions.

As a proposed release rule, treat any required value accepted without readable evidence as a blocker for this preparation workflow. Also treat unexpected output from a held case as a blocker. Other tolerances, such as how much unnecessary review is acceptable, need explicit agreement from the process owner rather than a universal percentage.

After changing a prompt, processor, preprocessing step or routing rule, rerun both the damaged cases and readable controls. A repair that sends every document to review is not a complete solution.

Give operators a runbook containing state definitions, evidence inspection steps, replacement-file handling, escalation contacts and override limits. Broader AI automation planning should include this handover. If discussing agent-led routing, use the custom AI agents resource without assuming an agent can authorise exceptions.

Worked walkthrough: readable, cropped and duplicate invoices

A hypothetical accounts-payable team tests a synthetic invoice with reference INV-041 and total R2,300. Both values are invented for this example, and the task is preparation for human review, not payment.

Normal case: The readable scan contains the reference and total. The tester compares extracted candidates with their visible locations. Under the proposed rules, the record reaches ready for preparation. Accounts payable still performs its separate checks before any payment decision.

Missing case: A copy cuts off the total. The extractor returns R2,300, perhaps as a plausible candidate. Even though it matches the answer sheet, the test fails if the workflow accepts it: the submitted image does not support that value. Expected handling is an unknown total, resubmission required, and no prepared record. The operator requests a complete page rather than typing the answer from memory.

Ambiguous case: Another copy makes the final reference characters indistinct. The reviewer inspects the image. If the characters cannot be resolved, they request a replacement; matching a familiar reference is not sufficient evidence.

Duplicate case: A second upload contains the same readable invoice under a new filename. The proposed workflow holds it for review and links the earlier case. A human determines whether it is a duplicate, a replacement or a legitimate separate submission before further preparation.

FAQs about damaged-scan testing

Use these answers to resolve specific test exceptions without weakening the evidence requirement. The AI automation glossary provides a shared starting point for terminology across technical and operational teams.

Should every rotated scan be rejected?

No. The proposed rule is to accept a rotated scan for preparation if the pipeline recovers orientation, retains required content and produces traceable, readable field evidence. Test the actual correction path. If the team has only demonstrated a manually repaired file, it has not demonstrated handling of the original rotated submission. Where recovery fails, route to review or request a correctly oriented replacement according to the agreed procedure.

Can a reviewer fill a cropped total from arithmetic?

Not as though it was extracted from the scan. A calculation may produce a separate candidate, but it does not restore the missing printed total. The proposed handling is to request the complete document. If the business permits a specialist to use other evidence, record that evidence and the authorised decision separately. Financial and tax consequences require appropriate human judgement, not an extraction-system assumption.

What if confidence is high but the digits are unreadable?

Treat the field as unresolved under the proposed evidence rule. Confidence is not permission to accept text that an operator cannot verify in the submitted source. Preserve the candidate for inspection, flag the ambiguity and block the affected preparation step. Then test whether this behaviour holds across similar damaged cases. Do not introduce a universal confidence threshold without evaluating the chosen processor, document type and operational consequences.

If your business needs help defining these acceptance states and operator checks, get in touch to discuss a scoped poor-input test pack before connecting document extraction to live business actions.

Sources

Share this article

Bukhosi Moyo

Written by

Bukhosi Moyo

CEO & Founder

Bukhosi is the founder and lead SEO strategist at Symaxx. He architects search-first digital systems for South African businesses, combining technical engineering with commercial strategy to build long-term organic assets.

Feedback

Was this helpful?

Tell us how this article felt in one click.

Back to Insights

Need help executing this strategy?

Our team turns these insights into revenue-generating search architectures for your business.