Document automation should treat a PDF containing several different forms as a document pack. Classify pages, group them into separate document instances, preserve original page references and route each type to its own reviewer. Split the file only when the document boundaries are clear, and hold uncertain cases for human review.
The proposed workflow below separates identifying a document from extracting its fields. That matters because a correctly read value can still belong to the wrong form. It is a design for teams considering document processing, not a claim that any particular platform supplies the entire workflow automatically.
1. Register the original pack before processing it
Keep an unchanged original and create a pack record before classifying or splitting anything. This gives reviewers a stable reference when derived files, page order or extracted values change.
The proposed pack record should contain a unique pack identifier, original filename, receipt time, file checksum, page count and storage reference. Record who submitted it and which business process it belongs to, without assuming that every page belongs to that person or organisation.
Create working copies for rotation, image enhancement or OCR. Never overwrite the original with an improved scan. A reviewer may need to distinguish unreadable source material from an extraction problem.
Use physical PDF page positions as the primary reference. Printed page numbers are separate evidence: a cover sheet may have none, and several forms may each start at page one.
Before uploading sensitive packs, ask the responsible security and information-governance people to assess access, retention and provider handling. A supplier pack containing banking details needs different access decisions from a public application form. Processing convenience should not determine those permissions.
2. Classify pages before deciding document boundaries
Identify likely document types first, but do not assume that a page label proves where a document starts or ends. Page classification and document grouping answer different questions.
Source: Azure Document Intelligence's overview describes custom classifiers that identify designated document types before extraction. Source: Google Document AI's overview describes classification and splitting, including PDFs containing multiple documents. These capabilities support this design, but do not establish its accuracy on your packs.
Define an explicit vocabulary with your reviewers: invoice, delivery note, supplier form, supporting attachment and unknown, for example. Include cover sheets and blank pages rather than forcing every page into a business form category.
For each page, retain the proposed label, any model confidence supplied, and the visible evidence behind the label. Useful evidence might include a heading, document reference or recurring table layout.
A proposed operating rule is to send conflicting labels to a boundary reviewer. Do not let a filename such as “supplier form” override an invoice visibly embedded in the same PDF.
3. Group pages into document instances
Group pages using evidence of continuity, not merely matching document types. Two adjacent invoices are still two documents, while one invoice can span several pages.
Look for document references, party names, continuation headings, printed page sequences and repeated headers. Treat these as evidence rather than absolute rules. A changed reference may indicate a new document, a correction or a supporting attachment.
Assign each proposed group a document identifier and an ordered list of original page positions. Lists are safer than ranges when pages arrive out of order or a supporting sheet interrupts a form. Record the reason for each boundary so a reviewer can challenge it.
Keep logical groups without physically splitting when boundaries are uncertain, documents share attachments, or the reviewer needs the whole pack. Create separate PDFs when an extraction step or review queue requires them and the grouping has been accepted.
If two forms appear on the same scanned page, a page-level split is insufficient. A proposed exception route is manual region selection, retaining both the original page reference and crop location. Do not pretend the crop was a separate original page.
4. Preserve evidence through extraction
Every extracted field should point back to its document instance and original page, not just to a newly created PDF. Otherwise, splitting can make evidence harder to find.
Keep raw text alongside any normalised value. For a date, retain the printed text and the interpreted date separately. For an amount, retain the currency evidence rather than assuming rand because the business operates in South Africa.
The proposed field record contains the field name, raw text, normalised value, document identifier, original page position, location if available, extraction status and human correction history. Use null or an explicit missing status when the source does not support a value.
If an LLM prepares structured records, Source: OpenAI's structured outputs guide describes schema-constrained output and exceptions such as refusals or incomplete responses. A correctly shaped record is not proof that its contents match the PDF. Validate references and inspect consequential fields.
Keep any processor or model version with the processing record. Before selecting a named product, verify current processor availability, region support and account requirements. The supplied overviews are not a complete deployment checklist.
5. Route documents without granting decision authority
Route each accepted document instance to a named review role, while keeping business approval outside the classification step. Recognising an invoice is not permission to pay it.
Create a proposed routing table with document type, reviewer role, backup role and escalation condition. An accounts-payable reviewer might check invoices, an operations reviewer might check delivery notes, and a supplier administrator might check onboarding forms. Assign unknown documents to pack triage rather than the most convenient queue.
Send reviewers the document pages, extracted fields, source links and exception notes. Avoid giving every reviewer the entire sensitive pack by default; authorised pack-level access should be a separate decision.
Keep routing acknowledgement distinct from review completion. A queue accepting a document does not mean someone has checked it. Also define who owns the overall pack when several reviewers handle its parts.
The AI agents versus automation comparison can help frame whether variable interpretation needs an agent or whether fixed routing rules are sufficient. Payment, tax, employment and legal consequences still require the appropriate responsible person.
6. Hold exceptions at the smallest safe scope
Hold the affected document when an exception is local, and hold the pack when the exception undermines its grouping or completeness. This proposed distinction avoids treating every problem as either harmless or a total failure.
A missing signature on a clearly bounded form may be local. An unidentified page between two similar forms may affect both boundaries. A truncated upload may affect the entire pack.
Use separate exception reasons for unknown type, unclear boundary, unreadable page, missing continuation, suspected duplicate and conflicting field values. Each reason should specify a reviewer action, not merely display a warning.
For suspected duplicates, preserve both occurrences and their original positions. A checksum can help identify identical files or pages, while matching references can flag candidates. Neither establishes whether a document should be ignored: a revised invoice may retain its reference while changing its content.
If orchestration uses tools, the application's permissions must control which actions are available. The custom AI agent glossary entry provides terminology, but an agent label does not replace access controls or human judgement about security-sensitive actions.
7. Evaluate boundaries, evidence and routing separately
Evaluate the proposed process against human-labelled packs before widening its use. Field accuracy alone cannot show whether the system joined two forms incorrectly or sent a correct record to the wrong reviewer.
Build a representative evaluation set with reviewer-agreed page groups, document types, expected fields and exception outcomes. Include different layouts, poor scans, repeated references, blank sheets, interrupted forms and revised documents. Keep evaluation packs separate from material used to tune the classification approach.
Measure correct document boundaries, correct type assignments, source-reference accuracy, field accuracy and routing accuracy separately. Count false merges and false splits explicitly. Record reviewer corrections and unresolved exceptions rather than hiding them inside an overall success percentage.
Set acceptance criteria with the business owner and relevant risk specialists. No universal confidence threshold is proposed here: a score's meaning needs evaluation for the selected model and document family.
During an initial controlled rollout, a proposed rule is human confirmation of every group. Broader AI automation should follow evidence from that review, not vendor capability descriptions alone. Potential reductions in mixed-pack errors must be demonstrated on the team's own work.
Reusable mixed-pack handling checklist
Use this proposed checklist for each pack. It is complete only when every page has a recorded disposition and every exception has an owner.
| Stage | Required action | Hold condition or owner |
|---|---|---|
| Intake | Store unchanged original; record pack ID, checksum and page count. | Unreadable or incomplete file: intake owner. |
| Page inventory | Assign every original page a label, including blank, cover or unknown. | Uncertain type: pack triage. |
| Grouping | Assign document IDs and ordered original-page lists; record boundary evidence. | Conflicting continuity: boundary reviewer. |
| Splitting | Create derived files only from accepted groups; retain page maps. | Shared page or attachment: manual decision. |
| Extraction | Store raw and normalised values with original-page evidence; mark missing fields. | Unsupported or conflicting value: type reviewer. |
| Duplicate check | Link suspected copies; preserve both occurrences. | Deletion or exclusion decision: authorised reviewer. |
| Routing | Assign type reviewer, backup and pack owner; confirm queue receipt. | No suitable reviewer: pack owner. |
| Closure | Record corrections, outstanding requests and final page dispositions. | Unresolved grouping or missing pages: keep affected scope open. |
Completion check: reconcile the page inventory to the original PDF, verify source links, and record human decisions before releasing reviewed data to its next authorised step.
Worked walkthrough: a normal pack and an uncertain continuation
A clear pack can proceed as separate review tasks; an uncertain continuation should remain unresolved until a person checks it. All identifiers, page counts and contents in this walkthrough are hypothetical, and the handling rules are proposed.
Suppose pack PACK-A contains six pages. Original pages one and two contain invoice INV-A; page three contains a delivery note; pages four and five contain a supplier form; page six is blank. The page inventory accounts for all six pages, including the blank one.
The system proposes three document instances and records the invoice's original-page list as [1, 2]. A reviewer confirms the boundaries. The invoice goes to accounts payable, the delivery note to operations and the supplier form to supplier administration. Any generated invoice PDF still maps its local page two to original page two. These tasks prepare review; none approves payment or validates banking details.
Now suppose page five instead says “page 3 of 3”, while page four says “page 1 of 3”. The system marks a missing continuation rather than filling the gap from another form. The supplier reviewer checks the original and requests the missing page if needed. The pack owner decides whether unrelated review tasks may continue.
If page six repeats the invoice's first page, it becomes a suspected duplicate, not discarded evidence. A reviewer compares content and records whether it is a redundant scan or part of a different invoice instance.
FAQs about mixed-form PDF processing
Resolve these questions through explicit workflow rules rather than assumptions about the extraction tool.
Should we split the PDF before running OCR?
Not necessarily. In this proposed design, obtain enough text and layout evidence to identify boundaries before creating separate files. If the pack already contains reliable separators, a reviewer may confirm those groups first. When a processor requires separate inputs, use accepted groups and retain their page maps. OCR on the whole pack should not imply that all its fields belong to one document.
What if an attachment supports more than one form?
Keep one original attachment reference and link it to each relevant document instance. Do not silently assign it to whichever form appears immediately before it. A reviewer should decide whether the attachment is shared evidence, a separate document or part of a single form. If separate review copies are needed, make the shared relationship visible so reviewers do not mistake copies for distinct submissions.
Can a high confidence score bypass boundary review?
Only after the team has evaluated a proposed rule for that document family and agreed which risks remain acceptable. A type score does not necessarily express confidence in the boundary or completeness of the form. Keep consequential decisions with responsible reviewers, and continue sampling accepted groups. A new layout or unexplained continuation should trigger review even when the type prediction looks confident.
Choose a workflow that can explain every page
Choose the simplest design that can account for every page, explain its groupings and show reviewers the original evidence. A sophisticated extractor is not enough if the team cannot trace a field back to the correct form.
For orchestration planning, the custom AI agents workflow resource offers a related starting point. If your business receives mixed document packs and needs help defining boundaries, page maps and reviewer queues, get in touch with Symaxx to discuss a scoped document-processing workflow and evaluation plan.

