A monthly document archive and a customer uploading evidence during a live support conversation have different waiting tolerances. Treating both as immediate can consume scarce interactive capacity; treating both as batch can leave the customer waiting for a task that needs an answer now. The right choice begins with the next action and its deadline.
This article proposes a latency-versus-cost matrix for document processing. It uses the documented OpenAI Batch pattern where asynchronous completion fits, while keeping review and exception handling visible. It does not promise that a provider's processing window is the same as the business workflow's final turnaround.
Define what must be ready and when
Specify the required output: extracted fields, a reviewed document classification, a source-linked summary or an approved downstream action. A model response may be only one stage. Include validation, reviewer availability, source clarification and action verification in the readiness definition.
Separate upload acknowledgement from processing completion. A customer can receive a truthful receipt confirmation while the document remains queued. Do not call that processed if its contents have not been checked. The status should explain the selected path and the next expected business stage without inventing a completion promise.
Measure deadlines against the actual process. If a staff decision depends on a result during the current interaction, a long asynchronous window may be unsuitable even when model pricing is attractive. If the decision is scheduled well after the queue's tested turnaround, batching may be worth evaluating.
Understand the documented Batch boundary
OpenAI's Batch guide describes asynchronous groups of requests, a 24-hour completion window, separate rate-limit capacity and a 50% processing-price discount compared with synchronous APIs. It also documents batch status, output retrieval and individual result identification. These are provider-processing features, not measured savings for your document workflow. Source: OpenAI Batch API
The guide includes an expired state when work does not complete within the window. Design for that and other failure states rather than treating every batch as a fully successful file. Reconcile each input's result and error status before deciding what must be retried.
Confirm that the intended endpoint, model, input format and account configuration support the selected path. This article does not assert that every document-processing service or interactive tool can be placed in the same Batch operation.
Latency-versus-cost decision matrix
Use this complete proposed matrix for a fictional document intake service. Time tolerances are scenario requirements, not claims about observed provider performance.
| Work type and requirement | Candidate path | Required review and recovery | Decision evidence |
|---|---|---|---|
| Customer needs a checked answer during an interaction | Immediate processing or approved staff handover | Validate output and surface uncertainty within the actual interaction | Measured completed-task waiting time; Batch window does not fit this requirement |
| Overnight classification with a later staff review | Batch candidate if the whole process fits the approved deadline | Match every result to input, review exceptions and allow deadline margin | Observed queue, processing, review and recovery times |
| Monthly archive tagging with no immediate dependent action | Batch candidate | Reconcile missing or failed items and preserve source inventory | Total cost per accepted item and completeness evidence |
| Sensitive or ambiguous documents needing specialist review | Either path, chosen from handling and deadline rules | Approved access and processing; designated reviewer queue | Review capacity and handling decision, not price alone |
| Mixed intake containing urgent and deferrable items | Separate approved queues | Prevent bulk work from blocking urgent cases; preserve item identity | Tested classification and workload isolation |
For every candidate, record output definition, latest acceptable ready time, eligible volume, input and model constraints, processing configuration, reviewer owner, exception rules, reconciliation identity and total cost categories. Distinguish known requirements from planning assumptions.
Batch acceptance: each document request has a stable input reference; returned results are matched by that reference rather than output position; successful, failed, missing and expired items have explicit dispositions. Do not reprocess verified successes because another item failed. Before any retry that could trigger an action, check the earlier operation's outcome.
Cost comparison: include model processing, document extraction, storage, orchestration, retries, discarded outputs and review effort. In an entirely hypothetical example, a model-processing component costing R40 synchronously would be R20 under a matching 50% component discount. That arithmetic does not prove a R20 reduction in the total workflow, whose other components may differ.
Acceptance fixtures: a normal complete batch; one failed item; one missing result; expired processing; an unreadable document; a duplicate input; and an urgent item incorrectly routed into bulk work. Inspect input/result matching, reviewer queue and customer status wording. Use stubbed actions so tests do not update real records.
The decision is ready when the owner can show that the selected path meets the actual readiness requirement, handles partial outcomes and has a cost comparison grounded in observed or explicitly hypothetical components. A lower processing price is not sufficient acceptance evidence.
Validate results before releasing them to the next stage
OpenAI's Structured Outputs guide supports constrained formats and refusals. A schema can help validate field shape, but does not prove that the extracted values match the document. Batch and immediate outputs need the same factual and source checks appropriate to their purpose. Source: OpenAI structured outputs
Record refusals and unusable outputs as exceptions, not successful empty results. Keep the original source and input reference available to the authorised reviewer. If an attachment is unreadable, state what was not checked instead of treating the document as fully processed.
A result can be schema-valid while attached to the wrong input. Matching identity is therefore a separate acceptance check. Inspect that mapping before publishing summaries or creating downstream proposals.
Work through normal, missing and duplicate documents
In a hypothetical normal case, a deferrable archive batch returns every item, the application matches the results and validation accepts the fields. The designated team reviews the exception-free output under its approved process. Only the verified stage is marked complete.
In a missing-result case, one input has no reconciled output after the operation ends. The application records it as unresolved and investigates the batch result and error records. It does not shift subsequent outputs by position or assume that an empty value means no relevant information.
In a duplicate case, the same verified document request appears twice. The proposed input-identity rule preserves the source collection records but avoids duplicate downstream work. If the underlying documents differ, they remain separate until the owner resolves their relationship.
Include operational failures in the comparison
Where n8n is part of the implementation, its Error Trigger documentation describes linked workflows for automatic execution failures. It does not replace batch-item reconciliation or establish the same trigger behaviour in manual tests. Configure the intended notification and recovery path explicitly. Source: n8n Error Trigger
A failure alert should identify whether collection, submission, provider processing, retrieval or review failed. Those failures have different owners and retry risks. Do not use one generic failed status to imply that no work completed anywhere in the process.
FAQ about batching uploaded documents
Does the Batch window guarantee the final reviewed output is ready within that time?
No. It describes the provider-processing window and its documented states. Validation, reviewer availability and recovery add separate stages. Plan and test the whole readiness requirement.
Should the whole batch be retried when one document fails?
Reconcile individual outcomes first. Retry only the items and operations permitted by the recovery rule, preserving successful results and checking uncertain actions before another attempt.
Does a 50% processing discount mean the pilot costs half as much?
No. Compare the selected model-processing component and all other workflow costs. The article's R40-to-R20 example is hypothetical component arithmetic, not a measured total-cost saving.
If your business needs help defining this process, explore Document processing, the wider AI automation services, and our custom-agent workflow guide. The agents and automation comparison and custom AI agent glossary explain the terms. To discuss your records and approval rules, get in touch.

