How do we preserve the original document after staff correct an AI extraction?

Keep original files, AI outputs and staff corrections linked with a reusable correction-history model, clear version rules and practical exception handling.

AI Automation
6 October 2026Updated 06 Oct 202610 min readBukhosi Moyo

Quick Answer

Preserve the original file separately from the extracted data, and record staff corrections as new events rather than overwriting the AI output. Link each correction to its extraction run, document version, reviewer and supporting page or passage. Build the current view from that history. If a supplier sends a replacement document, store it as a new source version, not as a correction to the original.

Key Takeaways

  • Keep original files separate from corrected field values.
  • Record corrections with before-values, after-values, reviewers, reasons and evidence locations.
  • Link replacement documents and extraction reruns without erasing earlier records.
  • Resolve missing evidence and conflicting edits through accountable human review.

Want the full breakdown? Scroll below.

Person planning a workflow on a whiteboard
On this pageJump to a section
  1. 11. Separate the source file from the working data
  2. 22. Record every extraction attempt before review
  3. 33. Capture corrections as accountable events
  4. 44. Use this correction-history data model
  5. 55. Keep the current view separate from the history
  6. 66. Protect originals without promising permanent retention
  7. 77. Route exceptions through a controlled save process
  8. 8Worked example: a corrected invoice and a replacement
  9. 9Evaluate whether the history can actually be reconstructed
  10. 10FAQs
  11. 11Sources

Share this article

Bukhosi Moyo

Growth Partner

Need help growing your company?

We build SEO-first websites and growth systems for South African businesses.

Get Started

Preserve the original document as a separate, protected file, and save staff corrections as linked records rather than edits to that file. Keep the first AI extraction too. A later reviewer should be able to see what arrived, what the AI read, what staff changed and which version was used downstream.

The proposed workflow below separates those layers. It is suitable for planning a document processing system where corrected data must remain traceable to its source. It does not establish that a document is authentic or that a corrected value is legally valid.

1. Separate the source file from the working data

Keep the source file unchanged, even when its extracted fields are wrong. Editing a captured invoice number should change the working data, not the uploaded PDF.

Use separate records for the document family, each received file version, each extraction run and each correction. A document family groups related submissions. A source version identifies one particular received file. An extraction run records one attempt to read that version.

Store the original bytes under a unique storage key. Record the received time, original filename, file type, byte size and a cryptographic hash. The hash can help identify byte-for-byte changes; it does not prove who created the document or whether its contents are true.

Keep OCR text, rotated images and compressed previews as derived files linked to the source. Never label a cleaned preview as the original. If processing requires a converted PDF, retain both the incoming file and the conversion relationship.

This separation gives the team a clear answer to a dispute: “This is the file we received, and these are the changes made to its captured data.”

2. Record every extraction attempt before review

Save the initial extraction before anyone corrects it. A corrected field alone cannot explain whether the error came from the document, the extraction or a later edit.

For each run, record the source-version ID, start and completion times, extractor identifier, configuration version, schema version and raw response location. Save failures and incomplete responses as run outcomes, not as valid extractions with empty fields.

Record both raw text and normalised values where useful. For example, a date printed as 05/10/2026 and a normalised date are different representations. Preserve the printed form so a reviewer can check the interpretation.

Source: OpenAI's Structured Outputs documentation describes schema-constrained responses. A correctly shaped response should still be checked against the document: format is not evidence that a field was read correctly.

In the proposed schema, distinguish missing, unreadable, ambiguous and present. Do not turn all four into an empty string. If the extraction supplies no usable evidence location, record that limitation instead of inventing a page reference.

3. Capture corrections as accountable events

Create a new correction event whenever a reviewer changes a value. Preserve the previous event and the original extracted value.

The correction screen should show the source page alongside the captured field. Require the reviewer to select the field, enter the replacement value, choose a reason and identify the supporting passage. Record reviewer identity from the authenticated session, not from a name typed into a text box.

Separate three proposed correction types:

  • Reading correction: the source says something different from the extraction.
  • Normalisation correction: the reading is right, but the stored format needs changing.
  • External clarification: the replacement comes from another document or communication.

External clarification needs its own evidence link. A telephone explanation should not silently become “what the PDF says”. Record who supplied it and route consequential changes for appropriate human judgement.

Save the correction and its resulting revision in one database transaction. Include an idempotency key so a repeated save request does not create duplicate events. If saving fails, tell the reviewer that the correction was not committed.

4. Use this correction-history data model

Use linked records with explicit keys, rather than one editable row containing the latest values. The following reusable dictionary is a proposed minimum model, not a native feature of any named product.

Proposed correction-history data dictionary

Record Required fields Proposed preservation rule
Document family document_id, business_case_id, created_at Groups related versions; does not identify file contents.
Source version source_version_id, document_id, storage_key, original_name, received_at, sha256, byte_size, supersedes_source_id Preserve original bytes; replacements receive new IDs.
Extraction run run_id, source_version_id, extractor_version, config_version, schema_version, started_at, outcome, raw_output_key Retain every attempt, including failed runs.
Extracted field field_id, run_id, field_path, raw_text, normalised_value, value_state, evidence_locator Preserve initial values; locator may explicitly be unavailable.
Correction event event_id, field_id, base_revision_id, before_value, after_value, correction_type, reason, evidence_reference, reviewer_id, recorded_at, idempotency_key Append events; never replace earlier corrections.
Working revision revision_id, document_id, source_version_id, run_id, parent_revision_id, included_event_ids, review_state, created_at Identifies a reproducible field snapshot.
Export receipt export_id, revision_id, destination, sent_at, result, external_reference Links downstream data to the revision actually sent.

Use server-generated IDs and timestamps. Enforce foreign keys and unique idempotency keys. Store evidence locators as page plus passage or coordinates where available. Represent missing values explicitly. Check that each correction's before-value matches its base revision before committing it.

An implementation may need extra tables for line items, attachments or approval events. Preserve the same relationships. For repeating fields, use a stable line-item ID rather than relying only on an array position that could change after a rerun.

5. Keep the current view separate from the history

Build the current view from a named revision, while keeping earlier revisions available. “Latest” should never be the only way to identify data sent to another system.

Under the proposed rules, an accepted correction creates a child revision. The interface displays that revision's values, but the timeline still exposes the initial extraction and earlier edits. Reversing a correction creates another event; it does not delete the mistaken event.

Handle concurrent edits by checking the base revision. If two reviewers start from the same revision, the second save should not silently overwrite the first. Show the intervening change and ask the reviewer to reconcile it.

An extraction rerun creates a new run, not a replacement for the old run. Compare its output with the current revision before adopting it. Otherwise, rerunning the extractor could undo human corrections.

Keep downstream exports tied to a revision ID. If a later correction affects exported data, notify the accountable owner and record the follow-up. A corrected invoice field must not automatically authorise a payment or change a tax treatment.

6. Protect originals without promising permanent retention

Restrict source-file changes through storage permissions and an agreed retention process. Calling a folder “originals” does not make its contents protected.

Source: Supabase's storage access-control documentation describes policies for storage operations, including the extra permissions needed for overwriting files. It also explains that service keys bypass row-level security. Those capabilities require deliberate application and operational controls.

For this proposed workflow, separate permission to upload a new source from permission to overwrite or delete an existing source. Give reviewers controlled read access and correction rights without ordinary source-overwrite rights. Have a security specialist assess privileged access, backups, logging and recovery.

Do not keep personal or commercially sensitive documents indefinitely by default. Ask the responsible legal, privacy and records-management people to decide retention periods, deletion authority and treatment of disputes. Apply approved retention rules to source files, derived files, history and backups where appropriate.

If an authorised deletion occurs, retain only the deletion record and metadata that the approved policy permits. A history screen must clearly say that the source is no longer available rather than showing a broken link as if evidence still exists.

7. Route exceptions through a controlled save process

Make the application responsible for validating and committing changes. Do not let document text or a model suggestion decide which source records can be changed.

Source: OpenAI's function-calling guide describes model tool requests followed by application-side execution. In this proposed design, application code checks the user's permission, record relationship, revision and required evidence before saving anything. A tool request is not permission to write.

Use these proposed exception rules:

  • Missing evidence: leave the field unresolved and request the missing page or attachment.
  • Ambiguous evidence: preserve the competing interpretations and refer them to the responsible reviewer.
  • Identical upload: link a new receipt event to the existing file where policy permits; do not assume the business submission is redundant.
  • Changed replacement: create a new source version and request comparison.
  • Conflicting edits: block silent replacement and require reconciliation.

A fixed correction workflow may be enough. The AI agents versus automation comparison helps frame that design choice. If considering a custom AI agent, keep its proposed evidence-retrieval role separate from authority to accept consequential corrections.

Worked example: a corrected invoice and a replacement

A normal correction should leave a visible chain from source to export. All identifiers, amounts and events in this example are hypothetical.

A Johannesburg accounts team receives invoice source S-A, linked to document family D-A. Extraction run R-A captures the total as R1,850.00. Reviewer U-A opens the original page and sees R1,350.00.

The reviewer records a reading correction with the before-value, after-value, reason and page passage. The application checks the base revision and creates V-B. The original file and R-A remain unchanged. A later export receipt references V-B; the invoice still requires the team's separate financial review.

Now suppose the bank-detail page is missing. The reviewer leaves those fields unresolved and asks for the attachment rather than copying details from a previous invoice. Any payment-related judgement stays with an authorised person.

The supplier then sends another PDF with the same invoice number but a changed total. Its bytes differ, so the team stores S-B as a linked replacement, not an overwrite of S-A. A reviewer compares both sources and records why one is selected for further processing.

If the supplier resends identical bytes, the team records the new receipt and checks whether it represents a duplicate submission. File equality alone does not settle that business question.

Evaluate whether the history can actually be reconstructed

Evaluate the design by asking someone other than the original reviewer to reconstruct a case. Do this before relying on the history for operational disputes.

Use a proposed evaluation set containing a normal correction, reversal, missing page, conflicting edit, identical upload, changed replacement and extraction rerun. For each case, ask the evaluator to retrieve the original bytes, locate the relevant passage, explain every accepted change and identify the revision exported.

Record failed retrievals, missing relationships, unexplained changes and exports without revision IDs. Set acceptance criteria with the business owner and technical team; do not substitute a reassuring timeline display for checking the underlying records. Also exercise restore procedures and rejected write attempts with the security team.

The custom AI agents workflow guide can support wider workflow planning, but evidence preservation needs explicit storage and database rules regardless of orchestration style.

If your business needs help defining these relationships within an AI automation workflow, get in touch with a sample document and a redacted correction scenario. Start with a reviewable design, not a promise that the history will resolve every dispute.

FAQs

Should staff annotate the original PDF when correcting a field?

Keep annotations in a separate overlay or derived copy. Link them to the source version, page and correction event. If staff must produce a marked-up PDF for another team, label it as an annotated derivative and retain the original separately under the approved retention policy. An annotation can explain a reading, but it should not replace the structured record of who changed which field and why.

What happens when a reviewer corrects the wrong value?

Record a reversal or superseding correction against the current revision. Preserve the mistaken correction, its author and its reason. If that revision was exported, identify the affected destination through the export receipt and ask its accountable owner how to correct it. Do not assume changing the capture screen repairs another system or reverses a consequential action already taken.

Can a new extraction run replace all staff corrections?

Not silently. Preserve the new run and compare it with the current working revision. Present changed fields, retained corrections and unresolved evidence for review. The responsible person should decide which values enter a new revision. Keep the old run available so someone can distinguish an extractor change from a reviewer change. Use stable field and line-item identifiers to make that comparison meaningful.

Sources

Share this article

Bukhosi Moyo

Written by

Bukhosi Moyo

CEO & Founder

Bukhosi is the founder and lead SEO strategist at Symaxx. He architects search-first digital systems for South African businesses, combining technical engineering with commercial strategy to build long-term organic assets.

Feedback

Was this helpful?

Tell us how this article felt in one click.

Back to Insights

Need help executing this strategy?

Our team turns these insights into revenue-generating search architectures for your business.