Verify preservation by treating references as text and comparing their exact characters from the original document through to the destination record. Include spreadsheet exports and imports in the test. A correct extraction is insufficient if a later step changes 000742 to 742, removes a slash or attaches the reference to the wrong customer.
For workflow automation, acceptance should depend on evidence from the whole route. The proposed test set below gives an operator clear expected results, including when to stop and ask a person.
Define what must stay identical
Write a short field agreement before testing. Identify the invoice reference and customer reference separately, specify where each appears on the source, and name its destination field. Avoid a single vague field called “number” when documents contain several possible references.
For each identifier, preserve leading zeros, letters, punctuation and significant spaces. Do not assume a fixed length because several samples look alike. In these hypothetical examples, 000742, 0742 and 742 are distinct identifiers unless an authoritative business rule establishes otherwise.
Keep the captured original separate from any approved matching variant. If staff need a search value without spaces, store that as an additional value with a documented transformation. It must not replace the reference supplied by the customer or invoice issuer.
Also define the scope of uniqueness. An invoice reference may need its supplier identity alongside it for matching. Ask the process owner which combination identifies a record; do not infer uniqueness from the reference alone.
Make extraction return identifiers as text
Where AI extracts document fields, specify reference fields as strings and provide a separate status for missing or uncertain readings. OpenAI’s Structured Outputs documentation describes schema adherence, including prescribed field types. This supports a text-shaped result, but does not establish that the extracted characters match the document. Source: OpenAI Structured Outputs
A string containing 742 still fails when the source says 000742. Therefore, test both type and content. Retain the source document and page or field location so a reviewer can resolve disagreements without searching the entire batch.
Use ordinary comparison logic for exact equality; reserve human judgement for unclear evidence. A custom AI agent can assist with interpretation, but the acceptance check should not ask another model whether two clearly available strings are “close enough”.
Our AI agents versus automation comparison helps frame this choice: use a defined check where the rule is exact, and reviewed interpretation where the source is uncertain.
Check every spreadsheet and import boundary
Draw the actual route on one line: document → extracted record → spreadsheet → exported file → imported record. Remove unused stages and add any manual copy-and-paste step. Each arrow needs a comparison because approval at one stage says nothing about the next.
For spreadsheets, ask the implementer to demonstrate both the stored value and its displayed appearance. Your test should reject a numeric value merely displayed with extra zeros when the agreed reference field requires text. Inspect the exported file as well as reopening it through the operator’s normal import process.
Google’s documentation distinguishes cell value data from spreadsheet formatting and describes batch operations that can change values, formats and validation. That distinction is why a formatting screenshot alone is weak acceptance evidence. Source: Google Sheets batch updates
Test the precise file format and import settings staff will use. Do not assume a different spreadsheet application, connector or import wizard will interpret the same file identically. Include a save, close and reopen cycle rather than inspecting only the freshly generated sheet.
Where an agent requests an import through function calling, the application executes the requested function. Place the proposed reference checks before that write and verify the saved result afterwards; these checks are application work, not an automatic consequence of tool calling. Source: OpenAI function calling
Reusable identifier-preservation test set
All identifiers below are hypothetical. These are proposed acceptance rules; confirm field-specific exceptions with the process owner before running the set.
| Case | Source input | Expected result | Human handling |
|---|---|---|---|
| Ordinary invoice | 000742 |
Exact text 000742 at every stage |
Accept after destination readback |
| Customer reference | 00128 |
Exact text 00128 on the intended customer |
Confirm customer association |
| Punctuation | INV/000742-A |
Preserve slash, hyphen, letters and zeros | Reject any unapproved change |
| Different lengths | 00742 and 000742 |
Keep as distinct values | Do not pad either value |
| Significant space | AC 00128 |
Preserve the internal space | Change only under an approved field rule |
| Missing reference | Blank source field | Missing status; no invented identifier | Request the reference from its owner |
| Ambiguous scan | Unclear 00O742 or 000742 |
Uncertain status; hold import | Inspect clearer evidence; record resolution |
| Repeated submission | Same supplier and 000742 twice |
Flag potential duplicate; no second write pending review | Confirm whether it is the same document |
| Shared reference | Different suppliers both use 000742 |
Preserve both; match using approved supplier context | Verify separate records are intended |
| Long identifier | 0001234567890123456789 |
Preserve every character as text | Reject rounding, shortening or substitution |
For each case, record: source file and location; field name; expected text; extracted text; spreadsheet stored value and type; exported text; destination stored value; destination record identity; pass, fail or held status; reviewer and resolution.
Proposed completion rule: every valid case matches exactly through destination readback, every exception follows its stated hold or review path, and no unexplained reference change or unintended duplicate remains.
Work through ordinary and difficult results
Consider a hypothetical distributor receiving an invoice with reference INV/000742-A and customer reference 00128. The operator checks both source locations, then follows the values through extraction, the working spreadsheet and export. After a test import, they open the destination record and verify both stored references and the customer association. Matching text on the wrong customer fails.
Now suppose the extracted customer reference is correct, but the export contains 128. Record the export boundary as the first failure. Hold that row, investigate the conversion and regenerate it from the preserved original. Do not repair it by adding zeros to every short reference: the different-length test deliberately checks that mistake.
For a missing reference, the reviewer should see the blank source field, rather than a plausible identifier generated from the filename. They request confirmation from the document owner and record the supplied evidence. Any temporary tracking label must remain separate from the business reference.
For the ambiguous scan in the table, a clearer document or authoritative record may resolve whether the character is zero or the letter O. If neither does, leave the item held. A reviewer’s approval should document the evidence used, rather than turn an unresolved guess into an accepted value.
Finally, two submissions with the same supplier and reference require duplicate review. Two different suppliers using the same reference require contextual matching. The operator checks document identity and the agreed uniqueness rule before deciding whether either record should proceed. Neither outcome authorises payment.
Turn the results into an operator handover
Ask the supplier or implementer to deliver the completed test record with the source samples, actual exports, destination readbacks and import configuration used. Record workflow and template versions so a later failure can be compared with a known run.
Give the operator a practical failure instruction: hold affected records, retain their originals, identify the first changed value and send that evidence to the named maintainer. After a correction, rerun the failed case and the remaining test set. Check existing destination records before retrying an interrupted import.
Repeat this acceptance exercise when extraction rules, spreadsheet templates, export settings or destination mappings change. Broader custom AI agent workflows should include this same evidence trail wherever references cross systems.
Before deployment, check current product support and account, plan and region eligibility for the chosen integrations. The proposed test proves only the cases and route actually exercised.
FAQ: testing reference preservation
Can we simply add zeros back after import?
Only where an authoritative field rule defines the exact width and the underlying identity is unambiguous. Otherwise, restore from the verified source and investigate where information was lost. Padding can collapse genuinely different references into the same value.
Is a correct-looking spreadsheet enough evidence?
No. Check its stored value and type, then the export and destination readback. Include the normal reopening and import steps. Keep screenshots as supporting evidence, alongside the actual values and record identities being compared.
What if the destination accepts only numeric references?
Treat that as a mapping conflict before deployment. Ask the system owner whether an appropriate text field or approved mapping exists. Keep the original identifier linked to any internal key, and test retrieval and reconciliation before approving the arrangement.
If your business needs help checking references across an AI automation workflow, get in touch with representative documents and the intended import route. Those examples provide a concrete starting point for agreeing acceptance evidence.

