Ask the supplier for a traceable demonstration of a job resembling your real process: original inputs, agreed expected outcomes, execution records, destination results, exceptions and operator instructions. Include awkward cases, then let your intended operator handle a stopped item. This gives you evidence for deciding whether to commission a scoped pilot.
A polished response on screen does not show everything you need to assess. Your buying question is whether the proposed workflow automation can complete the agreed task, recognise when it cannot proceed and leave your team able to manage the result.
Define the job before watching the demonstration
Send each supplier the same short process brief. Describe what starts the job, where information arrives, which outcome matters and who currently checks it. Choose a bounded task, such as turning incoming service requests into draft work items for a coordinator.
Agree on expected outcomes before running samples. Specify required information, permitted actions, review points and what counts as an exception. Otherwise, a supplier can explain an unexpected result after the event without showing that it followed an agreed rule.
Ask which steps use AI, which follow fixed rules and which require a person. Our AI agents versus automation comparison can help you frame that discussion. Evaluate the proposed design against the job, rather than assuming that more autonomous behaviour is preferable.
Require a visible chain from input to destination
For each sample, ask to see the source exactly as received, the information extracted, the action requested, the execution result and the resulting destination record. Give the sample a reference that appears throughout the evidence so you can follow it without relying on the presenter’s narration.
This distinction matters where a model requests an external action. OpenAI’s function-calling documentation describes separate stages: the model requests a tool call, application code executes it, and the tool output returns to the model. A requested action therefore needs execution evidence. Source: OpenAI function calling
Have the supplier open the destination and inspect the saved result. Check the selected customer, service location, request text and status against the source. If the demonstration only prepares drafts, inspect the draft and confirm what remains for a person to complete.
Ask the presenter to identify every simulated connection, prepared output or manual intervention. These can be useful in an early demonstration, but record the relevant step as unproven until it runs against a suitable test environment.
Use ordinary and awkward samples together
Consider a hypothetical maintenance business assessing an automation that prepares service work items. The following cases and handling rules are proposed examples, not measured supplier results.
Ordinary request: A customer supplies a recognised account reference, a specific branch and a clear description of a faulty door. Expected handling is to prepare a work item linked to that account and branch, preserving the reported fault. The coordinator checks the draft against the source before releasing it for scheduling. Evidence should include the source, saved draft and review status.
Missing information: The same request omits the branch. Under the proposed rule, the workflow holds the item for clarification and identifies the missing location. The coordinator contacts the requester, records the answer and resumes processing. Watch whether the original request remains attached and whether the corrected item reaches the proper branch.
Ambiguous match: The request names “Central branch”, but the test customer records contain more than one plausible location. Expected handling is to show the candidates and withhold the location-dependent action. The coordinator confirms the correct site with the requester. A plausible guess should fail this case, even if the resulting draft looks tidy.
Duplicate submission: Resubmit the ordinary request with the same source reference. The proposed expectation is that it links to the existing work item without creating another. Then submit similar wording with a different reference: the operator should review whether this is a separate fault. Ask the supplier to explain the duplicate rule and show its effect in the destination.
Retain unsuccessful runs as well as corrected reruns. If the supplier changes a rule during the session, label the new version and rerun the affected samples. This makes it possible to distinguish the original result from the repaired behaviour.
Demonstrate failure and human intervention
Ask the supplier to interrupt a test connection at an agreed safe point. Watch how the workflow reports the failure, what information reaches the operator and whether any earlier action already took effect. The operator should be able to determine whether retrying would repeat completed work.
The test method must match the product being demonstrated. For example, n8n documents that its Error Trigger runs when an automatic workflow errors; it cannot be tested by manually running a workflow. Some execution details also depend on saved execution data or where the failure occurred. Source: n8n Error Trigger
If approval is part of the proposal, demonstrate both approval and denial. n8n documents a human-review pattern in which approval allows a tool to execute with the specified input, while denial cancels the action. That capability still needs to be configured and demonstrated in the supplier’s workflow. Source: n8n human review for AI tool calls
Ask the reviewer to inspect the actual proposed action and its inputs. After denial, check the destination for unintended changes. Agree separately how the proposed workflow should handle unanswered requests; do not assume silence means approval.
Buyer’s demonstration checklist
Copy this checklist into the supplier meeting notes. Use “shown”, “failed” or “not shown” for each item, and attach an evidence reference. These are proposed buying criteria to adapt to your process.
Demonstration record
- Supplier, demonstration date and workflow version: ______
- Process being tested and boundaries: ______
- Buyer’s process owner and intended operator: ______
- Environment, connected systems and simulated steps: ______
- Expected outcomes agreed before the run: ______
Evidence to collect
- Original sample inputs are available and suitable for the demonstration.
- Each sample has a reference connecting its input, run and destination result.
- The ordinary case reaches the agreed destination with correct fields and status.
- Missing information produces the agreed hold or clarification request.
- Ambiguous matches reach a person without an unsupported selection.
- Duplicate input is handled according to the agreed rule.
- A connection failure produces a visible, usable operator notification.
- Approval and denial produce the expected destination outcomes, where applicable.
- Recovery checks for previous effects before repeating an action.
- Manual interventions, failed attempts and configuration changes are recorded.
- The intended operator finds and resolves an exception using the handover instructions.
- Outstanding account, plan, region and integration requirements are listed for current verification before deployment.
Decision record
- Evidence location and access confirmed by buyer: ______
- Unresolved gaps, responsible person and next verification: ______
- Buyer decision: proceed to scoped pilot / request another demonstration / decline.
- Reason, decision owner and date: ______
Make handover part of the buying decision
Ask for a short operator guide tied to the demonstrated workflow. It should explain how to find pending work, interpret an exception, check whether an action completed, correct permitted information, pause processing and seek supplier support.
Then ask your intended operator to use it while the presenter observes. Choose an exception already introduced during the demonstration. Record where the operator needed undocumented help. This exercise reveals whether the handover supports the actual job your team will inherit.
For a proposed custom AI agent, request the same evidence even if its route through the task can vary. The custom AI agents workflow guide provides context for defining that scope.
Use the completed checklist to compare suppliers. A missing screenshot is different from an incorrect destination update: judge gaps by their effect on your process. Proceed to a pilot only with explicit unresolved items, owners and verification steps. A successful demonstration supports that next decision; it does not establish performance across your full workload.
FAQ: assessing supplier demonstration evidence
Can a recorded demonstration be enough to shortlist a supplier?
It can support shortlisting if inputs, execution and destination results are visible. Before commissioning a pilot, request a witnessed run using buyer-selected samples. Record any steps the recording leaves unproven.
What if the supplier cannot connect to our system yet?
Ask for a clearly labelled test substitute and evidence of the intended integration requirements. Treat destination compatibility as outstanding. Before deployment, require a current check of account, plan, region and access eligibility for the proposed products.
Should we reject a supplier when a demonstration fails?
Assess the failure and response. A workflow that holds an incomplete request correctly may be behaving as intended. An unexplained wrong update needs investigation. Require the original evidence, a clear correction and a rerun before reconsidering that case.
If your business is comparing AI automation suppliers, get in touch with the process you want demonstrated and the evidence gaps you need resolved. A bounded brief gives the discussion a concrete starting point.

