Yes. Test browser automation on a controlled workflow copy before connecting the real account. Build representative interactive screens, use fictional records and route actions to an isolated test store. Agree what passing means before running the agent, including when it must stop. The resulting pilot should show which behaviours work in the copy and which still need authorised live verification.
Choose one task with a visible finishing point
Start with a task such as finding a service request and preparing an unsubmitted draft response. Define its starting screen, required reference, permitted changes and finishing point. Avoid copying the entire portal when the decision concerns one short journey.
Write the expected outcome in business terms: the correct request is selected, approved information appears in the draft, and nothing is sent. Include the evidence a reviewer needs, such as the selected record reference and the saved draft contents.
Decide whether the task needs an agent at all. A fixed sequence with stable fields may suit ordinary automation. Variable screens or interpretation may justify an agent. The AI agents and automation comparison helps frame that choice; the pilot should test the uncertain part.
Build a copy that reproduces behaviour
Screenshots can test whether the agent recognises a field, but an interactive copy is needed to test navigation and state changes. Represent search results, detail pages, editable fields, validation messages and saved drafts relevant to the task.
Make the copy resettable. Each case should start from a known state, so an earlier draft cannot accidentally help a later run pass. Preserve a separate version for replay tests where the existing result deliberately remains available.
OpenAI's 10 September 2026 Agents API public-beta announcement describes a managed harness and environment options. That dated announcement provides context for pilot infrastructure; it does not establish that a particular copied portal is isolated. Source: OpenAI Agents API announcement
Ask the implementer to demonstrate where browser navigation, downloads and action requests actually go. Remove routes to real messaging, purchasing or record-changing destinations. If an action is simulated, label its outcome as simulated rather than presenting it as a real portal response.
Make fictional cases representative
Invent records from the workflow requirements instead of lightly disguising customer records. Preserve useful formats, such as long references, similar organisation names and empty fields, without importing real identities or unusual personal histories.
Prepare complete, missing, ambiguous and repeated-request cases. Also include a changed button label and an interrupted save where these could affect the proposed task. Give each case an expected result before the agent sees it.
For a South African business, use the date and number formats staff actually expect. If the proposed copy displays day/month/year dates, include a case where interpretation needs confirmation. Test how uncertainty is handled rather than rewarding a plausible guess.
Record what the copy omits: other account roles, attachments, additional languages or alternative layouts. These omissions define the pilot's reach. They should guide later testing without turning a small pilot into an attempted replica of everything.
Check actions against actual test state
Separate what the agent proposes from what the application executes. OpenAI's function-calling documentation explains that application code executes model-requested functions and returns their outputs. In this proposed workflow, the implementer must route those requests to controlled test actions. Source: OpenAI function calling
Inspect the saved test record as well as the agent's final message. A statement that a draft exists is insufficient if the store contains no draft, the wrong reference or an extra record.
Structured Outputs can constrain a result record to an agreed schema. Use fields for the observed reference, disposition and evidence location, then verify their contents independently. Consistent formatting cannot establish correct field selection. Source: OpenAI Structured Outputs
Where a tool needs human approval, n8n documents a review step that pauses execution and allows approval or denial. If that approach is chosen, test denial and confirm no write occurs. It is a configured workflow pattern, not an automatic safeguard for every browser action. Source: n8n human review for AI tool calls
Use this bounded pilot brief
Complete the brief before running cases. Roles may belong to the same named person, but reconciliation and live-scope decisions must have explicit owners. The acceptance rules below are proposed rules for this pilot.
Bounded browser-automation pilot
Workflow/version: ______ | Fixture environment/version: ______
Test owner: ______ | Action reviewer: ______ | Reconciliation owner: ______ | Live-scope approver: ______
Task and finishing point: ______ | Permitted test actions: ______ | Forbidden actions: ______
| Case or boundary | Proposed acceptance rule | Required evidence |
|---|---|---|
| Isolation | Actions reach only approved test destinations; no production credentials are loaded. | Checked destinations, access configuration and test-store identity. |
| Ordinary task | Correct target and permitted result; no action beyond the finishing point. | Selected reference, action inputs and actual saved state. |
| Missing or ambiguous target | Hold for clarification; no guessed selection or write. | Missing field or candidate references and clarification request. |
| Changed screen | Reinspect within scope or stop when required evidence is unavailable. | Observed difference, disposition and action record. |
| Uncertain save | Reconcile existing state before another write; hold if unresolved. | Attempt record, saved-state inspection and reconciliation decision. |
| Repeated request | Identify the existing result without creating another record; hold if identity or state cannot be reconciled. | Request reference, existing result reference and before/after record comparison. |
| Unauthorised extra write | Fail the fixture, including an additional write during replay. | Attempt history, additional record or modification, and defect assignment. |
Per-case record: case/version; starting state; expected result; observed fields; proposed and executed actions; destination; resulting state; evidence location; pass/fail/hold; reviewer; defect owner.
Reconciliation handoff: send the request reference, attempted action, available state evidence and unresolved question to the named reconciliation owner. Record their decision before resuming.
Live-scope handoff: send case results, unresolved defects, omitted behaviours, isolation evidence and proposed live access/action limits to the named live-scope approver. Record their decision separately from fixture acceptance.
Completion rule: every case has a reviewed result; replay has no unauthorised additional write; unresolved states are held; both handoffs have named owners and evidence. Fixture completion does not authorise connection to production.
Work through ordinary, missing and duplicate cases
Consider a hypothetical service-request portal. Every reference and quantity in this example is invented. Request TEST-104 asks for a revised appointment window, and the approved task is to prepare a draft acknowledgement without sending it.
In the ordinary case, the copied search returns one exact reference. The agent opens it, checks the organisation and request details, then saves the draft. The action reviewer compares its text with the brief and checks the test store. Passing requires the expected draft and no sent message.
In the missing case, the request reference is absent. The organisation name appears on several rows. The agent should hold and identify the missing reference. A staff member supplies it or decides the request cannot proceed; selecting the first row fails the case.
In the ambiguous case, two rows share a reference but show different branches. The agent presents both candidates. The reviewer checks an agreed distinguishing field and records the selection. Without that evidence, the hold remains.
For replay, run TEST-104 again with its accepted draft still present. The expected outcome is to identify that draft without creating another record. If its relationship to the request cannot be established, hand the evidence to the reconciliation owner and hold. Any unauthorised additional write fails, even if both drafts look correct.
Finally, simulate a save that creates a draft but returns no confirmation. The agent must inspect state before retrying. The reconciliation owner reviews the request and draft references; absent confirmation alone is not permission to create again.
List the checks that still need the real account
The copy leaves authentication, account permissions and current production state unverified. It may also omit session expiry, delayed processing, competing staff edits and differences between account roles. List each relevant gap alongside the observation needed to close it.
Before deployment, check current product account, plan and region eligibility, plus the portal owner's permission and applicable use terms. September announcements and copied screens cannot establish these conditions.
A later live evaluation should specify the authorised account, permitted records, stopping point and named observer. Start with the narrowest useful observation. If draft creation is proposed, establish how its real effects will be checked before approving that step.
FAQ about testing a copied browser workflow
Can we start with screenshots only?
Yes, for limited visual interpretation. Move to interactive screens when testing clicks, validation, saving or replay. Record which behaviours screenshots cannot demonstrate.
What should happen when the same request is run twice?
The proposed rule is to identify the existing result without another record. Hold if identity or state is unclear; fail any unauthorised additional write.
Does passing the copy mean we can connect production?
It supports a separate live-scope decision. The approver still needs the remaining gaps, proposed access and action limits, and evidence required from the real portal.
If your business needs help defining this pilot, explore Custom AI agents and AI automation. The custom-agent workflow guide and custom AI agent glossary provide further context. To discuss a bounded test for your portal, get in touch.

