Start with Gemini 3.8 Live as the candidate for simple status enquiries, and Extended Thinking as the candidate for bookings involving several connected decisions. Then test both on the same workflow. Choose the model that reaches the correct outcome within an acceptable waiting time, including when information is missing or a caller changes their mind.
For a business considering AI chatbots, the useful comparison is a working call with a verified result. A fluent demonstration cannot establish whether your booking records, tool responses and approval process will work together.
What the September announcement establishes
Google’s announcement, published on 15 September and updated on 17 September 2026, describes Gemini 3.8 Live as designed for scale and cost efficiency, and Extended Thinking as designed for complex tasks and multi-step reasoning. It also describes background tool execution and simultaneous reasoning and speech. These are dated product claims, rather than evidence that either model will handle your calls correctly. Source: Google’s Gemini 3.8 Live announcement
That distinction gives you a starting hypothesis: try Live where the task is a bounded lookup, and Extended Thinking where several constraints must be reconciled. It does not settle response time, cost or accuracy in your implementation.
Before committing, check the exact model’s current availability, account eligibility, plan, region and supported integration route. Access through a consumer application does not establish access for your business voice integration.
Separate status enquiries from booking decisions
A status enquiry might ask, “Has my service appointment been confirmed?” The proposed workflow identifies the caller using your approved process, retrieves the matching record and reads the recorded status. Success means reporting that record accurately without changing it.
A complex booking might ask, “Move my appointment to next week, keep the same technician, and make it after lunch unless Thursday morning is the only option.” Success requires resolving dates, checking availability, preserving preferences and confirming the chosen option before a change is approved.
Write these task boundaries down before testing. Otherwise, a model can appear successful by quietly dropping a condition. Record whether each preference is compulsory or flexible: “same technician” and “after lunch if possible” require different handling.
The distinction between AI agents and automation helps here. Fixed rules can govern permitted booking changes, while the conversational component interprets the request and asks for clarification. More reasoning should not replace a clear booking policy.
Measure useful response time and tool correctness
Use separate timings for the first acknowledgement, the first useful answer and the verified outcome. “I’m checking” may make a pause easier to understand, but it does not answer the status question or confirm a booking.
For each call, inspect the requested action, supplied fields, returned record and spoken answer. A correct lookup followed by the wrong spoken date is a failure. So is a convincing booking confirmation when the booking system returned an error.
Tool execution needs its own controls. OpenAI’s function-calling documentation describes a flow in which the model requests a tool and application code executes it. That is useful background for designing this pilot, not documentation of Gemini’s specific interface. Your developer must verify the Gemini integration separately. Source: OpenAI function-calling guide
Keep the pilot’s business rules identical across both models. Use the same records, available actions and approval requirements. Include interruptions, slower tool responses and the actual audio channel, so you compare complete workflows rather than different demonstrations.
Copy this task-based voice pilot brief
The following brief is a proposed operating plan. Its release gates are business choices to agree before testing, not vendor guarantees or regulatory requirements.
Task-based voice model pilot brief
- Decision: Select Live, Extended Thinking, or neither for each task category.
- Access check: Confirm the exact model, account, plan, region and integration route before testing; repeat before deployment.
- Task A , status: Retrieve an authorised appointment record and report its status without changing it.
- Task B , booking: Prepare an appointment change that respects the caller’s confirmed date, technician and time constraints.
- Shared setup: Give both models the same approved instructions, test records, tools, audio channel and business rules.
- Test cases: Include an ordinary enquiry, constrained booking, missing identifier, ambiguous date, duplicate request, interruption and tool failure.
- Expected answer: Write the correct record, clarification or handoff for each case before running it.
- Timing record: Capture acknowledgement time, first useful answer time and time to a verified outcome separately.
- Correctness record: Check caller authority, record selection, requested action, field values, tool result and spoken summary.
- Booking control: Require caller confirmation of the final details and staff approval before the pilot changes a booking.
- Failure control: Do not claim completion after an error or uncertain result; check the record before any retry.
- Evidence sheet: Record task, model, run, expected outcome, actual outcome, timings, error, human handling and final record state.
- Proposed release gate: Resolve every observed unauthorised change, duplicate booking and false confirmation before live use; agree task-specific waiting-time limits.
- Final recommendation: State the selected model per task, unresolved failures, measured operating cost and responsible business approver.
Use staff role-play and test records first. A pilot can test the conversation and prepare proposed actions without touching customer appointments. Keep only the call evidence needed for review under your approved recording and retention arrangements.
Work through ordinary and difficult calls
All identifiers, dates and circumstances below are hypothetical. They illustrate expected handling rather than measured model results.
Ordinary status enquiry: A caller supplies appointment reference EXAMPLE-A. The authorised lookup returns “confirmed” for 14 October at 14:00. Either model should report those details accurately and make no change. The reviewer checks that the spoken answer matches the returned record. If both succeed, compare useful answer time and observed operating cost.
Complex booking: The caller wants EXAMPLE-A moved to the following week, preferably after lunch, with the same technician. The hypothetical availability record offers Tuesday afternoon with a different technician or Thursday morning with the original technician. The model should explain the trade-off and ask which constraint can change. Staff review the final proposal against the caller’s confirmed choice before approving it.
Missing identifier: The caller gives a first name but no appointment reference. The expected outcome is an approved identification step or staff handoff. Neither model should pick a record because it looks likely. A reviewer checks that no unrelated appointment details were disclosed.
Ambiguous date: “Next Friday” could be understood differently. The model should ask for confirmation using the full calendar date, then carry that confirmed date into the booking proposal. Staff should return an unresolved proposal for clarification rather than approving a guessed date.
Duplicate or interrupted request: The booking tool times out, and the caller says, “Try again.” The application should check whether the first action completed before attempting another change. Staff reconcile any uncertain result with the booking system. The correct spoken response states that confirmation is pending until the record is verified.
Score these cases independently. Strong performance on ordinary enquiries must not conceal a duplicate-action failure in booking calls.
Choose from the evidence, then set the human handoff
Choose Live for status enquiries if it reliably retrieves and reports the right record with acceptable waiting times. Choose Extended Thinking for complex bookings only if the pilot demonstrates better handling of connected constraints without unacceptable delays or new action errors. A split choice is reasonable; neither passing is also a valid result.
If both fail on missing identifiers, fix identification or record access first. If both announce success after tool errors, fix completion checks. Those defects may sit in the workflow rather than the model.
Use a consistent evidence sheet, but validate its contents. OpenAI’s Structured Outputs documentation describes schema-constrained output; correct formatting alone should not be treated as proof of a correct booking. This is a separate implementation reference, not a claimed Gemini capability. Source: OpenAI Structured Outputs guide
For human approval, n8n documents selected tool calls pausing until a person approves or denies them. That is one possible integration pattern, subject to compatibility checks. Source: n8n human review for tool calls
A custom AI agent should have explicit task boundaries. Our guide to custom AI agent workflows provides context for that broader design decision.
FAQ: choosing a model for your voice workflow
Should every booking request use Extended Thinking?
No. A single available slot with clear details may be straightforward. Test the tasks with conflicting preferences and dependencies separately before assigning a model.
What if the model speaks quickly but the booking remains pending?
Record the acknowledgement and completion times separately. The caller should hear that the request is pending, with a staff handoff if the result cannot be verified.
Can we choose a model before connecting the booking system?
You can shortlist one using role-play, but final selection needs realistic tool responses and failure cases. A conversation-only trial cannot prove that a booking change completes correctly.
If your business needs help designing this comparison, get in touch about AI automation. Bring your enquiry types, booking rules and expected outcomes so the pilot answers a specific operational decision.
Sources
- Google: Gemini 3.8 Live and Extended Thinking announcement , published 15 September 2026; updated 17 September 2026.
- OpenAI: Function calling , supplied living documentation; publication date unspecified.
- OpenAI: Structured Outputs , supplied living documentation; publication date unspecified.
- n8n: Human review for AI tool calls , supplied living documentation; publication date unspecified.

