All segments

Voice AI Customer Service: Setup for Shopify Teams

A step-by-step setup for voice AI customer service on Shopify: scope, order data, caller verification, handoff rules, recording and the launch test.

  • Published
  • Reading time 15 min read
  • Author Nafiul Hasan
Voice AI Customer Service: Setup for Shopify Teams. Diagram: work crossing a boundary. AI FOR ECOMMERCE Voice AI Customer Service: Setupfor Shopify Teams YOURSTHEIRS pointerflow.com

Short answer

Voice AI customer service answers phone calls with a synthetic voice that looks up Shopify order data, resolves a short list of routine questions and hands everything else to a person with context. Setup means limiting scope, connecting read-only order data, verifying callers, writing handoff rules, and testing before launch.

What this page says that the ranking pages do not: the setting values to decide before launch, and the one step most teams get wrong, which is writing handoff rules by topic when they should be written by caller state.

Voice AI customer service is a phone agent that answers calls in a synthetic voice, looks up Shopify order data and resolves a narrow set of routine questions, then passes everything else to a person. It suits operators at $3M–$30M in revenue, on Shopify Plus or a paid subscription platform, who already have a support team and a phone number customers actually ring. It does not suit a brand below the $3M floor, or one where phone is a rounding error in the contact mix.

Most guides on this topic are vendor tours. They describe what the voice sounds like and skip the parts that hurt: what data the agent may read, when it must stop talking, and what a caller hears when something breaks. This guide is the setup order, with the settings to decide at each stage.

Should your store run voice AI customer service at all?

Check your helpdesk before you buy anything. Pull a month of contacts by channel and by reason. If calls are a small share of contacts, a phone agent adds a channel you must monitor without removing much work, and your effort is better spent on chat and email. Our note on helpdesk automation tools covers that first.

Voice earns its place when three things are true. Callers ring for the same few reasons, usually where an order is, whether it can be changed, and how a return works. The answers live in data the agent can read, meaning order status, fulfilment and tracking. And a wrong answer costs a human minute, not a chargeback.

Voice is a poor fit where callers are angry, where the answer needs judgement, or where your data is unreliable. If your tracking numbers arrive late or your order statuses are edited by hand, an agent will read those errors aloud with total confidence. Fix the data first. A bad answer in text is a screenshot. A bad answer by phone is a conversation you cannot see.

Also decide who this is not for. Stores with a handful of calls a week, stores whose returns policy changes case by case, and stores that have not written down what an agent may promise should wait. The agent will do exactly what the policy says, and if the policy is a shrug, so is the agent.

How do you set up voice AI customer service, step by step?

The order matters. Scope comes first because every later setting depends on it, and the handoff rules come before the script because the script is only the happy path. Each step names the settings to decide. Setting labels vary between vendors, so treat the names as the decisions to make, not menu items to hunt for.

Step 1: Choose the call reasons the agent may own

Export a month of phone contacts and group them by reason. Then sort each reason into one of three piles: read-only lookups the agent can answer alone, actions that need a human decision, and calls the agent should never touch.

For an online store, the read-only pile usually holds order status, tracking location, delivery estimate, and return policy questions. The human pile holds address changes after dispatch, damaged goods, and anything involving an exception. The never pile holds refunds, chargeback threats, fraud suspicions and press or legal enquiries.

Write this as a one-page scope document with an explicit “may” and “may not” column. The “may not” column matters more. Every capability you leave open will be exercised by a caller within a week, and a capability you did not test is a capability you will meet first on a live call. Start with two or three call reasons. You can widen scope after a review, but narrowing it after a bad call is a worse conversation to have.

One rule for this step: the agent may read, and may create a ticket, but may not change money. No refunds, no store credit, no discount codes it invents. If a caller wants one, the agent’s job is to collect the request cleanly and route it.

Step 2: Connect order data with read-only access

Give the agent read-only access to orders, fulfilments and tracking. On Shopify this normally means an app or integration with read scopes on orders and customers, not a staff account and never a full-access token. Grant the narrowest scope the vendor’s setup will accept, and if it asks for write access to orders, ask why in writing.

Then test the data path, not the voice. Take twenty real orders across your messy cases: split shipments, backorders, orders with a subscription attached, orders edited after purchase, and orders fulfilled by a third-party warehouse. For each, compare what the agent’s data source returns with what your helpdesk shows to a human. Where they differ, the agent will contradict your team on a live call.

Two data problems show up repeatedly. Tracking status lags because carriers post updates on their own schedule, so the agent says “in transit” while the parcel sits on a doorstep. And subscription orders live partly in a subscription platform, so a caller asking about “my next order” gets an answer from the wrong system. If a question spans two systems, either connect both or take that reason out of scope. Our overview of Shopify helpdesk options covers how the ticket side connects.

Step 3: Set caller verification before any order lookup

Verification is a setting, and a strict default beats a friendly one. Decide three values: which identifiers the caller must supply, how many attempts they get, and what the agent does after the last failure.

Never treat the calling number as proof of identity. Phone numbers are shared by households, forwarded, and spoofed. A safer pattern is two identifiers the caller says aloud, such as the order number plus the billing postcode or the email on the order, both matched against the same order record.

As a starting point for tuning, not a benchmark, allow two attempts and then transfer to a person. Do not let the agent reveal which identifier failed, because that helps someone guessing. Do not read back the full address or the last four digits of a card as a courtesy. The agent should confirm details the caller already stated, and never volunteer new personal data.

What gets read aloud is the second half of this step. Decide exactly which fields the agent may say: order status, carrier, an estimated delivery date and item names might be fine. A full shipping address on an unverified line is not. Write the allowed fields down, because the vendor’s default is usually generous.

Step 4: Write handoff rules by caller state

Handoff rules are the step most teams get wrong, and it is worth slowing down for.

The usual setup writes handoff rules by topic: if the caller says “refund”, transfer. That catches the obvious cases and misses the ones that hurt. A caller asks about tracking, gets an answer, and asks again in different words because the answer did not help. No topic rule fires because the topic is still in scope. The agent answers politely for the third time. The caller hangs up and the next thing you hear is a chargeback or a one-star review.

Write the rules on caller state instead, alongside the topic list. The states to cover:

  • Repetition. The caller asks the same thing twice, or the agent has answered and the caller re-asks. Transfer on the second re-ask.
  • Failed verification. The caller cannot supply the identifiers after the attempts you set in Step 3.
  • Distress or anger. Raised voice, swearing, or phrases like “speak to a person” or “this is the third time I’ve called”. A request for a human is honoured at once, without a retention script.
  • Low recognition. The agent has failed to understand the same input more than once, common with poor lines and accents.
  • Out-of-scope drift. The conversation moves to a call reason from your “may not” column midway through.
  • Silence and confusion. Long pauses, or a caller talking over the agent repeatedly.

Then set what the handoff carries. A transfer with no context forces the caller to start again, which is the failure people remember. The agent should pass a summary that contains the caller’s verified or unverified status, the order in question, what was asked, and what the agent has already said. Have your team write the summary template, and read ten of them before launch.

Finally, decide what happens when no human is available. A transfer into an empty queue with no fallback ends in a dead line or a loop. Options are a call-back promise with a captured number, a voicemail that creates a ticket, or an email hand-off. Pick one and test it at 2 a.m.

Step 5: Tune the script and voice settings

The script is the smallest part of the work and gets the most attention. Keep it plain.

Decide these settings. Response length: short, with one idea per turn, because callers cannot scroll back. Interruption: allow it, since people talk over automated voices, and test that the agent stops speaking rather than finishing its sentence. Read-back: confirm any number the caller gives, whether an order number, a postcode or a phone number, before acting on it. Silence timeout: prompt gently after a pause and offer a person, and do not hang up on a caller who is looking for an email.

On voice choice, pick one that sounds calm and unhurried, and disclose what it is. Do not name the agent with a human first name or script fake typing sounds. Callers who feel tricked are harder to serve than callers who know they are talking to software. Latency deserves a listen too. A pause after every question sounds like a broken line, so place the agent on the vendor’s fastest configuration you can afford and test on a mobile network, not office broadband.

Keep the tone rule simple: no upsell, no cheerful filler, no promises the policy does not make. The agent may say “I can’t do that on this call, I’ll put you through to someone who can.” That sentence prevents more complaints than any clever phrasing.

Step 6: Set recording disclosure and retention

A voice agent creates recordings and transcripts, and that is personal data. Decide three things: what the caller hears at the start of the call, how long audio and transcripts are kept, and who on your team can open them.

The opening line should say the caller is speaking to an automated assistant and that the call may be recorded. Rules on call recording, automated voices and consent differ by state and country and keep changing, and marketing-style outbound calls fall under separate telephone rules. This guide covers inbound support only, and you should confirm the wording and retention period with counsel before launch. Do not let the vendor’s default template be the decision.

For retention, keep transcripts only as long as you need them to review quality and resolve disputes, and delete after that. Restrict access to the support lead and the person tuning the agent. If a caller asks for their data to be removed, you need to know where the audio, the transcript and any summary in your helpdesk live. Map that before the first call, not after the first request.

Step 7: Route the phone number, hours and overflow

Now connect the line. You either forward your existing number to the vendor’s number or port it, and forwarding is the safer first move because you can switch back in minutes. Porting is slower and harder to reverse, so leave it until the setup has earned trust.

Decide the routing values: which hours the agent answers alone, which hours it answers and transfers to staff, and what it does in an overflow. A sensible pattern is staff first during opening hours with the agent as overflow, then agent-first only where you have proven the lookup reasons work. Reverse that order only after review.

Test the failure modes on purpose. Unplug the integration and call. Take the helpdesk offline and call. Call from a withheld number. Call and say nothing. Each of these should end somewhere a human can find it, never in silence. Also record the rollback in writing: who flips the forwarding back, and how a caller reaches a person if the vendor is down.

Step 8: Run a shadow period before going live

Before real callers hear the agent, replay real calls against it. Take a sample of recorded past calls covering each in-scope reason, plus the awkward ones: a caller who rambles, a caller with the wrong order number, and a caller who is furious. Read every transcript where the agent transferred, hung up or contradicted your data.

Then go live on a share of calls, with a person watching the transcripts the same day. Widen the share only when a full review of your busiest day shows no unverified reads and no dead ends. Length is set by coverage of call types, not a calendar number. End the shadow period when every reason in scope has been replayed and reviewed.

Keep one rule through the first weeks: any call that ends in a transfer gets read by a person within a day, because those transcripts show your policy gaps. Once you see the same gap twice, fix the policy, not the prompt.

What does the setup look like in one table?

The setup reduces to eight decisions, each with a value to set and a failure if skipped. Take from the table that the first four rows are risk controls, while the rest are quality controls.

DecisionValue to setWhat breaks if skipped
ScopeTwo or three read-only reasonsAgent improvises on refunds
Data accessRead-only orders and fulfilmentsAgent can alter live orders
VerificationTwo spoken identifiers, two attemptsPersonal data read to the wrong caller
HandoffRules by caller state, with a written summaryLoops and repeat callers
VoiceShort turns, interruption on, read-backWrong numbers acted on
RecordingSpoken disclosure, set retentionUnclear consent and stored data
RoutingForward first, defined overflowDead lines
ShadowReplay every in-scope reasonLive callers become the test

How do you know the agent is working?

Verify with your own baseline, not a vendor’s headline. Pull last month’s contacts before launch so you know your own call volume, transfer patterns and repeat-contact rate. Then track four things after launch: calls resolved without a person, transfers by reason, repeat calls from the same number within a short window, and transcripts that ended in a hang-up.

Vendors publish containment and deflection figures, and helpdesk vendors such as Gorgias publish customer case studies. Read them as vendor-reported. No independent measure exists that maps to your catalogue, your caller mix and your policy, so any figure you plan around is a metric to confirm from your own data, not something to lift from a sales page.

Repeat calls are the number to watch. A call the agent “resolved” that comes back tomorrow was not resolved. Count it as a failure, and read the transcript of the first call to find out why.

Also sample calls yourself. Listen to a handful each week, including ones marked successful. Automated scoring misses tone, and a caller who was answered correctly but left annoyed still costs you.

Where does voice AI not belong?

Keep the agent out of refunds without a human, out of anything a chargeback could follow, and out of any conversation where a wrong answer costs more than a human minute. Keep it out of accounts where the underlying data is unreliable. If you want an agent that acts inside your systems, that is a bigger design question, and our piece on conversational AI for ecommerce sets out the difference between answering and acting. For a wider view of the tooling, the list of ecommerce AI bot platforms compares the packaging shapes.

Callers who are vulnerable or upset are the other exclusion. A bereaved customer chasing an order or a caller reporting a safety problem needs a person who can bend the policy. Your handoff rules should treat those as an instant transfer.

What about running this at volume?

Volume exposes what small tests hide. Peak days bring concurrent calls, and the vendor’s line limits and your helpdesk’s queue both cap what happens next. Ask what happens when every line is busy, and whether callers hear a busy tone, a queue or a call-back offer. Promotions and delivery disruptions change your call mix overnight, so keep a switch that narrows the agent’s scope, for example pausing delivery-date answers when a carrier is failing. Cost is worth a look too: voice vendors price in shapes such as per-minute, per-call, per-resolution or bundled seats, and you should check the current pricing on their own pages and model it against your call volume before signing.

Ownership matters as much as the tooling. Someone must own the agent weekly, reading transcripts and changing the scope document, or quality drifts as the policy and catalogue change. If no one on your team has that time, that is a cost of the project, not a footnote.

Who should own the setup?

Voice AI customer service is a customer service AI problem, and it lives or dies on scope, data access and handoff design rather than on the voice itself. If you would rather have that designed and monitored than built in-house, that is the work we describe on customer service automation. For agents that take actions beyond answering, see AI agents.

Sources

  • No external figures are quoted. The article is written from the documented behaviour of phone agents, Shopify order data and helpdesk hand-offs; vendor case studies are referred to only as vendor-reported and none is relied on for a number.

Frequently asked

Can a voice agent process refunds on a phone call?

It should not. A refund moves money, and a misheard order number or a spoofed caller turns a small mistake into a payout. Let the agent collect the request, verify the caller and log a ticket, then have a person approve and issue the refund.

Does voice AI customer service work for a store with no phone line today?

It can, but the case is weaker. If customers already email and chat and few call, a phone agent adds a channel to run rather than removing load. Check your helpdesk for call volume first. If phone is a small share of contacts, fix chat and email first.

Which helpdesk should the voice agent sit on top of?

Use the helpdesk your team already works in, so transfers land in a queue humans watch. Zendesk, Gorgias and similar tools each publish their own voice or telephony options and partner integrations. Check the current documentation for the one you run instead of assuming parity.

How do I stop the agent reading order details to the wrong person?

Never trust the calling number alone, because numbers are shared and spoofed. Ask for two identifiers the caller must say, such as the order number plus the billing postcode, and match both against the order. After repeated mismatches, stop and transfer to a person.

What happens when a caller has an accent or a bad line?

Speech recognition degrades on both, and the agent may mishear digits. Read every number back for confirmation, and count repeated re-asks as a handoff trigger. Test with recordings of your real callers, including mobile calls made from a street, before you trust vendor demos.

Do I need to tell callers they are talking to an AI?

Plan on yes. Disclosure rules on AI and recorded calls vary by state and country and are changing, so state plainly at the start that the caller is speaking to an automated assistant and that the call may be recorded. Confirm exact wording with counsel.

Should the voice agent upsell or offer discounts?

Not at first. A caller with a late parcel is not in a buying frame of mind, and a discount the agent invents costs margin you did not approve. Keep the launch scope to lookups and routing. Revisit offers later, with fixed rules a human wrote.

How long should the shadow period last?

Long enough to see your weekly rhythm, including your busiest day and a post-sale spike if you run promotions. The right length is set by coverage of call types, not by a calendar number. End it when every call reason in scope has been replayed and reviewed.

What should a human receive when the agent transfers?

A short written summary: who called, whether they were verified, what they asked for, what the agent already told them and the order number in question. Without it the caller repeats everything, which is the single most common complaint about automated phone systems.

Is a voice agent worth it for a brand under $3M revenue?

Probably not yet. Below the $3M floor most stores have too few calls for a phone agent to pay back the setup and monitoring effort, and a shared inbox with good macros does more. Come back when phone volume forces a person to answer the line.

What metrics should I watch after launch?

Track containment (calls resolved without a person), transfer rate by reason, repeat calls from the same number within a short window, and the transcripts that ended in a hang-up. Set your own baseline from your helpdesk first. Published vendor figures are vendor-reported and rarely match your mix.

Next step

Is this your customer service ai problem, or a symptom of another one?

Bring your numbers — the churn split, the decline rate, whatever your flows are earning — and we will tell you which of them is the expensive one.

Book a call →