Outsourcing customer service goes wrong in a predictable way. A brand signs with a vendor, forwards the support inbox, sends over a folder of macros and waits. Three weeks later refunds are up, replies sound like someone else’s company, and the founder is answering escalations at 11pm. The vendor is rarely the cause. The handover was missing its most important document.
This page is the setup we would follow to outsource ecommerce customer service on Shopify: the order of the steps, the setting values to write down, and the one step most teams skip. What it adds beyond a vendor checklist is a rule for writing refund authority as a dollar limit per contact reason before the vendor sees a ticket, because that is where the money and the brand voice are lost.
Should you outsource ecommerce customer service at all?
Outsourcing customer service makes sense when your inbox is dominated by repeatable questions and your team is spending its judgement on them. It makes less sense when most of your contacts are complex, high-value or emotionally charged. The test is your own ticket data, not a vendor’s sales deck.
A useful frame is the cost of a wrong answer against the cost of a human minute. A “where is my order?” reply that is slightly off costs you a follow-up. A wrongly denied warranty claim on a high-ticket item can cost you a customer, a chargeback and a public review. The first belongs with an outsourced team from day one. The second stays with you until you have written down exactly how it is decided.
There is also a timing point that vendors will not raise. Outsourcing a broken process does not fix it, it scales it. If your macros contradict each other and your returns policy is unclear, an outsourced team will answer inconsistently at a higher volume. Fix the policy first. The vendor is then executing your rules, not inventing them.
Nothing in this article assumes you are a small shop. It is written for operators on Shopify Plus or a paid subscription platform at roughly $3M in revenue and above, where ticket volume is high enough that a vendor contract is worth negotiating and the cost of a bad handover is measured in real repeat-purchase revenue.
Who is this setup not for?
Brands below the published $3M floor are not the audience. At that size a founder answering support personally is often better for retention than any vendor, and the setup below would be more process than the volume justifies.
It is also not for teams that want to remove humans entirely. This page assumes people answer customers, with tooling behind them. If you want the routine layer handled by software, that is a different build, covered near the end.
Finally, it is not for brands whose support is the product: concierge services, made-to-order goods with heavy back-and-forth, or anything where the conversation is the sale. Handing those conversations to a third party changes what you are selling.
How do you outsource customer service on Shopify, step by step?
The six steps below run in order. Each produces something you can hand to the vendor as a document or a setting, which is the point: by the time a vendor agent opens their first ticket, every decision they might need already has an owner and a value.
Step 1: Classify your tickets by contact reason
Export a representative sample of your own tickets from your helpdesk. Include a peak week, because that is when a vendor will be tested. If you use Shopify’s own helpdesk options or a third-party one, the export is a standard function.
Tag every ticket with exactly one contact reason. A workable starting list has these tags: order-status, address-change, cancel-before-ship, damaged-item, wrong-item, return-request, refund-status, discount-code-problem, product-question, subscription-change, payment-issue and other. Keep the list short. If a tag needs a paragraph to explain, split or merge it.
Count the tickets per tag. You are looking for two groups. The first is high volume with a single correct answer that lives in your order data: order status, address changes, refund status. The second is low volume with real consequences: damaged items, payment issues, anything that mentions a lawyer, a regulator or a chargeback. The first group goes to the vendor. The second stays in-house or goes through your escalation queue.
Do not skip the counting. Teams that pick contact reasons from memory usually overestimate the interesting tickets and underestimate the dull, repetitive ones, and the dull ones are what you are outsourcing.
Step 2: Write authority limits before the vendor sees a ticket
This is the step most teams get wrong. They hand over macros and a tone guide but never say how much an agent may give away. The vendor then does one of two things: escalates everything, which defeats the purpose, or improvises, which costs you money.
Write a table with one row per contact reason and four columns: the most an agent may refund without approval, the most they may replace or reship without approval, the most they may offer as a goodwill discount, and the condition that forces escalation regardless of amount. The dollar figures are yours to choose, and they should come from your own average order value and margin, not from a template. To work them out, take your typical order value for each product line, decide what loss per incident you would accept without a conversation, and set the limit at or below that.
Then add the non-money triggers. Escalate any message that mentions injury, an allergic reaction, a regulator, a chargeback or a legal threat. Escalate any customer the helpdesk has already marked as a repeat claimant. Escalate any request from someone who says the order was placed under another person’s name. These rules cost nothing to write and prevent the incidents that hurt most. For anything touching disputes, see how chargebacks work for ecommerce brands before you finalise the list.
Put the table into the playbook and, where your helpdesk supports it, into the tool itself. Approval queues and refund permissions enforced by the system are far more reliable than a document an agent has to remember at 3am.
Step 3: Scope helpdesk and Shopify access
Create a named seat for every vendor agent in your helpdesk. Shared logins hide who did what, and you will want that audit trail the first time a refund looks wrong.
In Shopify, create a dedicated staff account for each vendor agent with a restricted permission set. The values to set: allow viewing and editing orders, viewing customers, and issuing refunds up to what your plan and integration allow; do not allow access to settings, apps, payouts, staff management, or the theme. Shopify’s permission names and staff limits change, so check the current list in the Shopify admin’s staff permissions screen and in Shopify’s help documentation before you finalise it.
Turn on two-factor authentication for every seat, and agree in writing that a departing agent’s access is removed on the same day. If the vendor connects to your store through an app or an integration, read what data that app can reach and remove any scope you would not grant to a person.
Keep the tickets, tags and macros in an account you own. If the vendor’s contract requires you to use their platform, price the migration cost of leaving into your decision, because your customer history is part of what you would be leaving behind.
Step 4: Build the playbook, macros and escalation queue
Write one macro per contact reason, using your real policy wording. A macro is a starting point, not a script: agents should be told to change it to fit the customer’s message, and a reply that is only the macro should be treated as a quality failure.
Build a tone guide from real conversations, not adjectives. Choose three replies you are proud of and three that went badly, and write one line under each saying why. Add a short list of phrases your brand never uses. Vendors follow examples far better than they follow the word “friendly”.
Set up a single escalation queue. In the helpdesk it should be a view or a tag such as escalate-brand, assigned to a named person on your side. Agree a response window in writing, and set it against the working hours of that person, not an ideal. An escalation that waits until Monday because nobody owned the queue on Friday is the commonest way an outsourced setup embarrasses its owner.
Finish with a short list of things the vendor must never do: promise a delivery date you do not control, confirm stock you have not checked, or discuss a competitor. Put it on one page.
Step 5: Run a supervised pilot
Start with one or two contact reasons from the high-volume group, not the whole inbox. Order status and address changes are the usual choice because the correct answer sits in Shopify data and can be checked in seconds.
During the pilot, someone on your team reads every reply before or shortly after it goes out. Keep a running list of corrections. If the same correction appears twice, it goes into the playbook; if it appears three times, it goes into the tone guide or the macro.
Set the exit rule before the pilot starts. A rule like “widen scope when reviewed replies stop needing edits across a normal week and an awkward one” works. A calendar date alone lets a weak pilot pass because time ran out. Include one awkward week on purpose: a delayed shipment, a sale, a stock-out. A vendor that is only tested in a quiet week has not been tested.
Widen one contact reason at a time. Adding refund handling and damaged-item claims in the same week makes it impossible to tell which one caused a problem.
Step 6: Audit a weekly sample and re-tune
Once live, read a random sample of closed conversations every week. Random matters: if the vendor picks the sample, you are reading their best work. Score each conversation against a written scorecard with four questions: was the answer correct, was it within authority, did it sound like your brand, and did it need a second contact to resolve?
Reconcile refund and reship totals against your authority table. Where the totals exceed what the limits should allow, find out whether an agent exceeded the limit or the limit was ambiguous. It is usually the second.
Track reopen rate and repeat purchase for contacted customers next to response time. Speed alone is the metric vendors find easiest to hit and the one that most often masks thin answers. If you do not have these numbers, mark them metric to confirm and build the report from your helpdesk and Shopify customer data before you sign a long contract.
Feed everything back. The weekly review is where the playbook improves, not the kick-off call.
What breaks once volume grows?
Volume exposes three problems that a calm pilot hides.
Macro drift. After a few months of edits by many agents, the macros in use no longer match the policy. The fix is a named owner on your side, a monthly diff of macros against your live returns and shipping policy, and one rule: policy changes are published to the vendor the same day they go live on the store.
Promotion and shipping-delay spikes. A sale or a carrier problem produces a wave of one contact reason. Agree in advance how the vendor scales for it, whether by overtime, a surge pool or a change to the macro that answers the question once for everyone. Ask what the contract says about response times during a spike, since that is the week the contract matters.
Escalation queue backlog. If escalations rise, the vendor may be escalating out of caution because the limits are unclear, or your own queue owner may be overloaded. Look at which contact reasons dominate the queue. A pile of the same reason is a missing rule, not a difficult customer.
There is also a slower failure: knowledge loss. Agents leave, and a vendor retrains their replacements from the playbook. If the playbook is thin, quality falls each time someone new joins. The remedy is the same as everything above: keep the playbook as a living document that you own.
Returns are the area where handover most often fails, because a return touches the warehouse, the payment and the customer at once. If returns are a large share of your tickets, read how returns management works before you decide how much of it an outsourced team may touch.
How do you verify the vendor is doing the work well?
Verification is a routine, not an event. Each week you read a sample, reconcile money and read the reopen numbers. Each month you compare macros to policy and review the escalation queue by reason. Each quarter you repeat Step 1 on fresh data, because your contact mix moves as your catalogue and shipping partners change.
Ask the vendor for the same numbers you compute yourself, and treat any gap as a question. Vendor case studies, including those published by helpdesk platforms such as Gorgias, describe results for their own customers and are vendor-reported. No independent measure of outsourced ecommerce support quality exists that you can compare against, so your own baseline, taken before handover, is the only fair benchmark. Record it in Step 1.
Do not judge the vendor on a single month. A sample of a few conversations is anecdote. A steady weekly read over a full sales cycle is evidence.
Agree at the start what happens when the numbers are bad: a written improvement period, a right to pause new contact reasons, and a clear exit with your data returned. A contract without an exit clause turns every quality problem into a negotiation.
Where does AI belong in an outsourced setup?
Software and people each have a job here, and an outsourced team does not remove the case for automation. The routine layer, such as order status, tracking lookups and address-change requests, is exactly where software can answer from Shopify data without a person. The judgement layer stays human.
AI does not belong in refunds without a human, in any conversation where a wrong answer costs more than a human minute, or anywhere the underlying data is unreliable. If your order data is wrong, an assistant will state the wrong thing confidently and at speed. Fix the data before you automate anything on top of it, and design the handover between software and people before you buy either. For more on the tooling side, see helpdesk automation tools.
The way to think about the whole job is as a customer service AI and operations problem: which contacts software should resolve, which a trained team should handle within written limits, and which you keep. That is the work our customer service automation service is built around, and the six steps above are the input to it.
Sources
- No external figures are quoted. This article is written from long-standing helpdesk and Shopify staff-permission practice, and vendor case studies (such as those Gorgias publishes) are noted as vendor-reported. Confirm current plan limits and permission names in Shopify’s own documentation.