When Shopify shows “high risk of fraud detected” on an order, it’s tempting to treat the label as the decision. It isn’t. It’s a flag built from a handful of underlying signals, and the operator’s job — the part no fraud-analysis panel does for you — is reading those signals and choosing between three real options: cancel, hold for review, or ship anyway. Get the call wrong in one direction and you eat a chargeback. Get it wrong in the other direction and you cancel a customer who was never a fraud risk at all, on an order you’d already have shipped if the flag hadn’t scared you off it.
What Does “Shopify High Risk of Fraud Detected” Actually Mean?
The flag is a summary label sitting on top of several individual signals, and the label alone tells you almost nothing about which signals fired or how strongly. Shopify’s fraud analysis typically draws on factors including a mismatch between the billing and shipping address, unusually high order velocity from the same device, IP address or card, use of a proxy or VPN at checkout, and a shipping address that has a history of disputed orders on the platform. The specific weighting and the exact list of contributing signals aren’t fully published, and they change, so treat the mechanism as directional rather than a fixed formula you can reverse-engineer once and rely on forever.
What matters operationally is that “high risk” is not “confirmed fraudulent.” A first-time customer shipping a gift to a different address, paying with a card issued in a different country to where they’re currently browsing from, on a work VPN, can trigger several of the same signals as an actual fraud attempt — the flag doesn’t distinguish intent, only pattern. That’s precisely why the label can’t be the whole decision. The individual signal panel, which shows what specifically contributed to the score, is the part worth reading every time, not the summary badge.
It’s also worth being clear about what this flag is not. It is not a chargeback determination, it is not a payment hold placed by your bank, and it is not the same mechanism as Shopify Protect, which reimburses certain disputes after an order has already shipped and been charged back. That’s a separate, later-stage problem with its own evidence requirements. This article stays on the decision you make before the order ships.
Before You Decide: What You Need in Front of You
Before you act on any high-risk flag, pull together four things, because a decision made without them is a guess dressed up as a review.
The individual signal breakdown, not just the risk label. Every major fraud-analysis tool on Shopify — the built-in fraud analysis, or a third-party app if you’re running one — exposes which specific factors contributed to the score. Open that panel before doing anything else.
The customer’s order history, if they have one. A repeat customer with three clean prior orders who trips a flag on their fourth carries a different risk profile than a first-time buyer with the same trigger and nothing to check it against. Shopify’s admin shows prior orders tied to the same customer profile, email, or payment method; check it.
Your own baseline for this product category and price point. As an illustrative comparison, a $40 order and a $2,000 order shouldn’t be reviewed against the same instinct. High-value, easily resold categories — electronics, designer apparel, gift cards — warrant more scrutiny on a weaker signal than a low-value, made-to-order item does.
A defined escalation point, meaning someone specific who makes the manual call on holds, and a rough time budget for how long a hold can sit before it either ships or gets cancelled. Without this, holds accumulate until a shipping deadline forces a rushed decision instead of a reviewed one.
Step 1: Read the Fraud Analysis Signal Before You Do Anything Else
Open the order and go straight to the fraud analysis panel rather than the order list’s summary badge. The summary badge shows one of a small number of risk levels; the panel underneath shows what specifically drove it. This distinction is the single most important habit in the whole process, because two orders both labelled high risk can be flagged for entirely different reasons, needing entirely different responses.
An order flagged for a billing-shipping mismatch alone, with no velocity signal and no proxy use, is a materially weaker case for cancellation than an order flagged for velocity, proxy use and a disputed-address history stacked together. Reading only the badge collapses that distinction and pushes you toward treating every high-risk order the same way, which is exactly how good orders end up cancelled alongside real fraud.
If you’re running a third-party fraud app alongside Shopify’s own analysis, check both panels rather than assuming they agree — they draw on different signal sets and can disagree on the same order. A conflict between the two is itself useful information: it usually means the order sits in genuinely ambiguous territory and belongs in the hold category rather than an automatic cancel or ship.
Step 2: Check the Order Against Your Own Risk Baseline
The flag is generic; your business isn’t. A billing-shipping mismatch that would be alarming for a first-time customer buying a single high-value item is far less alarming for a returning customer who has simply moved house or is shipping a gift, and your own order history tells you which situation you’re looking at.
Pull the customer’s prior order count, their account age if they have one, and whether the current shipping address has appeared on any previous order from this account. Cross-reference the product category against your own chargeback history if you track it: a store that’s seen repeat fraud attempts on a specific high-margin SKU has a stronger case for treating a flag on that SKU seriously than the same flag on a low-value staple item.
This step is also where you weigh order value against review cost. As an illustrative comparison, a $35 order doesn’t justify the same manual-review time as a $1,200 order, even at an identical risk score, simply because the downside of getting it wrong is proportionally smaller. Set an internal value threshold below which flagged orders get a lighter, faster check rather than a full manual review, so your review capacity goes where the exposure actually is.
Step 3: Decide Between Cancel, Hold and Ship
Cancel, hold or ship is the decision the whole process exists to support, and it has three outcomes, not two. Treating “high risk” as binary — ship or cancel — throws away the option that catches the most ambiguous cases correctly.
Cancel outright when multiple hard signals stack: velocity plus proxy use plus a shipping address with a prior dispute history, for instance. Stacked hard signals are the closest this process gets to a confident call, and even then, cancelling politely with a clear customer-facing reason (a payment verification issue, not an accusation) preserves the relationship if it turns out to be a false positive.
Hold for manual review when signals conflict, or when a single strong signal has no corroborating pattern and no obvious innocent explanation. This is where a human check — a quick address or phone verification, a look at whether the same customer has ordered successfully before — resolves more ambiguity per minute spent than any automated rule will.
Ship when the flag traces to a single weak signal with a plausible, checkable explanation: a returning customer’s new shipping address, a VPN that’s explained by their stated location, an order value consistent with their history. Shipping a weakly flagged order from a known-good customer is very often the correct call, and it’s the option teams under-use because the flag itself feels like a reason to act.
The order value and the cost of a wrong decision in each direction should shape where you draw these lines, not a fixed rule copied from another store. A low-margin, easily-returned product can tolerate a more permissive ship threshold than a high-margin, hard-to-recover one.
Step 4: Document the Decision and the Evidence Behind It
Whichever way the order goes, record what drove the decision. For a cancellation, note which signals were stacked and who made the call. For a hold that resolved to ship, note what the manual check found — a confirmed phone number, a verified address, a customer service reply. For a ship decision made without a manual check, note that too, along with the specific weak signal that was judged insufficient to hold.
This record does two things. First, it’s your evidence if the order later disputes as unauthorised: a documented, signal-based decision to ship is a far stronger position in a representment case than no record at all. Second, and just as useful, it’s how you spot a pattern in your own decision-making over time — if every held order that got a manual phone check turned out fine, that’s a signal your phone-verification step is doing real work and worth keeping as a standard part of the hold process, not an occasional extra.
Keep this record somewhere searchable by order, not buried in a private note only one person can find. A support or fulfilment teammate picking up a customer enquiry about a delayed or cancelled order needs to see the reasoning without pulling in whoever made the original call.
Step 5: Verify the Outcome After Fulfilment
The loop closes here, and it’s the step most stores skip entirely. Set a recurring check — weekly is reasonable for most order volumes — where you review what happened to orders that went through this process: how many held orders were confirmed good and shipped, how many cancelled orders generated a customer complaint or a repeat order attempt that also looked suspicious, and how many shipped orders later came back as a dispute.
Reviewing outcomes on a schedule is the only way to know whether your cancel threshold is set correctly. A store that cancels ten high-risk orders a week and sees zero related disputes and zero customer complaints about wrongful cancellation might be running a threshold that’s too permissive, or might genuinely be catching real fraud — you can’t tell the difference without this review. A store that cancels ten and gets three angry emails from confirmed real customers has clear evidence the threshold or the process is too aggressive.
Where you can, follow up directly with customers whose orders were cancelled, briefly and without accusation, to ask whether the order was genuinely theirs. Not every store has the support capacity for this, but even a sample check on a portion of cancellations tells you more about your real false-positive rate than any dashboard number will, because the dashboard has no way to know a cancelled order was actually legitimate.
The Step Most Teams Get Wrong
The single most common mistake is auto-cancelling on the summary label without opening the signal panel underneath it. It’s understandable — a queue of flagged orders at the end of a busy day, a rule that says “cancel anything marked high risk,” and the whole review collapses into one click per order. It’s fast, and it feels safe, because cancelling a flagged order never generates a chargeback.
What it actually does is convert an unknown number of good orders into cancelled ones, silently, with no visibility into how many of them were real customers. Unlike a chargeback, a wrongful cancellation rarely generates a complaint loud enough to be noticed — most customers who get quietly cancelled simply don’t come back, and you never see the lost order as a lost order. It shows up, if it shows up at all, as a slightly lower repeat-purchase rate you can’t trace to a cause.
The fix isn’t complicated, but it does cost time: read the signal panel, check the order against customer history, and reserve the automatic cancel rule for the narrow case of multiple stacked hard signals with no mitigating history. Everything else earns at least a look before a decision.
The Cost of Over-Cancelling Good Orders
There’s no universal figure for what over-cancellation costs a given store — it depends on average order value, repeat-purchase rate, and how aggressive the threshold is — so build the number for your own business rather than borrowing one. A workable, illustrative method: take your current weekly count of cancelled high-risk orders, estimate what share you believe were likely legitimate based on a sample follow-up (Step 5 above gives you this), multiply by average order value, and then multiply again by your typical repeat-purchase rate, since a wrongly cancelled first-time customer is a lost future customer, not just a lost single sale.
As a hypothetical, illustrative only: a store cancelling twenty high-risk orders a week, estimating a quarter of them were likely legitimate based on a follow-up sample, at an illustrative average order value of $85, is looking at a roughly illustrative $425 a week in lost revenue from wrongful cancellations alone, before accounting for any repeat purchases those customers would otherwise have made. That figure is entirely illustrative and depends on numbers you’d need to measure in your own store — the exercise is the method, not the number.
What this method makes visible is that over-cancellation has a cost structure similar to a chargeback: money you don’t see leaving, attributed to nothing in particular, until you go looking for it specifically.
Where This Breaks at Volume
A single ops person reviewing five flagged orders a day can genuinely open the signal panel and cross-check history on every one. That doesn’t hold once volume climbs, particularly around a seasonal peak, when order volume rises and fraud attempts rise alongside it, and the flagged-order queue grows faster than review capacity does.
Two things happen under that pressure. First, review quality drops: signal panels stop getting opened, and the process quietly reverts to cancelling on the summary label, the same mistake covered in the section on what teams get wrong, now happening at scale instead of occasionally. Second, holds start expiring under shipping-deadline pressure rather than resolving on review — an order held for manual check on day one, still unreviewed by day three, either ships un-reviewed to hit a shipping promise or gets cancelled by default because nobody got to it. Neither outcome is a decision; both are a process failing under load.
The fix at volume is triage, not more manual reviewers indefinitely: use the stacked-signal and single-weak-signal ends of the decision (Step 3) to auto-resolve the clearest cases — a very small number of hard-stacked signals to auto-cancel, and a larger number of genuinely low-risk single-signal flags to auto-ship — and route only the genuinely ambiguous middle to a human queue. That middle category is usually a much smaller share of total flagged volume than it first appears, once the clear ends are peeled off.
Who Should Never Auto-Cancel on a High-Risk Flag Alone
This process assumes you’re an operator with enough order volume and margin to justify a manual review step at all — realistically, brands doing $3M-plus in revenue on Shopify Plus or a comparable paid subscription platform, where the cost of a wrongful cancellation and the cost of a chargeback are both large enough to warrant deliberate handling. If you’re running a small store with a handful of orders a day, a lighter-touch process — a quick personal look at every flagged order, rather than a formal signal-panel-and-baseline review — may genuinely be proportionate, and this level of process would be overbuilt for that volume.
At the other end, a business selling low-value, low-margin, hard-to-resell goods (an illustrative $12 consumable, say) has a weaker case for heavy manual review on every flag, because the review cost can exceed the exposure. That’s a legitimate reason to lean more heavily on automatic rules at that price point, provided you’re still tracking outcomes per Step 5 rather than setting the rule once and forgetting it.
What to Do When Fraud Analysis Disagrees With Your Own Read
Occasionally the fraud analysis panel flags an order your own judgement says is fine, or clears an order that feels wrong for reasons the tool doesn’t capture — a product combination that doesn’t match your typical customer, an unusually specific and rehearsed-sounding support message before the order was even placed. Treat this disagreement as information, not noise.
When the tool flags and your read says ship: this is exactly the hold category. Use the manual check — a verification contact, an order-history look — rather than overriding the flag on instinct alone, and document why you overrode it if you do.
When the tool clears an order and your own read says something’s off: fraud-analysis tools don’t see everything, particularly signals from outside the order itself, like a support conversation or a pattern across a small number of orders that individually look fine. Trust a specific, articulable reason over a vague feeling, but don’t discount it just because the automated score came back clean — the score is an input to your decision, not a replacement for it.
Fraud analysis exists to speed up the common cases, not to remove judgement from the uncommon ones, and the uncommon ones are disproportionately where the real cost sits, in both directions.
The cancel-hold-ship call, made correctly order by order, is the fraud-and-chargebacks problem in its rawest form: a wrong call in either direction costs real money, and neither direction shows up as clearly on a dashboard as a chargeback does. Pointerflow’s fraud and chargebacks work covers this decision layer alongside the representment work that follows when a shipped order is later disputed.
Sources
- Stripe: 25% of lapsed subscriptions trace to a failed payment rather than fraud or dissatisfaction (vendor-reported).