What does shopify attribution actually measure?
Shopify attribution is the decision, made after the fact, about which marketing touch gets credit for a completed order. Shopify’s own order record is not in dispute: a customer paid, the order exists, the ledger has it. What’s in dispute is which channel caused it, and every tool, platform and model answers that question with a different set of rules.
That’s the part worth saying plainly before anything else: there is no neutral, objectively correct attribution number waiting to be found. There’s a ledger, which is factual, and a set of models built on top of it, each making assumptions about which touches count, over what window, with how much credit split between them. The work isn’t finding the “real” number. It’s choosing a model, documenting why, and being able to explain the gap between what it says and what the ledger says.
Why do platform-reported conversions exceed your order ledger?
Every major ad platform counts a conversion using its own attribution window and its own rules for what counts as a qualifying touch, entirely independently of what any other platform reports for the same customer journey. A single order that involved a paid social ad, a paid search click and an email open can get claimed, in full, by all three, because none of the three platforms knows or checks what the others are claiming.
That overlap is the largest structural reason platform totals sum to more than your order count, and it isn’t a bug in any one platform’s tracking. View-through conversions widen the gap further: a customer who saw an ad, didn’t click it, and purchased later through an unrelated path can still get credited to that ad under a platform’s view-through window, a form of credit your order ledger has no equivalent concept for at all. Refunded and cancelled orders add a third source of drift, because most platforms record a conversion at the moment of purchase and don’t automatically reverse it when Shopify later processes a refund, so a channel can keep “credit” for revenue that no longer exists on your books.
Bot traffic, accidental clicks and a platform’s own generous definition of what counts as an engaged session add a smaller, harder-to-quantify layer on top. None of this means the platforms are lying. It means each one is answering a narrower question than “which channel actually drove this specific order,” and treating their combined totals as if they summed to your revenue was never going to hold up.
Cross-device journeys make the gap harder to close after the fact, not just wider to begin with. A customer who clicks a paid social ad on a phone, closes the app, and completes the order later on a laptop through a saved cart severs the click-to-purchase chain a single-device attribution window assumes, forcing the platform to rely on probabilistic identity matching that it does not disclose the accuracy of. That matching happens inside the platform, invisibly, and there’s no way for you to audit it order by order; you can only see its effect in aggregate, as part of the same gap already being discussed here.
Why does trusting the ad platform’s own dashboard fail as a strategy?
The obvious response, reading a platform’s dashboard and trusting the number it reports, fails because every platform’s dashboard is structurally built to report favourably on itself. That’s not a conspiracy; it’s a direct consequence of each platform only seeing its own touches and having every incentive to attribute broadly within its own window rather than narrowly.
Compare two platforms’ dashboards side by side for the same period and you’re not comparing two measurements of the same thing. You’re comparing two different models, run over two different windows, against two different definitions of a qualifying touch, both of which claim more revenue between them than actually landed in the bank. Add each channel’s budget owner reading their own platform’s number as the ground truth, and you get a business where every channel looks like it’s working, the sum of “working” channels exceeds total revenue, and nobody can say with confidence which spend to cut.
Whether conversion events even reach each platform reliably in the first place is a separate, plumbing-level failure, covered in a companion piece on shopify server side tracking, with its own separate fixes. Assume event delivery is solid for the rest of what follows here; the crediting-overstatement problem covered in this section still exists on top of it, because it’s a modelling and crediting issue, not a data-loss one.
What does a defensible attribution model actually require?
A model you can defend starts from the order ledger, not from any platform’s dashboard, because the ledger is the one number in the whole exercise that isn’t a modelling choice. Everything else, the window, the touch definitions, the credit-splitting rule, is a decision your business made, and a decision has to be written down to be defensible later.
Three concrete things follow from that. First, a documented methodology: which touchpoints count, over what window, how credit splits when more than one touch is present, and why that window and that split were chosen rather than a platform default. Second, a fixed reconciliation cadence, comparing modelled or platform-reported figures against the actual ledger on a schedule, not only when a number looks wrong. Third, version control on the methodology itself: when the window or the model changes, the change and the date it took effect get recorded, so a quarter-over-quarter comparison isn’t silently comparing two different methodologies as if they were the same one.
| Model | What it measures | Strength | Weakness | When it fits |
|---|---|---|---|---|
| Last-click (platform default) | Credit to the final tracked touch before purchase | Simple, matches most platform dashboards by default | Ignores every touch that built awareness earlier in the journey | Early-stage measurement, single dominant channel |
| Linear / multi-touch | Credit split across several touches in a journey | Recognises that most purchases involve more than one touch | The splitting rule is itself a modelling choice, not a measured fact | Multiple active channels, longer consideration periods |
| Media mix modelling / incrementality | What a channel actually added, tested against a holdout or baseline | Answers causation, not just correlation of touches | Needs enough spend and volume to produce a reliable read | Larger budgets, mature channel mix, real test-and-control capacity |
| Ledger reconciliation | The actual revenue that landed, net of refunds, against modelled or platform claims | The one figure that isn’t a modelling assumption | Says nothing on its own about which channel caused what | The baseline every other model in this table should be checked against |
Read this table as a stack, not a menu: ledger reconciliation is the floor every other model sits on, not an alternative to pick instead of them.
What should the written methodology document actually contain?
A methodology document earns its keep the first time someone questions a number, which means it has to answer questions a spreadsheet of totals can’t. Name the attribution window explicitly, in days, for each channel, along with the reason that window was chosen rather than left at a platform default. Name what counts as a qualifying touch, an ad click, a view-through impression, an email open, and state plainly which of these your model includes and which it deliberately excludes.
State the credit-splitting rule in words a non-technical stakeholder can follow, not just as a formula: does the first touch get more weight than the last, are all touches weighted equally, is there a decay applied to older touches. Record how refunds and cancellations are handled, and on what schedule they’re subtracted from reported figures. Record the reconciliation cadence itself, so anyone reading the document knows how current the numbers in front of them actually are.
Date every version of the document and keep prior versions rather than overwriting them, because the most common way a methodology fails under questioning isn’t that it’s wrong, it’s that nobody can say when or why it changed. A model that quietly moved from a 7-day to a 14-day click window between two quarters, with no written record of the change, produces a quarter-over-quarter comparison that looks like a performance shift when it’s actually a methodology shift.
How often should you reconcile platform numbers against the ledger?
A weekly directional check, comparing each platform’s reported conversions against Shopify’s order count for the same period, catches a channel drifting further out of line than usual early enough that it’s still cheap to investigate. This doesn’t need to be precise to the dollar; it needs to flag a gap that’s grown compared to its own recent pattern.
A full monthly reconciliation, refunds subtracted from the ledger side and each platform’s window-driven overlap accounted for, is what any number used for a real spend decision or reported upward should be built on. Waiting until quarter-end to reconcile means finding a three-month-old problem after three months of decisions were already made on the wrong number, which is a more expensive place to discover it than a Tuesday in week two.
Campaign-level checks deserve a lower threshold than account-level ones, because a launch or a promotional push is exactly when webhook volume, test traffic and short-window view-through claims are all higher than usual at once, and a gap that would be background noise on an ordinary week can represent a meaningfully wrong budget call during a launch.
How does attribution differ for brand versus non-brand paid search?
A branded search term, a shopper typing your store’s name directly into Google, gets credited by most platforms the same way a cold non-brand click does, even though the two represent almost opposite levels of marketing effort. A customer who already knew your name and searched for it was very likely going to land on your site through some path regardless of whether a paid ad was sitting above the organic result; a non-brand click on a generic category term is doing genuine discovery work a branded click isn’t.
Blending the two into a single paid search return figure overstates the channel’s real contribution, because branded clicks tend to convert at a materially higher rate and cost less per click, pulling the blended number up and making the whole channel look more efficient than the non-brand spend within it actually is. Splitting brand and non-brand into separate campaigns, with separate reporting lines in your reconciliation, is the only way to see whether non-brand spend is earning its place on its own terms rather than riding on a branded conversion rate it didn’t create.
This split matters more, not less, as a store’s brand recognition grows: a $3M store with limited name recognition sees a small brand-term share, so the blended number is a reasonably honest read of paid search performance, while a store nearer the top of the $3M-$30M range, with a well-known name, can see branded clicks carry a large enough share of paid search’s reported conversions that the channel’s “efficiency” is mostly measuring how well-known the brand already was, not how well the campaigns are running.
How do subscriptions and repeat orders change the attribution picture?
A subscription renewal, processed through an app like Recharge rather than through a fresh checkout, still generates an order in Shopify, and a naive attribution setup will credit that renewal to whatever channel or touch the model last saw, often a channel the customer never interacted with again after their first purchase months earlier. Left uncorrected, this inflates the apparent ongoing performance of whichever channel acquired the customer originally, well past the point that channel had anything to do with the renewal decision.
The fix is separating first-order attribution from renewal reporting entirely: attribute the initial subscription order to the marketing touches that earned it, and report renewal revenue as a distinct line, driven by retention and product experience rather than by any paid channel, unless a specific win-back campaign demonstrably triggered that particular renewal. Blending the two into one attribution total makes an acquisition channel look responsible for revenue it stopped influencing the moment the first order shipped, and it’s one of the more common reasons a channel that looks strong on a blended attribution report turns out to be far weaker once first-order and repeat revenue are split apart and checked separately.
What does it cost to run this properly?
Triple Whale and Northbeam both publish pricing and plan tiers on their own sites, and both change often enough that a figure quoted here would likely be stale by the time you read it; check their current pricing pages directly rather than trusting a number from any third-party article, including this one. What a published price doesn’t include is the internal cost of actually using the tool well: someone has to own the methodology document, run the reconciliation cadence, and be the person a CFO calls when the numbers don’t match.
As an illustrative example, not a reported figure: a brand paying a monthly platform fee but spending zero analyst hours on reconciliation is, in practice, paying for a more detailed version of the same self-reporting bias every platform dashboard already has, just with better visualisation layered on top. The tool’s cost is the smaller half of the total cost of running defensible attribution; the larger half is the discipline of using it as an input to a documented methodology rather than as the methodology itself.
Where a store chooses to build reconciliation on Shopify’s own order and UTM data instead of paying for a dedicated attribution platform, the cost shows up entirely as analyst time rather than a subscription line, and that trade is worth making deliberately rather than by default in either direction. A data warehouse or reporting layer to hold reconciled figures over time is a further cost worth planning for once monthly reconciliation becomes a standing process rather than a one-off exercise, since a methodology that only lives in a single spreadsheet is one edit away from losing its own history.
Who doesn’t need a dedicated attribution model yet?
This article targets brands doing $3M to $30M in revenue on Shopify Plus or a comparable paid subscription platform, running enough channels and enough spend that the gap between platform totals and ledger revenue is large enough in dollar terms to change a real decision. Below that floor, the calculation is different, and it isn’t softened here on purpose.
A store running one or two channels at modest spend, without the volume to make a monthly reconciliation cadence worth a dedicated analyst hour, doesn’t need a documented multi-touch methodology or a paid attribution platform. A simple monthly comparison of total ad spend against total revenue, by channel, checked against the order ledger, answers most of the decisions that store actually has to make. Build the fuller model described in this article when the channel mix and the spend have grown enough that a wrong allocation decision costs more than the model does.
Reading attribution correctly, once the data has already landed, is a reporting and analytics problem at its core: the whole exercise exists to give a business a number it can act on and defend, and that’s the specific ground Pointerflow’s reporting and analytics service is built to cover.
Sources
No specific pricing figures or third-party statistics are quoted in this article. Triple Whale and Northbeam are referenced only as publishers of their own current pricing, which readers are directed to check on each vendor’s own site rather than trust as quoted here. The mechanism explanation for why platform-reported conversions exceed the order ledger, and the reconciliation method, are written from first-hand analysis, not from a published study.