Free calculator · no email required
A/B test significance, including the money
Conversion rate is half the test. A variant that converts better and earns less per visitor is a loss that every conversion-only calculator reports as a win — so this one tests both, and tells you when they disagree.
What statistical significance means here
A result is statistically significant when the difference between two arms is larger than sampling variation comfortably explains — conventionally a p-value under 0.05. For an ecommerce test that has to be checked on two metrics: conversion rate, and revenue per visitor. They can move in opposite directions, and only the second one pays for anything.
Your test.
Prefilled with an illustrative test where the variant converts better and earns less — the exact case this page exists for.
Verdict
Do not ship on the conversion result. The variant converts better and earns less per visitor — which is the outcome a conversion-only test is blind to.
Conversion rate
2.50% 2.85% +14.1%
p = 0.0364 · significant
Revenue per visitor
$2.10 $1.94 -7.6%
p = 0.2753 · not significant
Conversion rate and revenue per visitor disagree. The variant is converting more people and earning less from them — this is the case a conversion-only calculator reports as a win.
- Visitors per arm neededfor a lift this size, 80% power
- 32,909
- Days to get thereat your current traffic
- 26
How a winning test loses money
The worked example above is not a contrived case. It is the most common shape of a false win in ecommerce.
- 1
You test something that lowers the barrier
A discount banner, a cheaper default bundle, a smaller pack size, free shipping at a lower threshold. All of them make buying easier.
- 2
Conversion rate rises, clearly and significantly
It should. You removed friction, or money, from the decision. The test tool turns green and the screenshot goes in the deck.
- 3
Order value falls by more than the conversion gained
Invisible to a conversion-only test, because it never looks at what the extra orders were worth. In the example: 14% more conversions, 19% less per order, and revenue per visitor down.
- 4
And the damage compounds
A lower order value also raises your effective payment cost and lengthens acquisition payback — see the margin calculator. The test that looked like a win moved three numbers the wrong way.
The maths, published
Three tests, and one thing this calculator refuses to do.
| Test | Method | Needs |
|---|---|---|
| Conversion rate | Two-proportion z-test, two-tailed, p < 0.05 | Visitors and conversions per arm |
| Revenue per visitor | Welch-style test on RPV, with variance derived as p·(σ² + µ²) − (p·µ)² | Mean and standard deviation of order value per arm |
| Required sample | Two-proportion sample size at 80% power, α = 0.05 | The observed rates |
| Assumed order-value spread | Never. Withheld and marked metric to confirm | — |
Revenue-per-visitor significance is undefined without variance. Inventing a coefficient of variation would produce a confident-looking verdict from an assumption the operator never made — which is the exact failure this page criticises in conversion-only tools. Reviewed 2026-09-09.
What this cannot tell you
Whether you stopped at the right moment. P-values assume a sample size fixed in advance. Checking daily and stopping the first time a threshold is crossed inflates false positives substantially — the fix is deciding the runtime before the test starts, not a better calculator.
Whether the effect lasts. Novelty moves early numbers, and seasonality moves everything. A two-week test in the last week of November is measuring November.
Whether the segments agree. A result that holds in aggregate can reverse for returning customers or on mobile. Aggregate significance is where the analysis starts.
Downstream effects. A discount that lifts first orders may lower repeat rate — a cost that lands months later, in the cohort curve rather than in the test.
Definitions
- p-value
- The probability of seeing a difference this large if there were no real effect. Under 0.05 is the convention, not a law.
- Revenue per visitor
- Conversion rate × average order value. The metric that pays.
- Statistical power
- The chance of detecting a real effect of a given size. 80% is the usual target and what the sample-size figure assumes.
- Peeking
- Checking results repeatedly and stopping when they look good. The most common way an ecommerce test lies.
- Standard deviation
- How spread out order values are. Two stores with the same average need very different sample sizes if one is far more spread.
Questions about A/B testing
How do I know if an A/B test is statistically significant?
Compare the conversion rates of the two arms against the variation you would expect from sampling alone. If the probability of seeing a difference this large by chance — the p-value — is under 0.05, the result is conventionally called significant. That is a statement about noise, not about money, which is why this calculator also tests revenue per visitor.
Why test revenue per visitor as well as conversion rate?
Because in ecommerce they can point in opposite directions. A cheaper bundle, a prominent discount or a simplified basket will often convert more people and earn less from each of them. A conversion-only test reports that as a clean win and you ship a change that costs money. Revenue per visitor is conversion rate multiplied by order value, and it is the number the business actually banks.
Why does the revenue test need a standard deviation?
Because significance is a question about spread, not just averages. Two stores with the same £84 average order value behave completely differently if one sells everything at £80–£90 and the other mixes £20 and £400 orders — the second needs far more traffic to distinguish a real change. Without the spread the test is undefined, so this calculator withholds the verdict rather than assuming a distribution to produce a confident-looking answer.
Where do I find the standard deviation of my order values?
Export the order values for the test period and run a standard deviation over them — one column in a spreadsheet, one function. Most analytics tools will not show it, which is precisely why almost nobody tests revenue per visitor and almost everybody ships conversion winners.
My test is not significant. Should I run it longer?
Only if the required sample is reachable. The calculator shows how many visitors per arm a lift of the size you observed would need, and how many days that takes at your current traffic. If the honest answer is eleven weeks, the useful decision is usually to test something with a bigger expected effect rather than to wait — small effects need enormous samples, and most stores do not have them.
Can I stop the test as soon as it hits significance?
No, and this is the most common way ecommerce tests mislead. Checking repeatedly and stopping the moment a threshold is crossed inflates the false-positive rate substantially — p-values assume a sample size fixed in advance. Decide the runtime before you start, and read the result at the end.
Is anything I enter sent to you?
No. The arithmetic runs in this browser tab, with no endpoint behind the page.
Related
Find out what you’re losing.
Before you commit to anything, we tell you exactly what you’re losing and what it costs to stop it. Two weeks. Fixed fee. Credited in full against any build you go ahead with.
- Fee
- $1,500–$3,000, fixed
- Duration
- Two weeks
- Credited
- In full, against any build
- You supply
- Read access + one 45-minute call