Your A/B test wasn't significant. It was noise.
Most Shopify stores don't have the traffic to detect the uplifts they're chasing. Here's the sample-size maths nobody in this industry particularly wants to show you.
Written by Alex Price

Here's a conversation we have most months. A client shows us a test they ran. Variant B is up 14%. Everyone's delighted. Then we ask how many orders were in it, and the answer is about ninety.
Ninety orders can't tell you anything about a 14% difference. It can barely tell you about a 40% one. That test didn't win — it wobbled, and it happened to wobble upward on the day somebody looked.
This isn't a criticism of anyone. The tools show you a big green number and a confidence percentage, and it's entirely reasonable to believe them. But the maths underneath is less forgiving than the dashboard suggests, and it's worth knowing where the floor is.
How much traffic you actually need
The number of visitors a test needs goes up roughly with the square of how small the effect is. Halve the uplift you're hunting for and you need about four times the traffic. That single fact explains almost every disappointing testing programme we've inherited.
As a rough guide, starting from a 2% conversion rate and looking for a decent chance of detecting the change:
- A 50% uplift — around 3,000 visitors per variant. Rare in the wild, but it happens on a genuinely broken page.
- A 20% uplift — around 20,000 per variant. This is a very good result.
- A 10% uplift — around 80,000 per variant. This is a realistic good result.
- A 5% uplift — around 310,000 per variant. Most stores will never measure this reliably.
Two variants, so double those. If you get 30,000 sessions a month, a 10% uplift takes you something like five months to call properly. Not five days.
What people do instead, and why it goes wrong
Stopping when it looks good
If you check a running test every morning and stop the first time it crosses significance, you will find winners. You'll find them in tests where both variants are literally identical, because you've given yourself dozens of chances to catch a random high point. Decide the duration before you start, and don't look at the result until you get there.
Running too many tests at once
Twenty tests a year at a one-in-twenty false positive rate means one confident, well-presented, entirely fictional winner every year. It'll get built into the theme and quoted in a deck for the next three.
Testing things that are too small
Button colours and microcopy are cheap to test and almost never move anything by enough to measure. If a change couldn't plausibly shift behaviour by a fifth, most stores can't detect it — so either make a bigger change or just ship the small one on judgement and stop pretending it was measured.
If you can't say in advance what you expect to happen and roughly how big it should be, it isn't a test. It's a change with a chart attached.
What to do if you're under the threshold
Plenty of stores we work with aren't big enough for classic A/B testing, and that's fine — it just means the programme looks different. Rather than pretending, we lean on four things.
- Research instead of testing. Session replays, support tickets and five usability sessions will find more in a fortnight than a year of underpowered tests.
- Bigger swings. Redesign the whole product page rather than moving one element, so the effect is large enough to see.
- Ship on judgement where the downside is small. If a change is obviously better and cheap to reverse, just make it.
- Measure at the top. Add-to-cart and checkout starts happen far more often than orders, so they reach significance much sooner.
None of that is as satisfying as a dashboard declaring a winner. It is, however, true — and it compounds, which the fictional winners never do.
If you want a second opinion on a test you've run, send it over. We'll tell you honestly whether it held up, including when the answer is that we'd have called it flat.


