Can You Even Learn This?
// gate three
This is the gate nobody walks through, and it's the one that decides whether you're running a CRO program or just redecorating on a schedule.
An A/B test is a measuring instrument, and every instrument has a resolution. Below it, the thing you're trying to see is smaller than the noise — and the test doesn't tell you that. It returns a number either way, and the number looks just as confident when it's meaningless.
minimum detectable effect
The smallest lift your traffic can distinguish from luck. It's set by three things, and only one of them is under your control:
| your conversion rate | Lower rates are harder. A 1% base rate needs far more traffic than a 10% one, because almost every visitor is a "no" and you're listening for a change in the rare yes. |
| the lift you'd accept | The only dial you have. Halving the lift you're willing to call a win roughly quadruples the traffic you need. |
| your traffic | Fixed, this month. It decides how long the instrument has to run. |
The rough version, good enough to kill bad ideas in a meeting:
visitors per variant = 16 × p × (1 - p) / (p × lift)² — where p is your current rate and lift is relative. That's 95% significance at 80% power.
// the calculator
// fill in all three
Trail Runner 2's real numbers are in the placeholders: 1,204 sessions a month at 1.1%. Try asking for a 20% lift. Then try 10%.
what you just found out
At 1.1% and a 20% lift, Trail Runner 2 needs about 36,000 visitors per variant. It gets 1,204 a month. That's roughly five years to run one test on your best-trafficked product page.
And it isn't a small-store problem. A store doing 18,000 sessions a month on a page converting at 3.8% still needs twenty weeks to detect a 10% lift. Nearly every A/B test ever run on a boutique Shopify store was underpowered. The ones that "won" mostly measured the weather.
This is also the honest answer to that 100% confidence column from earlier. Thirty human sessions makes Snapshot's progress bar full — it means the snapshot has enough data to describe the page. It has never meant you can prove a change to it. Two different questions, and only one of them is a statistics question.
so what do you do instead
Stop treating the A/B test as the only way to be rigorous. It's one instrument, with a traffic requirement most pages can't meet. The rest of the toolkit:
- ship it and watch the snapshot. Make the change, run a fresh snapshot of the same size, compare. You lose the clean causal claim and you keep the learning — and because a snapshot is a fixed number of real humans rather than a date range, the comparison is a lot fairer than a before-and-after date range would be
- raise the bar instead of the traffic. Don't test for 5% lifts you can't see. Make changes big enough to show up — a reworked gallery, not a button colour
- test on the page type, not the page. One PDP can't power a test; all forty of them together can. Roll the change across the template and measure the template
- measure the intermediate step. Add-to-cart happens twenty times more often than checkout, so it moves out of the noise twenty times sooner. Scroll depth and zone clicks sooner still
- just fix the broken thing. If the size guide is unreachable on mobile, that is a bug. Nobody A/B tests a bug
A program that ships twenty well-reasoned changes a year and measures them honestly will beat one that runs four underpowered tests and believes the results.