Can You Even Learn This?

// gate three

This is the gate nobody walks through, and it's the one that decides whether you're running a CRO program or just redecorating on a schedule.

An A/B test is a measuring instrument, and every instrument has a resolution. Below it, the thing you're trying to see is smaller than the noise — and the test doesn't tell you that. It returns a number either way, and the number looks just as confident when it's meaningless.

minimum detectable effect

The smallest lift your traffic can distinguish from luck. It's set by three things, and only one of them is under your control:

your conversion rateLower rates are harder. A 1% base rate needs far more traffic than a 10% one, because almost every visitor is a "no" and you're listening for a change in the rare yes.
the lift you'd acceptThe only dial you have. Halving the lift you're willing to call a win roughly quadruples the traffic you need.
your trafficFixed, this month. It decides how long the instrument has to run.

The rough version, good enough to kill bad ideas in a meeting:

visitors per variant = 16 × p × (1 - p) / (p × lift)² — where p is your current rate and lift is relative. That's 95% significance at 80% power.

// the calculator

// fill in all three

Trail Runner 2's real numbers are in the placeholders: 1,204 sessions a month at 1.1%. Try asking for a 20% lift. Then try 10%.

what you just found out

At 1.1% and a 20% lift, Trail Runner 2 needs about 36,000 visitors per variant. It gets 1,204 a month. That's roughly five years to run one test on your best-trafficked product page.

And it isn't a small-store problem. A store doing 18,000 sessions a month on a page converting at 3.8% still needs twenty weeks to detect a 10% lift. Nearly every A/B test ever run on a boutique Shopify store was underpowered. The ones that "won" mostly measured the weather.

This is also the honest answer to that 100% confidence column from earlier. Thirty human sessions makes Snapshot's progress bar full — it means the snapshot has enough data to describe the page. It has never meant you can prove a change to it. Two different questions, and only one of them is a statistics question.

so what do you do instead

Stop treating the A/B test as the only way to be rigorous. It's one instrument, with a traffic requirement most pages can't meet. The rest of the toolkit:

A program that ships twenty well-reasoned changes a year and measures them honestly will beat one that runs four underpowered tests and believes the results.