The Loop
// the cadence
A backlog with no cadence becomes a graveyard. The third thing a program needs is the boring one: a loop that keeps turning when nobody is excited about it.
| 1 // snapshot | A store snapshot, same size every time. Where is the traffic, where is the friction, what changed since the last one. |
| 2 // review | Expert review the page that wins gate one. That's session three — PACE, one stage at a time. |
| 3 // backlog | Findings become hypotheses, hypotheses get scored and routed. That's today. |
| 4 // ship or test | One change at a time, by the route you gave it. |
| 5 // snapshot again | Same size, same page. The comparison is the whole point, and it's only fair if the size matches. |
Then back to one. The loop is the deliverable — not any single test in it.
running it without it dying
- one change at a time on a page. Two at once and you've learned that something worked. This is the rule people break first and regret longest
- fixed snapshot size, always. A snapshot is a fixed number of real humans, not a date range — that's what makes two of them comparable. Change the size and you've thrown away your baseline
- write the prediction down before you ship. Memory is generous about what you expected, and it's generous in whichever direction makes you look better
- keep the losers. A written record of what didn't work is the only thing that stops the same idea being re-proposed every quarter with a new name
- re-snapshot the store quarterly, not weekly. Less than a quarter and you're reading noise and chasing it
what to do while something is running
Work a different page. This is why the backlog is a queue and not a stack — the next item is usually somewhere else entirely, and a program that can only do one thing at a time stalls every time it ships.
Instrument the things you found you couldn't measure. Build the next review. The waiting is not downtime, and treating it as downtime is how teams end up stopping a test early because they got bored — which is, statistically, the same as not running it.
the honest version of all this
You will not run many A/B tests. You will ship a lot of well-reasoned changes, measure them as carefully as your traffic allows, be right about most of them, and be wrong in public about the rest. That is what a CRO program looks like on a store this size, and it compounds.
The alternative — waiting until you have the traffic to test everything properly — is just not doing CRO, with better excuses.