Test It or Ship It
// gate three, continued
Every item in a real backlog carries a route as well as a score. The score says how much you want it; the route says how you'll find out whether you were right.
Three routes. Most of your backlog is not the first one.
| test it | Enough traffic to detect the lift you'd accept, in a time you'd actually wait. Rare, and worth protecting when you have it. |
| just ship it | You believe it, the evidence is good, the downside is small — and the page can't power a test. Ship it, snapshot it, move on. |
| instrument first | You can't tell whether it's broken because nothing is measuring it. No test, no fix — go make it measurable, then come back. |
routing it honestly
- a change you'd ship anyway if the test won doesn't need a test — it needs shipping
- a change you'd roll back if it lost does need a test, and you should find the traffic for it
- a fix to something plainly broken is never a test. Rage-clicking an unreachable size guide is not a preference to be honoured
- if the honest route is "instrument first", say so and put it at the top — everything behind it is blocked, and it's usually a day of work
The third route is the one that saves programs. On Trail Runner 2, nothing is tracking size-guide opens or variant changes — so two of the likeliest hypotheses can't be evaluated at all. That's not a reason to guess harder. It's a ticket.
one more thing about shipping
"Just ship it" is not the lazy option, and it shouldn't feel like a defeat. It's the honest one when the alternative is a five-year test. What makes it rigorous rather than reckless is the follow-through: a fresh snapshot of the same size afterwards, compared to the one you have. Change one thing at a time and the comparison still teaches you something, even without a control group.
What makes it reckless is shipping six changes at once and declaring victory on whatever the next month's revenue does.