Test It or Ship It

// gate three, continued

Every item in a real backlog carries a route as well as a score. The score says how much you want it; the route says how you'll find out whether you were right.

Three routes. Most of your backlog is not the first one.

test itEnough traffic to detect the lift you'd accept, in a time you'd actually wait. Rare, and worth protecting when you have it.
just ship itYou believe it, the evidence is good, the downside is small — and the page can't power a test. Ship it, snapshot it, move on.
instrument firstYou can't tell whether it's broken because nothing is measuring it. No test, no fix — go make it measurable, then come back.

routing it honestly

  • a change you'd ship anyway if the test won doesn't need a test — it needs shipping
  • a change you'd roll back if it lost does need a test, and you should find the traffic for it
  • a fix to something plainly broken is never a test. Rage-clicking an unreachable size guide is not a preference to be honoured
  • if the honest route is "instrument first", say so and put it at the top — everything behind it is blocked, and it's usually a day of work

The third route is the one that saves programs. On Trail Runner 2, nothing is tracking size-guide opens or variant changes — so two of the likeliest hypotheses can't be evaluated at all. That's not a reason to guess harder. It's a ticket.

one more thing about shipping

"Just ship it" is not the lazy option, and it shouldn't feel like a defeat. It's the honest one when the alternative is a five-year test. What makes it rigorous rather than reckless is the follow-through: a fresh snapshot of the same size afterwards, compared to the one you have. Change one thing at a time and the comparison still teaches you something, even without a control group.

What makes it reckless is shipping six changes at once and declaring victory on whatever the next month's revenue does.