Scoring the Queue

// the model

Three hypotheses is a conversation. Thirty is a backlog, and a backlog needs an order that survives the loudest person in the room.

ICE is the simplest model that does that: score every item on three things, one to ten, and sort by the average.

ImpactIf this works, how much does it move the metric?
ConfidenceHow sure are you it will work, given the evidence you actually have?
EaseHow cheap is it to ship? Ten is an afternoon, one is a quarter.

what it's really for

Not precision. A 7.3 is not better than a 7.0 in any meaningful sense, and anyone defending that gap is doing numerology. What the model buys you is three specific arguments instead of one vague one — when two people disagree on the order, ICE tells you which of the three they disagree about, and that argument is usually resolvable.

where it goes wrong

Score all three hypotheses on one criterion before moving to the next — all three impacts, then all three confidences. Scoring item by item makes you anchor on whatever you scored first.

// hypothesis one

Impact — how much does it move the metric?

1 = barely measurable10 = changes the quarter

Confidence — how sure are you, from the evidence?

1 = a hunch10 = two cards agree

Ease — how cheap is it to ship and measure?

1 = a quarter10 = an afternoon

// score once all three are set

// hypothesis two

Impact — how much does it move the metric?

1 = barely measurable10 = changes the quarter

Confidence — how sure are you, from the evidence?

1 = a hunch10 = two cards agree

Ease — how cheap is it to ship and measure?

1 = a quarter10 = an afternoon

// score once all three are set

// hypothesis three

Impact — how much does it move the metric?

1 = barely measurable10 = changes the quarter

Confidence — how sure are you, from the evidence?

1 = a hunch10 = two cards agree

Ease — how cheap is it to ship and measure?

1 = a quarter10 = an afternoon

// score once all three are set

Hold onto the order these produce. On the next page you'll find out that the top-scoring one may not be the one you run — and that's not a flaw in the scoring.