asteep.
How-To Guides

How to Prioritise CRO Tests When Everyone Has a Favourite

Jonathan Whitaker Jonathan Whitaker · · 7 min read
How to Prioritise CRO Tests When Everyone Has a Favourite

Most teams I work with do not have a CRO problem. They have a queue problem. There are forty ideas in a spreadsheet, everyone has a favourite, and the one that gets built is whoever’s argument was loudest in the meeting.

Prioritising CRO tests properly is dull work that pays better than any single clever test. This is how I sort a backlog, which scoring frameworks are worth the effort, and the check that decides whether a test is worth running at all.

Why “start with the homepage” is bad advice

The homepage gets suggested first because it gets the most visitors. That reasoning skips the part that matters: how many of those visitors were going to convert anyway, and how many were never going to.

A homepage carries mixed intent. Existing customers looking for a login, job applicants, journalists, and somewhere in there the people you want to persuade. Changing a headline moves all of them slightly and none of them decisively. A checkout step, by contrast, is pure intent — everyone there has already decided to buy, and anything that removes friction converts directly into revenue.

Sort your pages by traffic multiplied by how close the page sits to the money, and the queue reorders itself immediately. Product pages, pricing, checkout and the highest-traffic landing pages come first. The homepage usually lands fourth or fifth, and the “About us” redesign somebody has been pushing for drops off the list.

Prioritising CRO tests by traffic and proximity to the conversion, with checkout and pricing ahead of the homepage
Traffic alone is a poor sort. Traffic against intent is a good one.

The two scoring frameworks, honestly assessed

Two frameworks come up constantly. Both are useful. Neither is as objective as it looks, and knowing where the softness sits is what stops them being theatre.

What you scoreBest forWhere it goes soft
ICEImpact, Confidence, Ease — each 1 to 10A quick sort of a long backlogImpact and Confidence are both guesses, and the same person usually supplies both
PIEPotential, Importance, EaseChoosing between pages rather than between ideasPotential and Importance overlap heavily; scores cluster and stop separating anything

My working compromise: score with ICE, but require evidence for Confidence. A number above seven has to point at something — session recordings, a funnel drop-off, support tickets, a previous test on a similar page. Without that rule, Confidence becomes a measure of enthusiasm, and the loudest voice wins by a different route.

Ease deserves more weight than it usually gets. A test that takes two days to build and runs for three weeks beats one that needs a developer sprint and runs for the same three weeks, even when its projected impact is smaller — because you get to run four of them in the time the big one takes. Velocity compounds; single tests do not.

The check that comes before scoring

Before anything gets scored, one question decides whether it belongs in an A/B queue at all: does this page get enough conversions for a test to finish?

Not enough traffic — enough conversions. A page with 20,000 monthly visitors and 40 conversions cannot detect anything short of an enormous swing, and it will take months to say so. Teams run these tests anyway, call the result at three weeks because a colour looks ahead, and ship a change that did nothing.

Put the page’s current conversion count and rate into any sample size calculator before it enters the backlog. If the required runtime lands beyond six weeks, the honest answer is that this page is not testable right now, and the idea moves to a different queue.

Scoring CRO tests with ICE, requiring evidence for confidence and weighting ease for velocity
Score the idea only after the page passes the volume check.

What to do when the numbers are too small

Most sites are in this position, and it is not a reason to stop optimising. It is a reason to stop pretending that A/B testing is the only method.

  • Ship obvious fixes without testing them. A broken mobile form, a checkout that fails on one browser, a page that takes nine seconds to load. There is nothing to learn from an experiment here — the current version is defective.
  • Watch session recordings instead. Twenty recordings of people abandoning the same step tell you more than an underpowered test ever will, and they take an afternoon.
  • Ask five customers. The reason for the drop-off is often something nobody on the team considered, and it usually cannot be inferred from click data.
  • Test at the funnel level rather than the page level. Aggregating a step across all product pages gives you enough volume to read, where each page alone gives you noise. My guide to conversion funnel analysis covers how to structure that.

The sites that improve fastest at low volume are the ones that stop running experiments and start fixing defects. Testing is for questions where reasonable people disagree, not for confirming that a broken thing is broken.

Keeping the backlog honest

A CRO backlog with hypothesis, evidence, ICE score, required runtime and outcome recorded for each test
Five columns keep a backlog from becoming a wish list.

A backlog that records only ideas turns into a wish list within a quarter. Five columns keep it useful: the hypothesis in one sentence, the evidence behind it, the ICE score, the runtime the volume check produced, and — once it has run — the outcome.

That last column is the one teams skip and the one that makes the whole thing worth maintaining. Recording losses stops the same idea returning every six months under a new name, and after a dozen entries you can see which kinds of change actually move your particular audience. That pattern is worth more than any framework’s score.

Review the queue monthly and delete freely. An idea that has sat unloved for three cycles is telling you something, and a shorter list gets acted on.

Common questions

How many tests should run at once?

As many as you can keep on separate pages or separate audiences. Two tests on the same page in the same period contaminate each other, and untangling the result afterwards is guesswork. Most teams are limited by build capacity long before they hit that ceiling.

Is ICE better than PIE?

ICE separates ideas better because Ease is genuinely objective, while PIE’s Potential and Importance overlap enough that scores bunch together. Either framework works if you enforce evidence for the subjective inputs. Neither works as a substitute for judgement.

Should a losing test be reverted immediately?

Revert, then write down why you thought it would win. The gap between the prediction and the result is the most useful thing a losing test produces, and it usually reshapes the next three hypotheses more than a win would.

What if leadership insists on a specific test?

Run it, and score it in the backlog like everything else. Then show the runtime the volume check produced. “This will take eleven weeks to read” is a conversation that reprioritises itself, and it lands better than an argument about which idea is cleverer.

Where to start on Monday

Take your existing list of ideas and run the volume check against each one before scoring anything. Half the queue usually disappears at that stage, which is the point — those pages need defect fixes and customer conversations, not experiments.

Score what survives with ICE, demand evidence for every Confidence above seven, and build the cheapest one first. The specific tactics that tend to earn their place in that queue are in my list of CRO practices that work — but the order you run them in will do more for your results than the tactics themselves.

Jonathan Whitaker

Jonathan Whitaker

Marketing analyst · Vancouver

Marketing analyst and CXL-certified optimizer with 6+ years of experience in web analytics, conversion optimization, and privacy-first data strategy. Former analytics lead for e-commerce and SaaS companies across North America, now focused on helping businesses make better decisions with less data. Specializes in Plausible, Umami, Matomo, and cookieless analytics. Based in Vancouver, BC.