“We’re optimizing the site” gets used for three different activities: a redesign, a round of ad-spend experiments, and an actual evidence-gated program. They produce different outcomes for different reasons, not just different amounts of effort. This article is the map: the five decisions an optimization program actually makes, in what order, on what evidence, and where the boundary sits against the other two. (For CRO’s own definition first, start with CRO for businesses with traffic, this article picks up from there.)
What is an optimization decision system? It’s the recurring process of deciding what’s actually broken, what deserves attention first, whether a proposed fix is true once tested, what the result teaches, and where a proven change should scale, run as a loop rather than a one-time redesign. Each decision is gated by the evidence the one before it produced.
| Decision | Question | Evidence gate | Output |
|---|---|---|---|
| Diagnose | Where is revenue actually leaking? | Analytics and session data showing a real, comparable gap | A named page, step, segment, or device |
| Prioritize | What earns the next slot? | The diagnosis, plus a rough size of the traffic or revenue affected | A ranked sequence, not a list |
| Test | Is this specific fix true? | A result read against a real-effect-vs-noise standard | A validated (or ruled-out) hypothesis |
| Learn | What does the result mean for what’s next? | The completed test, plus a deliberate review | A rule that narrows the next decision |
| Scale | Does this deserve more reach? | The learning, checked against where else the condition exists | A reset baseline for the next Diagnose |
The five decisions, in order
03 Test
Experiment or Personalize — the method the evidence calls for
03 Test
Experiment or Personalize — the method the evidence calls for
-
01 Diagnose
Question: where is revenue leaking, specifically
Gating evidence: analytics/session/funnel data
-
02 Prioritize
Question: what's worth fixing first
Gating evidence: diagnosis + rough sizing
-
03 Test (Experiment or Personalize — the method the evidence calls for)
Question: is this specific fix actually true
Gating evidence: result vs. real-effect-vs-noise standard
-
04 Learn
Question: what did the result actually mean
Gating evidence: completed result + post-mortem
-
05 Scale
Question: does this deserve more traffic, more surface area, or neither
Gating evidence: learning checked against where else the condition exists
- Stage 5 (Scale) feeds back into stage 1 (Diagnose) — the loop never dead-ends.
Each decision below states the question it answers, the evidence required before the program is allowed to move to the next one, and the mistake that skipping it produces. Skip the evidence and the label still says “diagnosis” or “test,” but the decision underneath it is really just a guess with a title.
1. Diagnose: what’s actually broken, and where
Where is revenue leaking, specifically, not generally. “Conversion is low” is a symptom broad enough to be true of almost any site, not a diagnosis. A real diagnosis names a page, a step, a segment, or a device where the gap is measurable against a comparable baseline, the same page, same traffic source, same time window, compared against itself, not an industry number nobody can verify applies to your site.
The evidence gate: analytics, session behavior, and funnel data showing where visitors actually drop, not where a stakeholder assumes they drop. Skip it, and every decision downstream inherits an opinion wearing a diagnosis’s clothes.
2. Prioritize: what’s worth fixing first
Of everything that’s broken, what earns the next slot in a queue that can only move one or two things at a time. A diagnosis routinely surfaces more problems than a program can act on at once, that’s normal, not a failure of the diagnosis. Prioritization turns that list into a sequence, using effort, expected impact, and confidence in the evidence as the ranking inputs, not whoever raised their hand loudest in a meeting.
The evidence gate: the diagnosis itself, plus a rough sizing of how much traffic or revenue touches the broken step. Skip it, and a program either tries to fix everything at once, which collapses into a redesign, or fixes whatever’s easiest regardless of whether it matters.
3. Test: is this specific fix actually true
Does the change we believe will help actually help, measured, not assumed. “Test” here means acting on a prioritized problem, and the method varies with the evidence available: a formal A/B test where there’s enough volume to trust a result, a segment-specific experience where the base page already converts but different visitor groups need different things, or a straight fix where the defect is obvious enough that no test is needed to confirm it. A high-volume page with an ambiguous hypothesis calls for a test; an obvious, low-ambiguity defect (a broken control, a dead link) usually doesn’t. What doesn’t vary is the decision itself: don’t roll a belief out everywhere until it’s been checked against reality on a contained slice of it.
The evidence gate: a result read against a standard that separates a real effect from noise. A change that “feels like it worked” without that separation is exactly the failure mode this decision exists to prevent.
4. Learn: what did the result actually mean
What does this specific result, win, loss, or flat, tell us about the next decision, not just this one. A losing test isn’t wasted; it rules something out, which narrows what gets tried next. A winning test isn’t just a win to bank; it’s a signal about what the underlying friction actually was, which often points at other pages carrying the same friction. Treating every test as a closed, isolated event is how a program burns through ideas without getting smarter between them.
The evidence gate: the completed test result, plus a deliberate post-mortem step, not a passive assumption that the number speaks for itself.
5. Scale: does this deserve more traffic, more surface area, or neither
Now that we know something works, where else does it apply, and does it apply enough to be worth extending. A validated change doesn’t automatically deserve a site-wide rollout, sometimes it worked on one page for reasons specific to that page, and forcing it everywhere is a smaller, quieter version of the redesign mistake: an unverified assumption applied at scale.
The evidence gate: the learning from the prior decision, checked against where else the same underlying condition, the friction, the segment, the device, actually exists, not copied to every page because it worked on one. Scale closes the loop back into Diagnose: a change rolled out at scale resets the baseline the next diagnosis measures against, which is what makes this a loop instead of a five-step list that ends.
How this maps to the Experience Optimization Framework
This is a compressed, decision-level view of AlexDesigns’ Experience Optimization Framework, Discover, Prioritize, Experiment, Personalize, Learn, Scale, not a stage-for-stage restatement of it. Four stages map one-to-one: Discover to Diagnose, Prioritize to Prioritize, Learn to Learn, Scale to Scale. Experiment and Personalize aren’t a fifth and sixth decision in sequence, they’re two alternative methods for acting on a prioritized problem, both folded into this article’s single “Test” step, since which method fits which situation is its own judgment call that doesn’t change the decision being made. Naming five decisions instead of six, honestly, as a compression rather than a false match, is what makes this operable at the level a business actually has to act on.
How is this different from a redesign or growth-marketing spend?
Both can look identical from outside, all three result in a page that’s different from before. The difference is the mechanism, and the mechanism determines whether anyone can trust the result.
When a redesign changes headline, layout, imagery, navigation, and copy simultaneously and conversion moves afterward, up or down, there’s no way to know which change did it, or whether they canceled each other out. That’s not a criticism of redesigns as a practice: a brand rebuild, a platform migration, or a genuinely broken information architecture is a real, legitimate need, and a decision system isn’t the right tool for it. Likewise, if the honest diagnosis is “not enough people are showing up,” that’s an acquisition problem, and spend-based growth tactics are the right lever, more spend won’t fix a checkout step that loses two-thirds of the people who already arrived; it just makes that leak more expensive per visitor. Neither comparison makes optimization inherently better. It solves a different problem than either one.
"We're optimizing the site" gets said about three completely different activities — a redesign, growth-marketing spend, and this decision system.
| Row label | Redesign | Growth-marketing spend | This decision system |
|---|---|---|---|
| What changes | how a page looks, all at once, on taste | how many people arrive | one thing at a time, isolating the decision being made and gated by evidence. |
| Shape | a project with a start and an end date | an acquisition lever — legitimate, but a different question than what happens once visitors are here | a recurring loop with no end date. |
| What goes wrong when it's confused with optimization | nobody isolated a variable, so nobody can attribute the result | it just buys more visitors to abandon the same broken step | — |
Redesign
Growth-marketing spend
This decision system
Worked example: one issue through the loop
(Illustrative, composite example, not a specific named client engagement.)
The loop compounds because it doesn’t stop: each pass starts from the baseline the previous pass changed, the opposite of a redesign’s one-time, static outcome. A program that runs this honestly will produce losses as often as wins, that’s evidence working correctly, not a sign the program is failing. A program that only ever reports wins is either not running real tests or not reporting the real results.
A B2B software company's demo-request page converts well on desktop but noticeably worse on mobile.
-
01 Diagnose
Mobile demo requests lag desktop.
Evidence
Session recordings show mobile visitors abandoning specifically when they tap the 'company size' dropdown, a native-feeling control that actually renders as a broken custom widget on iOS Safari.
-
02 Prioritize
High volume, clear mobile friction.
Evidence
This page carries a meaningful share of total demo requests and ranks above two other queued fixes once traffic volume is weighed in.
-
03 Test
Custom dropdown vs. native select.
Evidence
This is a genuine test candidate: enough mobile traffic exists to reach significance, so the team runs an A/B test replacing the broken widget with a native mobile select element.
-
04 Learn
Native select wins on mobile.
Evidence
The variant clearly outperforms the broken control on mobile with no effect on desktop, confirming the friction was rendering-specific, not a copy or offer problem.
-
05 Scale
Rolled out to every matching form.
Evidence
The native select rolls out to every form on the site that shares the same component, resetting the baseline the next diagnosis will measure against.
Paid-search mobile visitors responded even more strongly to the fix than organic mobile visitors did… that signal doesn't get absorbed into 'Learn,' it becomes the next loop's Diagnose, feeding a genuine Personalize candidate.
- 01 Diagnose: Mobile demo requests lag desktop. — Session recordings show mobile visitors abandoning specifically when they tap the 'company size' dropdown, a native-feeling control that actually renders as a broken custom widget on iOS Safari.
- 02 Prioritize: High volume, clear mobile friction. — This page carries a meaningful share of total demo requests and ranks above two other queued fixes once traffic volume is weighed in.
- 03 Test: Custom dropdown vs. native select. — This is a genuine test candidate: enough mobile traffic exists to reach significance, so the team runs an A/B test replacing the broken widget with a native mobile select element.
- 04 Learn: Native select wins on mobile. — The variant clearly outperforms the broken control on mobile with no effect on desktop, confirming the friction was rendering-specific, not a copy or offer problem.
- 05 Scale: Rolled out to every matching form. — The native select rolls out to every form on the site that shares the same component, resetting the baseline the next diagnosis will measure against.
- Carried forward — seeds the next loop's Diagnose: Paid-search mobile visitors responded even more strongly to the fix than organic mobile visitors did… that signal doesn't get absorbed into 'Learn,' it becomes the next loop's Diagnose, feeding a genuine Personalize candidate.
What to do next
- Write down what your last three site changes were actually deciding: a diagnosis, a prioritization call, a tested hypothesis, a learning, or a scale decision. If none of them fit, they were probably a redesign in disguise.
- Before your next change ships, name which of the five decisions it represents and what evidence is gating it. If the honest answer is “none, it just seemed right,” that’s the gap this system closes.
- Before calling anything “evidence,” check whether it would actually change what you do next. If a diagnosis, a test result, or a learning wouldn’t change your next move either way, it isn’t gating a decision, it’s decoration.
- If your actual problem is “not enough visitors,” this isn’t a conversion fix, that’s an acquisition-lever problem, not a decision-system problem.
This five-decision model is the operating system behind how AlexDesigns runs ongoing optimization work for every retainer client, not just a framework on a slide. If you’re not sure which of the five decisions your own program is stuck on, that’s a conversation, not a guess. Get Alex’s perspective on where your program is losing the loop.
Alex’s Perspective
Across the 100+ optimization and experience programs and 7,000+ experiments and personalization experiences I’ve been part of, the programs that stalled almost never stalled because the team ran out of ideas. They stalled because someone skipped a decision, usually prioritization, occasionally the learn step, and the program quietly turned into “testing stuff” without anyone deciding to let that happen. The five decisions aren’t bureaucracy; they’re the difference between a program that gets smarter every quarter and one that just stays busy.
Alex Harris leads AlexDesigns’ conversion optimization practice, where this five-decision model is the operating system behind every engagement, not just the framework on a slide.


