How to Find What to Improve on Your Website

15 min read

In this article

  • Key takeaways
  • Is pairing quantitative and qualitative data your actual next step?
  • Where do you start looking?
  • How do you figure out what people are actually doing?
  • What behaviors should you look for?

Finding what to improve on a website means pairing two kinds of data: quantitative analytics that show where a page leaks, and qualitative behavior, heatmaps, session recordings, that shows why visitors leave without converting. Closing that gap turns a guess into a testable hypothesis instead of a redesign based on opinion.

Most teams know roughly which pages matter, the ones with the most traffic, or the ones closest to a sale. The harder question is what to actually change on those pages, the core diagnostic work of conversion optimization. Knowing where people go tells you nothing about what they do once there.

Key takeaways

  • Quantitative data (analytics) shows which page is losing visitors; qualitative data (heatmaps, recordings) shows why.
  • Watch for people not clicking the CTA, abandoning a form out of order, or scrolling past the section you care about.
  • Pairing the where with the what turns a guess into a testable hypothesis, written in three parts: element, observed behavior, expected result.
  • A hypothesis can look structured and still not be testable, filling in a template with a guess isn’t the same as filling it in with an observation.
  • When two signals contradict each other, or two “confirming” signals turn out to be the same session pool restated, resolve it before you write the hypothesis, not after you’ve built the test.
  • The why, why people convert or don’t, is the next layer once you have the where and the what.

Is pairing quantitative and qualitative data your actual next step?

This layer only helps once you already know which pages matter and need to know what’s happening on them. Check whether that’s genuinely where you are before diving into heatmaps.

This isn’t your priority yet when:

  • You haven’t identified your highest-traffic, most important pages yet, in which case where to start your conversion optimization plan or where your funnel leaks comes first.
  • You already have a specific, well-formed hypothesis backed by real behavior data, in which case the honest next step is running the test, not gathering more qualitative evidence to confirm what you already know.
  • Your traffic on this page is too low for a qualitative tool to show a repeatable pattern yet, in which case the fix is accumulating enough representative sessions to see the pattern repeat, not switching tools. That takes longer on a low-traffic page only because it takes longer to accumulate enough visits, not because evidence quality improves with the calendar.

The division of labor: deciding where to look is where to start your conversion optimization plan‘s job. Everything from here on assumes that decision is made and is entirely about what’s wrong on the page you’ve already chosen.

If none of those describe where you are, keep reading.

Where do you start looking?

Start with the pages that have the most to gain: your highest-traffic pages and the ones tied most directly to a conversion, a purchase, a signup, a lead form. Your analytics tool (Google Analytics or whatever you already run) is good at telling you this. It’s quantitative data: it answers where people are and how many of them drop off. That’s the map, it tells you which rooms are crowded, but not what people are tripping over once they’re inside.

How do you figure out what people are actually doing?

Quantitative data tells you a page is leaking; it won’t tell you why. For that you need qualitative data, a record of real behavior on the page itself. The everyday tools here are heatmaps (where people click and how far they scroll), session recordings (a replay of a real visitor’s path through the page), and form-interaction tracking (which fields people fill, skip, or abandon).

A heatmap shows whether anyone reaches your call to action, or whether they all give up above it. A session recording lets you watch a visitor hesitate, scroll back up, hunt for something they can’t find, and leave. Form tracking shows the exact field where people stall. Tools like Hotjar bundle these together, but the value is in the behavior they expose, not the brand of the tool.

What behaviors should you look for?

Watch for three patterns, because each points straight at a fixable problem: people not clicking the call to action, people abandoning a form out of order, and people scrolling past the section you care about.

  • People aren’t clicking the call to action (or the links tied to your conversion goal). That’s a sign the action isn’t clear, isn’t compelling, or isn’t where their attention is. It’s a prompt to rethink the copy, the placement, or the offer.
  • People aren’t completing a form in a sensible order, jumping around, backtracking, abandoning partway. That usually means the form is too long or asks for something people aren’t ready to give. The fix is often removing fields, not redesigning the page.
  • People scroll right past the section you care about. If almost no one reaches it, the problem may be that it’s too far down, not that the content is wrong.

Each of these is an observation, not yet a conclusion. The point isn’t to react to one recording, it’s to spot a pattern across enough of them that you can trust it.

Pair two data sources into one testable hypothesis Quantitative analytics: answers WHERE a page leaks visitors Qualitative heatmaps, recordings: answers WHY visitors leave without converting Testable hypothesis grounded in real behavior
This figure shows how quantitative data (which page is losing visitors) and qualitative data (why visitors leave) combine into a single hypothesis specific enough to test.

Alt text: Two boxes, quantitative analytics and qualitative heatmaps/recordings, converging into a single testable-hypothesis box.

How does this become a test?

Once you’ve paired the where (this page leaks) with the what (people stall at this specific element), you have the raw material for a hypothesis, the difference between guessing and testing. Write it in three parts so it stays specific enough to actually test:

  1. Name the element. The exact thing you’d change, a form field, a headline, a CTA button, not “the page” in general.
  2. Name the observed behavior. The specific pattern from your heatmap, recording, or form data that points at this element as the problem.
  3. State the expected result. “If I change [element], because [observed behavior], then conversions should improve because [reasoning].”

A hypothesis grounded in real behavior is something you can run an experiment against and actually learn from, whichever way the result lands.

Two vague ideas, turned into testable hypotheses

Most starting ideas arrive vague: “improve the pricing page,” “fix the checkout.” Here’s what happens when you push each one through the element/behavior/expected-result structure instead of stopping at the vague version.

The pricing page: “The pricing page needs work”

That’s a hunch with no element named. Pull the two data sources apart before writing anything.

Quantitative data shows the pricing page has the highest exit rate in the signup funnel, that’s where, not what. Session recordings show the pattern: visitors scroll to the plan comparison table, scroll back up to the plan names, scroll back down, two or three times, before picking a plan or leaving. The heatmap confirms it, a band of repeated hover-and-scroll activity sits right at the boundary between the plan names and the feature rows below. That’s an observed behavior, not a guess: visitors can’t hold the plan name and its features in view at once, so they’re re-reading to reconstruct a comparison the page isn’t making for them.

  • Element: the plan comparison table’s layout, specifically that plan names scroll out of view while the visitor reads the feature rows below.
  • Observed behavior: repeated up-and-down scrolling between the plan-name row and the feature rows, visible in both the heatmap and the recordings, concentrated at that exact boundary.
  • Expected result: if the plan names stay visible while the visitor scrolls the feature rows (a sticky header on the comparison table), the back-and-forth scrolling should drop and more visitors should reach a plan-selection click, because the comparison they’re currently reconstructing by memory would instead stay in view.

“The pricing page needs work” could mean a dozen fixes, new copy, a different price, fewer plans. The version above names one element, ties it to one observed behavior, and states one measurable result, that’s the difference between the vague idea and something you can build a test around.

The checkout page: “Checkout feels clunky”

Same problem, different page. “Clunky” is a feeling, not an element, and a feeling doesn’t tell a designer what to build.

The funnel report shows checkout, after cart, before payment, has the sharpest single drop in the path. Form tracking narrows it further: most visitors reaching the shipping-address field fill in the street address, then stall on the phone-number field before finishing or abandoning. Recordings confirm it, visitors clicking into the phone field, clicking back out without typing, then clicking back in.

  • Element: the phone number field on the shipping-address step of checkout.
  • Observed behavior: visitors hesitate at the phone field specifically, shown by a stall in form tracking and repeated click-in/click-out behavior in recordings, not a uniform slowdown across the whole form.
  • Expected result: if the phone field is marked optional (or removed, if it isn’t used for delivery notifications), the stall should disappear and checkout completion should improve, because the hesitation is about being asked for contact information a delivery form isn’t expected to need, not about the form’s length overall.

Both examples start from the same kind of sentence anyone says out loud in a meeting. The difference between that sentence and a testable hypothesis is naming one element, tying it to one observed behavior, and stating one measurable result with the reasoning attached.

When a hypothesis looks structured but isn’t testable

Filling in the template doesn’t automatically produce a testable hypothesis. The most common failure mode isn’t skipping the structure, it’s filling each blank with something that sounds like an observation but is really a restated opinion.

Three patterns show up often enough to name:

  • The “observed behavior” is a guess wearing a data citation. “Visitors probably find the CTA unclear” isn’t an observed behavior, even with a heatmap screenshot attached. A real one names what people did (scrolled past a section, stalled on a field). If the sentence would still be true without ever watching a recording, it’s an opinion, not an observation.
  • The expected result has no way to fail. “If we improve the design, conversion should improve” is unfalsifiable, almost any outcome can be read as confirming it afterward. A testable result names a specific metric and direction tied to the specific behavior observed, so a flat or negative result is a real answer, not something you can explain away.
  • The element is too broad to change in one test. “The onboarding flow” isn’t an element, it’s a whole section of the product. If proving the hypothesis means touching five screens, the element hasn’t been named yet, it’s been gestured at. Push until it’s one field, one headline, one button, one layout decision.

The tell: a hypothesis that survives someone who watched the same recordings pushing back with “how do you know that’s why,” where you can point to the specific clip rather than restate your reasoning more confidently. If the only answer is a firmer tone, the template got filled in, but the hypothesis underneath it is still a guess.

When analytics and the heatmap disagree

Two real signals sometimes point in different directions, and the fix isn’t to trust whichever one you find more convincing. Say analytics shows a page’s exit rate is unremarkable, nothing flags it as a leak, but the heatmap shows almost nobody scrolls far enough to see the CTA. That looks like a contradiction. It usually isn’t one.

Work through it in order, before writing a hypothesis off either signal alone:

  1. Check whether they’re measuring the same slice of traffic. An aggregate exit rate blends mobile and desktop, paid and organic, new and returning visitors. A heatmap session pool is often narrower, sometimes a single device type or a single week. Segment the analytics along the same axis the heatmap used before assuming the two disagree.
  2. If they still disagree after segmenting, that’s real, not a triangulation failure. It usually means the two signals are describing different parts of the same problem, not the same part twice. Here, a normal exit rate paired with nobody reaching the CTA can mean people are leaving from a different, earlier point on the page, or converting through a path the heatmap isn’t capturing, not that either tool is wrong.
  3. Find a third signal that explains the gap rather than picking a side. A session recording, or a differently segmented analytics cut, usually resolves which read is right for this page, or shows that both are right for different visitors.

Don’t average the two readings or split the difference in the hypothesis. A hypothesis built on an unresolved contradiction is a guess with two data points attached instead of zero.

Two signals, or one signal counted twice?

A heatmap and a recording pulled from the same week of traffic aren’t two confirmations, they’re one observation viewed two ways. Triangulation only works when the signals could have failed independently, meaning a different collection method, a different session pool, or a different time slice, not the same underlying sessions replayed through a second tool.

Before citing “the heatmap and the recordings agree” as confirmation, check whether they were pulled from the same session pool. If forty recordings and the heatmap all come from the same three-day traffic spike, that’s one data point stated twice, not two independent ones. Real independence looks more like: a heatmap aggregated across a full month, a funnel report pulled from a separate analytics layer, and a form-abandonment log, three collection methods, three different vantage points on the same page, each of which could have shown a different answer and didn’t.

How confident do you need to be before you build a test around it?

A hypothesis is ready to test once at least two independent signals point at the same element and the pattern holds across a representative sample, enough sessions to see it repeat, not one dramatic clip. Below that, it’s a lead worth investigating further, not a test worth spending a slot on. Three things build confidence, and it’s worth checking a candidate hypothesis against all three before treating it as actionable:

  • Consistency. The pattern recurs across a meaningful, representative slice of sessions, not just the one recording you happened to watch and remember.
  • Independence. The pattern shows up in genuinely separate signals, not the same signal restated in a different tool’s dashboard (see above).
  • Specificity. The pattern points at one element, not a general sense that “something’s off” on the page.

If you only have one signal, or the pattern only showed up in a handful of recordings from a single afternoon, that’s a candidate worth watching more sessions for, or finding an independent second signal for, not something to build a test around yet. Testing a hypothesis you’re not actually confident in doesn’t save time, it spends a test slot finding out what more watching would have told you for free.

Frequently Asked Questions

What if I don’t have enough traffic for a heatmap tool to show a clear pattern?

What you need is more observed sessions, not more elapsed time, and those aren’t the same thing. A qualitative tool becomes trustworthy once it has seen enough representative traffic for a pattern to repeat, not once a certain number of days has passed. On a low-traffic page, accumulating that many sessions takes longer in calendar terms, which is why “wait longer” looks like the answer. But if a traffic spike or a new channel brings a flood of visits, the sample fills in days instead of weeks. Watch the session count, not the calendar.

The next layer is the why, why people convert or don’t, but the where and the what are what get you to a testable idea in the first place.

What to do next

  • Confirm the page is genuinely a top-traffic or top-conversion-value page before investing in qualitative research.
  • Watch session recordings until a behavior repeats, not just once, before treating it as a pattern.
  • Write the hypothesis in the three named parts and check each part against the failure-mode list above before calling it testable.

Continue based on what you need next

If you’ve got the analytics and the recordings but aren’t sure how to turn them into a test worth running, that’s exactly what a conversion review is for. Book a consultation and we’ll take a look.

Alex’s Perspective

The hypotheses that actually hold up under a test are never the ones built on a hunch about what “should” convert better. They’re the ones written after watching real visitors hesitate at the same field or scroll past the same section, over and over, until the pattern is undeniable. The template doesn’t do that work for you, it just forces you to show your evidence instead of your confidence.

Where this fits

This is primarily a Discover article, turning a page you know is underperforming into an evidenced diagnosis, that hands off directly into Experiment. It’s one instance of the evidence-gated decision loop described at the program level: Diagnose doesn’t clear until the evidence for a specific hypothesis actually exists, not before.


Written by Alex Harris, who runs AlexDesigns’ conversion optimization work directly, pairing analytics with real behavior data before any test gets built, not designing from opinion. Reviewed against the current hypothesis-writing process used on active client work.