The first thing to test is the part of the journey that is actually failing.

That sounds obvious. It is also why another subject-line test is often the wrong place to start.

If people are not opening, test what they see before the open. If they open but do not click, test the promise and the content. If they click but do not buy, look at the page, offer, product, and buying path. The email button has already done its job.

Start with the bottleneck. Run one clean test. Make one decision.

Why most email tests start in the wrong place

Subject lines are easy to test. Button colors are easy to change. Neither fact makes them important.

An A/B test is useful when it helps you make a decision. It is less useful when it produces a small chart, a winner badge, and absolutely no change in what you do next.

Before you build the test, ask:

  1. Where are people dropping out?
  2. What could reasonably explain it?
  3. What single change would test that explanation?
  4. What will we do with each possible result?

That last question matters. If A wins, what changes? If B wins, what changes? If neither result is clear, what happens next?

If the answer is “nothing,” save the audience.

Use the bottleneck-first test ladder

Move through the journey in order. Find the first meaningful break. Test there.

What you seeLikely bottleneckWhat to test firstMain metricGuardrails
Few people engage with the sendRecognition or expectationSender name, subject line, send contextClicks and qualified opensUnsubscribes, complaints
People open but do not clickPromise or contentLead, offer framing, product choice, CTA hierarchyClick rateOrders, unsubscribes
People click but do not buyPage, offer, product, or buying pathMessage match, landing page, price/offer clarityConversion rateRevenue per recipient, refunds
Orders happen but value is weakAudience or offer economicsSegment, product mix, threshold, bundleRevenue per recipientMargin, complaints
Unsubscribes or complaints riseRelevance or frequencyWho gets the email, why they get it, how oftenComplaint and unsubscribe ratesRevenue, clicks

This is not a promise that one metric tells the whole story. It is a way to choose the first useful question.

If people are not opening, test recognition

Open tracking is imperfect. Privacy features can record an open that a person did not make, and some real opens may not be recorded.

So do not treat open rate as a courtroom witness.

Still, a weak top of funnel can tell you that people do not recognise the sender, do not understand the subject, or do not expect the email.

Test one of those things at a time:

  • Sender name
  • Subject line
  • Preview text
  • The context set at signup
  • The timing of the send

Keep the rest of the email stable. Then look beyond opens. Did clicks improve? Did orders improve? Did unsubscribes or complaints move?

A subject line that wins more opens but sends the wrong people into the email is not much of a win.

If people open but do not click, test the promise

The subject line earned attention. The email did not earn the next step.

Now look at what readers found after the open:

  • Was the point clear in the first screen?
  • Did the content match the subject line?
  • Was there one obvious next step?
  • Was the product or offer relevant to this audience?
  • Did competing links turn the email into a tiny website?

Useful tests include the opening message, offer framing, product selection, proof, CTA wording, or the order of content.

This is where “make the button green” often appears. It might help. But first make sure the offer around the button gives anyone a reason to press it.

If people click but do not buy, stop blaming the email

A click means the email created enough interest for the next step.

If orders do not follow, check what happened after the click:

  • Did the landing page match the email?
  • Was the same product easy to find?
  • Was the offer still clear?
  • Did the page work properly on a phone?
  • Did shipping, price, stock, or checkout create friction?
  • Were the clicks from people—or security scanners and bots?

This may lead to a landing-page test, an offer test, or a product test. It may not lead to another email test at all.

That is fine. The goal is to find the problem, not keep the email team busy.

If people buy but revenue per recipient is weak, test the economics

Conversion rate can improve while revenue barely moves.

That can happen when the audience is too narrow, the product value is low, the discount is expensive, or the offer attracts orders that would have happened anyway.

Consider testing:

  • A broader or narrower audience
  • A different product category
  • A bundle instead of a blanket discount
  • A spend threshold
  • Full-price value against a promotional offer

Track revenue per recipient alongside total revenue. If margin data is available, use it. A test that produces more orders and less profit has not discovered free money.

If people leave or complain, test relevance

When unsubscribes or spam complaints rise, creative polish is rarely the first fix.

Check who received the email and why:

  • Were inactive people included?
  • Did customers receive overlapping campaigns and flows?
  • Was the same offer repeated too often?
  • Did the email match what people signed up to receive?
  • Did a low-quality signup source add people who never wanted the emails?

Test the audience, cadence, or reason to send. Inbox providers notice complaints. Customers notice being annoyed. Both are inconveniently observant.

Write the hypothesis before you build the test

Use this sentence:

Because we see [the problem], changing [one variable] for [this audience] should improve [the main metric] without hurting [the guardrails].

Example:

Because recent buyers open the campaign but rarely click, leading with the new product instead of the brand story should improve click rate without increasing unsubscribes.

This forces the team to name the problem, the change, the audience, the metric, and the risk.

It also makes weak ideas easier to spot before they reach the send button.

Change one thing

Klaviyo recommends testing one variable at a time so you can tell what caused the result. It also recommends using a large enough audience and limiting the number of variations. That is sensible. Four clever variations on a small list mostly create four small guesses.

Keep the test clean:

  • One meaningful variable
  • One main metric
  • A few guardrails
  • A clear audience
  • Enough time and volume to observe the result
  • No editing after the test is live

If you need to test a full creative direction, that can still be valid. Just be honest about what you learn. A complete redesign can tell you which version won. It cannot tell you which of the twelve changes caused it.

Pick the metric before the result exists

Choose the winning metric before the test runs.

Klaviyo reports win probability, lift, and the selected metric in campaign test results. It considers a result statistically significant when a variation reaches at least 90% win probability. That is useful evidence. It is not permission to ignore the business result.

Match the metric to the bottleneck:

  • Recognition problem: qualified engagement, not opens alone
  • Content problem: click rate
  • Buying-path problem: conversion rate
  • Economics problem: revenue per recipient or margin
  • Relevance problem: unsubscribes and complaints

Then keep guardrails. A click-rate winner that tanks revenue is not the winner you wanted.

Decide what happens next

Before launch, write three lines:

  • If A wins, we will…
  • If B wins, we will…
  • If the result is unclear, we will…

An unclear result is still a result. It may mean the change was too small, the sample was too small, the audience did not care, or the tested idea was not the real bottleneck.

Do not rerun the same test forever hoping the chart develops confidence.

Make the next decision:

  • Keep the current version
  • Roll out the winner
  • Test a bigger change
  • Move to the next bottleneck
  • Stop spending attention on a low-impact question

A simple pre-test checklist

Before you press start, confirm:

  • We know where the journey is breaking.
  • We have a clear reason for the test.
  • We are changing one meaningful thing.
  • The audience is large enough to learn something useful.
  • We chose the main metric in advance.
  • We named the guardrails.
  • We know what we will do after each possible result.

That is enough. You do not need a laboratory coat.

Test the problem, not the convenient part

Good A/B testing is not a calendar full of experiments. It is a decision system.

Find the bottleneck. Write the hypothesis. Change one thing. Measure the result that matches the problem. Then make the decision.

If you want a repeatable way to turn campaign results into weekly actions, get the Inbox Operating System. It gives you the scorecards, review rhythm, and decision rules to stop guessing what to fix next.

## Related resources

## Sources