Home Blog Contact
Home/Blog/How to Run Ecommerce Email A/B Tests That Mea…
How toEmail MarketingEmail MarketingA/B TestingEcommerce

How to Run Ecommerce Email A/B Tests That Mean Something

5 min readBy Miloš Mitrović

Most email A/B tests in ecommerce prove nothing. A store owner sends subject A to one group and subject B to another, sees A win by two points on open rate, declares victory, and rolls it out. The next week the same test would show B winning, because the difference was noise the whole time. Testing without the discipline behind it is worse than not testing, because it hands you false confidence.

Here is how to run tests whose results hold up.

Key takeaways

  • Change one variable per test, or you cannot repeat the win.
  • Measure revenue per recipient; open rate is inflated and easy to misread.
  • Conversion tests often need 5,000 or more recipients per version to be reliable.
  • Split randomly, run 24 to 48 hours, and roll out at 90 to 95 percent confidence.

Test one thing at a time

If you change the subject line and the send time and the hero image all at once, and revenue goes up, you have learned nothing about why. You cannot repeat what you cannot isolate. Change one variable per test. The other elements stay identical across both versions.

That does not mean every test is trivial. The variable can be a big idea, discount framing versus benefit framing, as long as everything else is held constant.

Test the things that move money

Rank your tests by how much they can change revenue, not by how easy they are to set up. Roughly in order of impact:

  • Offer and framing. Free shipping versus 15 percent off. A bundle versus a single item. This moves purchase behavior more than anything else.
  • Subject line and preview text. These decide whether the email gets opened at all, so they gate everything downstream.
  • Email structure. One product focus versus a grid. Where the call to action sits.
  • Send time and day. Real, but a smaller effect than most people expect. Test it after the higher-impact items.

Button color and font tweaks are near the bottom. They rarely change revenue enough to detect.

Measure the metric closest to money

This is where most tests go wrong. Open rate is easy to measure and almost meaningless on its own. A clickbait subject can win opens and lose sales because it drew the wrong attention. Since Apple Mail Privacy Protection inflates opens by auto-loading images, open rate is even shakier than it used to be.

Measure conversions and revenue per recipient. Revenue per recipient is total revenue from the version divided by how many people received it. It accounts for both how many bought and how much they spent. When the money metric and the open metric disagree, trust the money, and reconcile it against your store’s own numbers before you rely on it.

Track the funnel so you can read the story: opens, then clicks, then conversions, then revenue per recipient. A version can win opens, lose clicks, and still win revenue. The full funnel tells you what happened.

Get the sample size right

This is the part that separates a valid test from a guess. A narrow gap between two versions needs a large sample to detect. If version A converts at 2.0 percent and B at 2.4 percent, you need thousands of recipients per version to know the gap is not chance.

Two working rules:

  • For open-rate tests, aim for at least 1,000 recipients per version, and treat differences under one point as noise.
  • For conversion and revenue tests, you often need 5,000 or more per version, because purchases are rarer events and rare events are noisier.

If your list is 3,000 people, you cannot reliably test conversion rate on a single send. Do not pretend otherwise. Either test the higher-frequency metric, or accumulate results across sends before you call it.

Run the test clean

  • Split randomly and at the same time. If A goes out Tuesday morning and B goes out Thursday night, you tested the day, not the content. Send both at once to randomly assigned halves.
  • Let it run long enough. Ecommerce purchases trail the send. Someone opens at lunch and buys that evening. Give a campaign test at least 24 to 48 hours before reading revenue, and give flow tests longer because entries trickle in.
  • Do not peek and stop early. Watching a test and stopping the moment one side is ahead manufactures false winners. Decide your sample size and duration in advance, and read the result when you get there.

Know when a result is real

A difference is worth acting on when it clears two bars: it is large enough to matter to the business, and it is unlikely to be chance. Most email platforms report a confidence level or declare a winner using significance testing. Aim for 90 to 95 percent confidence before you roll out. Below that, log it as inconclusive and move on. Inconclusive is a legitimate outcome, not a failure. It tells you the change does not matter enough to detect, which is itself useful.

Build a testing calendar

One-off tests teach you one-off lessons. A calendar compounds. Run one structured test per week, write down the hypothesis before you send, and keep a log with the date, the variable, the sample size, the winner, and the confidence level. After three months you have a documented playbook for your specific customers, which is worth more than any generic best-practice list.

The log also stops you from re-testing settled questions and from forgetting what you learned when a new person touches the account.

A note on flows

Flow emails are where testing pays off most, because a winning version keeps earning every time the flow fires for months. Test your abandoned-checkout and welcome messages first. A subject-line win on a one-time campaign helps once. The same win on an automation helps thousands of times.

Takeaway

Test one variable at a time, measure revenue per recipient instead of opens, and respect sample size so you are reading signal and not noise. Split randomly, wait for purchases to land, and roll out at 90 percent confidence or higher. Log every test. Over a quarter, that discipline turns guesses into a playbook built on your own customers.

Sources

M
Miloš Mitrović
Email Marketing for Ecommerce

Have a question or a project?

Whether it is about this post or a system you want built, I'm happy to talk.

Get in touch

404

Post not found. It may have been moved or the link is incorrect.

← Back to the blog
Summarize with AI
ChatGPT, Perplexity, and Grok open with the prompt ready to run. Claude, Gemini, and Copilot open a chat with the prompt copied; press Ctrl+V (Cmd+V on Mac) to paste. The full text is included, so it works even without web access.