Benchmarks & Data

How to A/B Test Cold Email: Sample Sizes That Actually Mean Something

To prove a cold email reply rate moved from 3% to 4% you need about 5,301 contacts per variant. To prove 3% to 3.5%, 19,743. To prove 3% to 6%, 749. Those figures use the standard settings for comparing two proportions: two-sided 95% confidence and 80% power. Most cold email A/B tests run at 100 to 300 contacts per variant, which means most of them are reading noise.

This is for anyone who has declared a winning subject line after 40 sends. It gives the sample-size table, explains why small lifts cost so much to prove, sets out what to test first when volume is limited, and shows how to run a test whose result survives next month. The stack that produces the numbers you are testing is in our guide to outbound sales infrastructure.

How many contacts per variant do you need to A/B test cold email?

It depends on the rate you start from and the lift you want to see. From a 3% reply rate, proving a jump to 6% takes 749 contacts per variant, to 5% takes 1,506, to 4% takes 5,301, and to 3.5% takes 19,743. The table gives the full set at two-sided 95% confidence and 80% power.

MetricFromToContacts per variant
Reply rate3.0%3.5%19,743
Reply rate3.0%4.0%5,301
Reply rate3.0%4.5%2,518
Reply rate3.0%5.0%1,506
Reply rate3.0%6.0%749
Reply rate2.0%3.0%3,826
Reply rate2.0%4.0%1,141
Reply rate5.0%7.0%2,213
Reply rate5.0%10.0%435
Open rate40%45%1,534
Open rate40%50%388
Cold email A/B test sample sizes per variant at two-sided 95% confidence and 80% power, from the standard two-proportion calculation: proving a 3% to 3.5% reply lift needs 19,743 contacts per variant, 3% to 6% needs 749, and a 40% to 50% open-rate lift needs 388.

The figures are per variant, so a two-way test needs double. They come from the standard two-proportion sample-size calculation that any statistics textbook or online calculator reproduces; there is nothing proprietary in them. 95% confidence means a 1 in 20 chance of declaring a winner when there is none. 80% power means a 1 in 5 chance of missing a real lift. Loosen either and the numbers fall, and so does what the result is worth.

Why does a smaller lift need so many more contacts?

Because the sample grows with the square of the lift you are trying to see. Halve the lift and you need roughly four times the contacts. From 3%, a 2-point lift to 5% needs 1,506 per variant, a 1-point lift to 4% needs 5,301, and a half-point lift to 3.5% needs 19,743.

This is why 'we tested two subject lines on 200 people' cannot hold. At 100 per variant and a 3% base rate, each variant expects 3 replies. One extra reply is a 33% 'lift'. Two extra is a 'doubling'. Both sit inside normal variation. The base rate matters as well: meetings run at 1% to 2% of contacts across a sequence, so testing on meetings needs thousands per variant even for a doubling. We test on replies and read meetings as the check. The benchmarks for each stage are in average open, reply, and meeting-booked rates for outreach.

What should you test first when volume is small?

Big swings. Test the segment (who gets it), the angle (the problem you lead with), and the offer (what you ask for) before subject-line wording, greeting, or send time. A big swing can double a reply rate, and a doubling from 3% shows at 749 per variant. Wording moves fractions of a point, and fractions of a point need five figures.

  1. Segment. The same email to two ICP segments. Targeting and copy move reply rates from 2% to 9%; timing moves them a point or two. Start where the range is widest.
  2. Angle. Cost against risk against speed, or a pain against a trigger event. The first line and the subject follow from the angle, so it is one variable, not three.
  3. Offer. A 20-minute call, a one-page teardown, or a single question. Changing the ask changes who replies.
  4. Sequence shape. Three emails against five, or the breakup on day 12 against day 18. What goes into each follow-up is in the cold email follow-up sequence.
  5. Then wording. Subject lines, body length, one personalised line against none, send time. Worth testing at 5,000 per variant, not 500. Our cold email subject line examples are a starting set for when you get there.

One variable per test. Change the angle and the subject together and a win tells you the pair beat the old pair, not which part did it. The exception is the whole-message test at small volume: pit two complete emails with different angles against each other, accept that you learn 'this one wins' and not why, and pull the next pair from a bank organised by angle, which is how the 15 templates we use are grouped.

Should you measure opens, replies, or positive replies?

Positive replies where volume allows, raw replies with auto-replies removed as the leading read, and opens not at all. Positive replies are what become meetings, but their base rate of 3% to 6% is lower than the 8% to 11% raw rate of disciplined outreach, so they need more contacts to move with confidence.

Opens are inflated by Apple Mail Privacy Protection and similar prefetching, so the cheap open-rate rows in the table (40% to 50% needs only 388) are a cheap read of a polluted number. Raw replies include 'not interested' and opt-outs, which still tell you the email was read and understood, so they are a fair leading indicator. Decide the metric before the test, define a positive reply as one that asks a question or accepts a next step, and never re-classify after you see the split. The definitions we report against are in how to measure sales outreach performance.

How do you run a cold email A/B test that holds up?

Randomise contacts across variants within the same list, send in the same week from the same domains, fix the sample size before you start, and read the result once, after every sequence has finished. Anything else lets the calendar, the list, or your own optimism pick the winner.

  • Parallel, not sequential. A this month and B next month is confounded by list, season, and domain health. Both variants run in the same window.
  • Same domains, same mailboxes. Alternate variants evenly across mailboxes so a tired domain does not sink one side.
  • Sample size first. Take it from the table. If you cannot reach it, write down before you start that the result will be directional.
  • No peeking. Reading early and stopping when one side pulls ahead is stop-early bias, and it manufactures winners out of noise. Replies arrive across a 2 to 3 week sequence; day 3 tells you about email one only.
  • Log it. Variant, contacts, replies, positive replies, dates. A test you cannot reproduce is an anecdote.

When should you not A/B test cold email at all?

When you have fewer than about 1,000 contacts per variant and hope to see a small lift, or when the campaign is already broken. Below 1,000 per variant you can only detect large lifts; below 300 you cannot call anything a test. And a raw reply rate under 2% is a targeting or deliverability problem that no copy test will find.

Order of operations: fix deliverability and the list, get to 8% to 11% raw, then test. Below the sample sizes, run learning rather than testing. Change one big thing per campaign, log the direction, and let a pattern across three or four campaigns count as evidence. Volume compounds: 1,000 contacts a month is 12,000 a year, enough to prove by December what no single campaign could prove in March, provided the log exists.

How MarginSales approaches cold email testing

MarginSales provides sales outreach services for companies that want to extend their outbound capacity without building the entire sales development function internally. Every programme opens with big-swing tests, segment, angle, and offer, at 500 or more contacts per variant, and moves to wording tests only once the volume supports them. Raw replies are the leading read, positive replies decide, and meetings booked are the check. Results go into a per-client log that carries across campaigns, which is how a 0.8% reply rate becomes 9% over 90 days in the programmes where it happens, and why we do not promise that to every client.

AI writes variants quickly, and drafting is the simple part, so we use it there. Choosing which lift is worth chasing, reading whether a reply is positive, and writing the next angle from what the market said are human work. Automate the simple, keep humans on the meaningful.

Frequently asked questions

How many emails do you need to A/B test cold email?

It depends on the lift you want to prove. At two-sided 95% confidence and 80% power, showing a reply rate moved from 3% to 6% takes about 749 contacts per variant, 3% to 4% takes 5,301, and 3% to 3.5% takes 19,743. Below about 1,000 contacts per variant you can only detect large lifts, and below 300 you cannot call it a test.

Should you A/B test subject lines in cold email?

Not first. Subject-line wording usually moves reply rates by fractions of a point, which takes thousands of contacts per variant to prove. Test the segment, the angle, and the offer first; those can double a reply rate and show up at a few hundred contacts per variant. Subject lines are a test for when you send 5,000 or more per variant.

What is stop-early bias in A/B testing?

Reading results before the planned sample is reached and stopping when one variant looks ahead. Early swings are mostly noise, so the practice manufactures false winners. Fix the sample size in advance, wait for every sequence to finish its 2 to 3 weeks, then read the result once.

Should you measure raw replies or positive replies?

Positive replies where volume allows, because they are what turn into meetings. Their base rate of 3% to 6% is lower than the 8% to 11% raw reply rate of disciplined outreach, so they need more contacts to move with confidence. At small volume, use raw replies with auto-replies removed as the leading read and confirm on positive replies as volume grows.

Get your last test re-read

Send us the numbers from your last A/B test: contacts per variant, replies, positive replies, and dates. We will tell you whether the winner was real, what lift your volume can actually detect, and which big swing to run next. Book the test review. Twenty minutes, and you keep the sample-size sheet either way.