← Journal
PlaybookSeptember 15, 2026· Dimitar Petkov· 9 min read

Sample Size Calculator for LinkedIn Outreach Tests (Free Tool)

A free sample size calculator for LinkedIn outreach experiments, plus the methodology behind statistically valid A/B testing at every reply rate and effect size.

Research this article with AI

Follow Well Met on Google

Sample Size Calculator for LinkedIn Outreach Tests (Free Tool)

LinkedIn outreach testing fails when teams declare winners too early. A test running two days with 80 prospects per variant is not complete. It is confirmation bias dressed up as data.

Sample size determines whether your test result is signal or noise. Too few prospects and a 15% reply rate versus 18% tells you nothing. Enough prospects and that same gap becomes a decision you can act on.

This guide explains how to calculate the minimum sample size for valid LinkedIn outreach experiments, what factors drive the number up or down, and how to balance statistical rigor with the practical constraint that every prospect sent a losing message is a missed opportunity.

How many prospects do you need for a statistically significant outreach test?

The minimum sample size depends on four inputs: your baseline conversion rate, the size of the improvement you want to detect, your desired confidence level, and your target statistical power.

For reply rate tests where the baseline sits between 8% and 25%, plan for 200 prospects per variant to detect a 5 percentage point lift with 95% confidence and 80% power. This is the standard configuration for weekly iteration cycles.

Booking rate tests require larger samples because the event is rarer. When baseline booking rates run 2% to 8%, you need 400 to 500 prospects per variant to achieve the same statistical validity. Lower-frequency outcomes demand more data to separate real effects from random variation.

What is statistical power and why does it matter?

Statistical power is the probability that your test will correctly identify a real effect when one exists. Industry standard is 80% power, meaning if the variant truly performs better, you have an 80% chance of detecting it.

Power depends on sample size. Increase the number of prospects per variant and power rises. Decrease it and power falls, leaving you vulnerable to false negatives where a winning message is discarded because the test lacked the sensitivity to prove it.

The companion concept is the significance level, typically set at 0.05. This controls the risk of a false positive, declaring a winner when no real difference exists. Together, power and significance level define the twin risks of experimentation: missing a real winner or promoting a false one.

What factors increase the required sample size?

Smaller effect sizes demand larger samples. Detecting a 2 percentage point improvement in reply rate requires roughly four times the sample size needed to detect an 8 percentage point improvement. The closer the variants perform, the more data you need to confidently separate them.

Lower baseline rates require more prospects. A test starting from a 5% booking rate needs a larger sample than one starting from a 20% reply rate to achieve equivalent power. Rare events produce fewer data points per prospect contacted.

Higher confidence or power thresholds increase sample size. Moving from 80% power to 90% power, or from 95% confidence to 99% confidence, requires more prospects. Each increment of certainty costs more data.

Sample size per variant required to detect a 5 percentage point lift at 95% confidence, 80% power (prospects)0100200300400Reply rate …Reply rate …Positive re…Booking rat…Source: Aurium Research
Source: Aurium Research

How do you calculate sample size for an A/B test?

Sample size calculation for binary outcomes (replied or did not reply, booked or did not book) follows a standard formula that accounts for baseline rate, effect size, significance level, and power. The formula derives from the normal approximation to the binomial distribution.

Online calculators handle the math. You input the baseline conversion rate, the minimum detectable effect you care about, your desired confidence level (typically 95%), and your target power (typically 80%). The calculator returns the number of prospects required per variant.

Manual calculation requires specifying the z-scores for your chosen significance level and power, then solving for n. For a two-tailed test at 95% confidence and 80% power, the z-scores are 1.96 and 0.84 respectively. The formula incorporates the variance of the binomial distribution at both the baseline and variant rates.

Chart showing how sample size requirements increase with lower baseline rates and smaller effect sizes

Can you run valid tests with fewer than 200 prospects per variant?

Yes, if the effect size is large. Testing a completely new value proposition against your current message might produce a 15 to 20 percentage point difference in reply rate. In that scenario, 100 prospects per variant can provide a clear directional signal.

The trade-off is precision. Smaller samples widen your confidence intervals. A test with 100 per variant might tell you Variant A is better, but the confidence interval around the true lift will be wide enough that you cannot pin down whether the improvement is 10% or 25%.

Reserve smaller samples for high-impact variables where you expect large swings. Use the full 200+ sample for close-call tests where precision matters, such as optimizing an already-strong message or choosing between two well-performing CTAs.

What mistakes do teams make when sizing outreach tests?

Declaring winners before reaching statistical significance is the most common error. A test that has been live for three days with 90 prospects per variant has not generated enough data to support a conclusion, regardless of how the numbers look.

Testing too many variables at once creates confounded results. If you change the opening line, the CTA, and the message length simultaneously, you cannot isolate which element drove the change in performance. Test one variable at a time unless you have the volume for full multivariate designs.

Ignoring the cost of burning prospects on losing variants leads to list exhaustion. Every message sent to a suboptimal variant is a prospect who received a worse pitch. Track burn rate explicitly and shift to sequential testing when runway tightens.

How do you balance statistical rigor with practical constraints?

Real outreach programs operate under constraints that academic experiments do not face. Your total addressable market may be 5,000 prospects, not 50,000. Burning 400 prospects on a single test consumes 8% of your list.

Audience isolation prevents prospect waste. Assign each prospect to a single test cohort at the point of list building. A prospect who receives Variant A of your opening line test cannot simultaneously appear in your CTA test. Strict segmentation protects test validity and prevents messaging overlap.

Weekly sprint cycles compress learning timelines. A testing cadence that reviews results every Monday, formulates a new hypothesis on Tuesday, and launches the next test by Wednesday ensures continuous improvement without waiting for perfect certainty. Small, frequent tests beat large, infrequent ones when list size is constrained.

What tools exist for sample size calculation?

Statsig provides a power analysis tool that estimates the relationship between minimum detectable effect, experiment duration, and allocation. You select populations, metrics, and analysis types to tailor the calculation to your specific design.

Binary sample size calculators, such as the one documented by John D. Cook and referenced by Statsig, handle experiments with binary outcomes. These tools accept inputs for significance level, power, and participant allocation ratios to estimate required sample size.

Dedicated outreach platforms with built-in A/B testing capabilities, including Aurium, automate variant management, randomized assignment, and statistical significance calculations. These systems prevent common mistakes like early winner declarations and overlapping test cohorts.

Sample size calculation tools for outreach experiments
ToolBest forKey feature
Statsig Power AnalysisProduct and platform experimentsWarehouse-native integration, metric selection
Binary sample size calculatorsSimple two-variant reply rate testsFast estimates for binary outcomes
Aurium experimentation platformLinkedIn outreach at scaleAutomated significance testing, variant management

When should you extend a test instead of declaring a winner?

Extend the test when results trend in one direction but have not yet crossed the significance threshold. If Variant A shows 17% reply rate and Variant B shows 14%, but your calculator says you need 200 per variant and you have only sent 150, let the test run.

Patience is a competitive advantage in experimentation. Declaring a winner early and rolling it out feels productive, but if the result was noise rather than signal, you have just committed your entire list to a message that performs no better (or worse) than the control.

Stop early only when the result is directionally clear and the practical difference is large enough to act on. A 25 percentage point gap in reply rate after 100 prospects per variant is almost certainly real, even if the test has not technically reached the pre-specified sample size.

At least 200 prospects per variant are required for LinkedIn reply rate tests to detect a 5 percentage point difference with statistical validity

Aurium Research (accessed), 2026-09-15

Booking rate tests require 400 to 500 prospects per variant because baseline rates typically range from 2 to 8 percent

Aurium Research (accessed), 2026-09-15

Sample size calculations for binary outcomes depend on baseline rate, effect size, significance level, and statistical power

Statsig, 2024-11-13

Statistical power of 80 percent and significance level of 0.05 are standard parameters for A/B test design

Apple Machine Learning Research, 2023-09

Frequently asked questions

  • What is the minimum sample size for a LinkedIn reply rate A/B test?

    At least 200 prospects per variant to detect a 5 percentage point difference with 95% confidence and 80% power, assuming baseline reply rates between 8% and 25%. Smaller effect sizes or lower baseline rates require larger samples.

  • Why do booking rate tests need larger samples than reply rate tests?

    Booking rates (typically 2 to 8%) are lower-frequency events than reply rates (typically 8 to 25%). Rarer outcomes produce fewer data points per prospect contacted, requiring 400 to 500 prospects per variant to achieve the same statistical power.

  • Can I declare a winner after three days if one variant is clearly ahead?

    Not unless the test has reached the pre-calculated sample size and achieved statistical significance. A result that looks clear after three days with insufficient sample size is often noise. Extend the test until the required number of prospects per variant is reached.

  • What is statistical power and why does 80% matter?

    Statistical power is the probability that your test will correctly detect a real effect when one exists. Eighty percent power is the industry standard, meaning if the variant truly performs better, you have an 80% chance of identifying it. Lower power increases the risk of false negatives.

  • How do I avoid wasting prospects on losing test variants?

    Use strict audience isolation to ensure each prospect appears in only one test at a time. Track burn rate (prospects consumed per test per week) and shift to sequential testing when your total addressable list shrinks. Reserve full sample sizes for close-call tests and use smaller samples for high-impact variables with expected large effects.

Want warm pipeline without the hours?

Book a call