Free A/B Test Sample Size Calculator
This sample size calculator estimates how many visitors per variant you need for a statistically valid A/B test, based on your baseline conversion rate and minimum detectable effect. Built for marketers planning landing page or ad tests.
Required Sample Size (per variant) appears here — fill in the fields above.
About the A/B Test Sample Size Calculator
Calculate the sample size you need for a statistically valid A/B test from baseline conversion rate and minimum detectable effect. Try Creetr for ad tests.
Every calculation and generation runs entirely in your browser - free, instant, nothing uploaded. Numbers tell you whether an ad is worth making; they don't make it. When the maths says go, Creetr turns a product link into a ready-to-post UGC ad in about two minutes. Start with Creetr - no camera, no creator fees.
Sample size calculator tools estimate how many visitors per variant you need to run a statistically valid A/B test, based on your baseline conversion rate and the minimum detectable effect (the smallest lift you need to be able to reliably detect). Enter your numbers above for an instant sample size estimate at standard 95% confidence and 80% power.
What is a sample size calculator?
A sample size calculator determines the minimum number of visitors, per variant, required for an A/B test to reach a statistically significant, trustworthy conclusion. It's built on your baseline conversion rate (the control's current performance) and minimum detectable effect, or MDE (the smallest relative change in conversion rate you actually care about detecting — expressed as a percentage of the baseline, not a raw percentage-point difference).
The underlying statistics use a standard two-proportion sample size formula at 95% confidence and 80% statistical power — industry-standard defaults that balance test reliability against practical runtime. Smaller MDEs (trying to detect a subtle 5% lift) require dramatically larger sample sizes than larger MDEs (a bold 30%+ lift is easier to detect with fewer visitors).
How to use this sample size calculator
- Pull your baseline conversion rate — your control's current performance over a recent, representative period (use the conversion rate calculator if you need to compute it first).
- Decide your minimum detectable effect — the smallest relative lift that would actually be worth acting on. 20% relative lift is a common default for creative and landing page tests; smaller MDEs like 10% require much larger sample sizes and longer test runtimes.
- Enter both numbers and read your required sample size per variant.
- Multiply by the number of variants in your test (including control) to get total required traffic.
- Divide total required traffic by your average daily/weekly traffic to that page or ad set to estimate test runtime — don't call a test early just because one variant looks ahead after a few days.
- If required sample size exceeds what you can realistically get in a reasonable timeframe, either increase your MDE (accept you can only detect bigger lifts) or extend the test duration — don't shrink your sample size arbitrarily, since that inflates the risk of a false positive.
Why sample size matters
Calling A/B tests too early is one of the most common and costly mistakes in performance marketing. A variant that looks like a 15% winner after 200 visitors per side is very often noise, not signal — without adequate sample size, you have no statistical basis to trust the result, and shipping a "winning" creative or landing page change based on an underpowered test is essentially a coin flip dressed up as data.
This matters just as much for ad creative testing as for landing pages: if you're running multiple UGC video ad variants against each other in TikTok or Meta Ads Manager, the same sample size logic applies to conversion events, not just impressions or clicks — you need enough actual conversions per variant, not just enough traffic, before trusting which creative concept is genuinely outperforming.
Benchmarks
Approximate required sample size per variant at 95% confidence, 80% power, for common baseline conversion rates and MDEs:
| Baseline CVR | MDE (relative) | Approx. Sample Size / Variant |
|---|---|---|
| 2% | 20% | ~19,000 |
| 2% | 30% | ~8,500 |
| 5% | 20% | ~7,300 |
| 5% | 30% | ~3,300 |
| 10% | 20% | ~3,400 |
| 10% | 30% | ~1,500 |
As a rule of thumb, halving your MDE roughly quadruples required sample size — which is why teams testing bold creative swings (30%+ expected lift) can validate results far faster than teams chasing marginal 5-10% tweaks.
Common sample size mistakes
The most damaging mistake is peeking at results daily and stopping the test the moment a variant looks ahead — this dramatically inflates your false positive rate, since random noise naturally produces temporary leads throughout a test's run. A second mistake is testing too many variants at once against a fixed traffic budget, which stretches your per-variant sample size timeline out for weeks or months without anyone noticing until the test has already dragged on far longer than planned. Third, teams often ignore weekday-versus-weekend behavior differences and call a test after just 3-4 days — always plan for at least one full week, ideally two, to average out day-of-week variation in both traffic volume and conversion behavior.
Test more concepts, not just more variants of one idea
Most teams don't have a sample-size problem so much as an idea-generation problem — they run small, incremental tests (button color, headline tweak) that require huge sample sizes to detect, instead of testing bigger creative swings that produce bigger, faster-to-detect lifts.
Creetr generates multiple distinct UGC-style video ad concepts — different hooks, actors, and angles — from a single product link, so you can test bold creative differences instead of marginal tweaks, and reach statistical significance faster. Try Creetr free. For the metric you're testing toward, see the CTR calculator.
Turn this into a real UGC video ad
Paste a product link, pick an AI actor, and Creetr generates a ready-to-post ad.
Frequently asked questions
What sample size do I need for an A/B test?+
It depends on your baseline conversion rate and minimum detectable effect — smaller baseline rates and smaller effects both require larger sample sizes. Use this calculator with your specific numbers rather than relying on a rule-of-thumb figure like "1,000 visitors," which is often insufficient for low-conversion-rate pages.
What is minimum detectable effect (MDE)?+
MDE is the smallest relative change in conversion rate you want your test to be able to reliably detect, expressed as a percentage of your baseline. A 3% baseline with a 20% MDE means you're trying to detect a lift to roughly 3.6% or a drop to roughly 2.4%.
Why did my A/B test show a winner that didn't hold up later?+
This almost always means the test was called before reaching adequate sample size, or the "winning" variant's lead was within normal statistical noise. Always let a test run to its calculated required sample size before making a final call, even if one variant looks ahead early.
What confidence level should I use for A/B tests?+
95% confidence is the standard default across marketing and product testing, balancing reliability against practical test runtime. Higher confidence levels (99%) require larger sample sizes; lower levels (90%) run faster but carry more risk of a false positive.
How long should an A/B test run?+
At minimum, until you hit your calculated required sample size per variant — and ideally at least 1-2 full weeks regardless, to account for day-of-week variation in traffic and conversion behavior. Stopping a test mid-week based on early results is a common source of false conclusions.
Does sample size requirement change for multi-variant tests (A/B/C/D)?+
The per-variant sample size calculation stays the same, but total required traffic multiplies by the number of variants, and you should adjust your significance threshold (Bonferroni correction or similar) when comparing more than two variants to avoid false positives from multiple comparisons.