Calculators·Free tool

Free A/B Test Sample Size Calculator

This sample size calculator estimates how many visitors per variant you need for a statistically valid A/B test, based on your baseline conversion rate and minimum detectable effect. Built for marketers planning landing page or ad tests.

All tools

Required Sample Size (per variant) appears here — fill in the fields above.

About the A/B Test Sample Size Calculator

Calculate the sample size you need for a statistically valid A/B test from baseline conversion rate and minimum detectable effect. Try Creetr for ad tests.

Every calculation and generation runs entirely in your browser - free, instant, nothing uploaded. The hard part is what comes next: filming it. Once your a/b test sample size is ready, Creetr shoots it with an AI actor and exports a ready-to-post ad in about two minutes - no camera, no creator fees.

Sample size calculator tools estimate how many visitors per variant you need to run a statistically valid A/B test, based on your baseline conversion rate and the minimum detectable effect (the smallest lift you need to be able to reliably detect). Enter your numbers above for an instant sample size estimate at standard 95% confidence and 80% power.

What is a sample size calculator?

A sample size calculator determines the minimum number of visitors, per variant, required for an A/B test to reach a statistically significant, trustworthy conclusion. It's built on your baseline conversion rate (the control's current performance) and minimum detectable effect, or MDE (the smallest relative change in conversion rate you actually care about detecting — expressed as a percentage of the baseline, not a raw percentage-point difference).

The underlying statistics use a standard two-proportion sample size formula at 95% confidence and 80% statistical power — industry-standard defaults that balance test reliability against practical runtime. Smaller MDEs (trying to detect a subtle 5% lift) require dramatically larger sample sizes than larger MDEs (a bold 30%+ lift is easier to detect with fewer visitors).

How to use this sample size calculator

  1. Pull your baseline conversion rate — your control's current performance over a recent, representative period (use the conversion rate calculator if you need to compute it first).
  2. Decide your minimum detectable effect — the smallest relative lift that would actually be worth acting on. 20% relative lift is a common default for creative and landing page tests; smaller MDEs like 10% require much larger sample sizes and longer test runtimes.
  3. Enter both numbers and read your required sample size per variant.
  4. Multiply by the number of variants in your test (including control) to get total required traffic.
  5. Divide total required traffic by your average daily/weekly traffic to that page or ad set to estimate test runtime — don't call a test early just because one variant looks ahead after a few days.
  6. If required sample size exceeds what you can realistically get in a reasonable timeframe, either increase your MDE (accept you can only detect bigger lifts) or extend the test duration — don't shrink your sample size arbitrarily, since that inflates the risk of a false positive.

Why sample size matters

Calling A/B tests too early is one of the most common and costly mistakes in performance marketing. A variant that looks like a 15% winner after 200 visitors per side is very often noise, not signal — without adequate sample size, you have no statistical basis to trust the result, and shipping a "winning" creative or landing page change based on an underpowered test is essentially a coin flip dressed up as data.

This matters just as much for ad creative testing as for landing pages: if you're running multiple UGC video ad variants against each other in TikTok or Meta Ads Manager, the same sample size logic applies to conversion events, not just impressions or clicks — you need enough actual conversions per variant, not just enough traffic, before trusting which creative concept is genuinely outperforming.

Benchmarks

Approximate required sample size per variant at 95% confidence, 80% power, for common baseline conversion rates and MDEs:

Baseline CVRMDE (relative)Approx. Sample Size / Variant
2%20%~19,000
2%30%~8,500
5%20%~7,300
5%30%~3,300
10%20%~3,400
10%30%~1,500

As a rule of thumb, halving your MDE roughly quadruples required sample size — which is why teams testing bold creative swings (30%+ expected lift) can validate results far faster than teams chasing marginal 5-10% tweaks.

Common sample size mistakes

The most damaging mistake is peeking at results daily and stopping the test the moment a variant looks ahead — this dramatically inflates your false positive rate, since random noise naturally produces temporary leads throughout a test's run. A second mistake is testing too many variants at once against a fixed traffic budget, which stretches your per-variant sample size timeline out for weeks or months without anyone noticing until the test has already dragged on far longer than planned. Third, teams often ignore weekday-versus-weekend behavior differences and call a test after just 3-4 days — always plan for at least one full week, ideally two, to average out day-of-week variation in both traffic volume and conversion behavior.

Test more concepts, not just more variants of one idea

Most teams don't have a sample-size problem so much as an idea-generation problem — they run small, incremental tests (button color, headline tweak) that require huge sample sizes to detect, instead of testing bigger creative swings that produce bigger, faster-to-detect lifts.

Creetr generates multiple distinct UGC-style video ad concepts — different hooks, actors, and angles — from a single product link, so you can test bold creative differences instead of marginal tweaks, and reach statistical significance faster. Try Creetr free. For the metric you're testing toward, see the conversion rate calculator, and read the performance marketing guide for a full testing framework.

Turn this into a real UGC video ad

Paste a product link, pick an AI actor, and Creetr generates a ready-to-post ad.

Try Creetr free

Frequently asked questions

What sample size do I need for an A/B test?+

Your A/B test sample size depends on your current conversion rate and the smallest improvement you want to be able to detect. Lower baseline conversion rates and the desire to find smaller improvements both necessitate more data. Instead of guessing with a generic number, use a dedicated calculator, inputting your specific conversion rate and the minimum uplift you're aiming for. This ensures your test has enough power to reliably identify meaningful changes, preventing you from missing opportunities or drawing false conclusions.

What is minimum detectable effect (MDE)?+

The Minimum Detectable Effect (MDE) is the smallest change in your conversion rate that your test is designed to reliably identify. Think of it as the smallest improvement or decline you need to be confident your changes are making a real difference, not just random fluctuations. For example, if your current conversion rate (baseline) is 3%, and you set an MDE of 20%, your test will be able to detect if your conversion rate increases to around 3.6% or decreases to around 2.4%. Choosing the right MDE is crucial for balancing the sensitivity of your test with the time and resources required to run it on Creetr.

Why did my A/B test show a winner that didn't hold up later?+

Your A/B test winner may not have held up because it was concluded prematurely or the observed difference was due to random chance. To ensure reliable results on Creetr, always let your A/B test run until it achieves the statistically determined required sample size. Even if one variant appears to be performing significantly better early on, continuing the test to its full duration helps differentiate genuine performance gains from temporary fluctuations. This practice ensures you're making decisions based on robust data, not just an initial, potentially misleading trend.

What confidence level should I use for A/B tests?+

Aim for 95% confidence in your A/B tests on Creetr to ensure your results are reliable without extending test durations unnecessarily. This level means that if you were to run the same test many times, 95% of the time the winning variation would be the true winner, and 5% of the time you might see a winner by chance. While 99% confidence offers greater certainty, it demands significantly more data, potentially making your tests impractical for short-form video ad campaigns. Conversely, a 90% confidence level might allow for quicker results, but increases the chance of mistakenly believing a variation is better when it's not, leading to wasted ad spend. For most Creetr users, 95% strikes the right balance for making informed decisions on your ad creatives.

How long should an A/B test run?+

Your A/B test should run until you reach your calculated required sample size for each variant, and ideally for at least one to two full weeks. This duration is crucial for capturing variations in traffic and user behavior that occur across different days of the week. Running your test for this minimum period helps ensure your results are reliable and not skewed by short-term trends. Stopping prematurely, especially mid-week, can lead to inaccurate conclusions about which ad variant performs best on Creetr.

Does sample size requirement change for multi-variant tests (A/B/C/D)?+

Yes, while the sample size needed *per variant* remains consistent, your *total* traffic requirement increases proportionally with each additional variant in your multi-variant test. For example, an A/B/C/D test needs four times the total traffic of an A/B test to achieve the same statistical power for each individual comparison. Crucially, when you're comparing more than two variants, you must also adjust your significance threshold to account for the increased chance of finding a false positive due to making multiple comparisons. This is often done using methods like the Bonferroni correction, ensuring your results remain reliable.