Skip to main content

Sample size calculator

How many observations you need, decided before you collect any. For two conversion rates or two means — with a table showing what a smaller, more affordable test would actually buy you.

What are you comparing?

Hypothesis direction

Enter what you want to detect

Fixing the sample size before you start is what makes the p-value at the end mean what it claims to mean.

Why this comes first

Every other calculator on this site is for after the data exist. This one is for before, and the order matters more than it looks.

A p-value below 0.05 promises a 5% false-positive rate for a single, pre-planned test. Every part of that phrase is load-bearing. Decide your sample size after watching the numbers move and the promise evaporates: you are no longer running one test, you are running one per look and reporting the flattering one. Fixing n in advance is not bureaucratic diligence — it is the thing that makes the final p-value mean what it says.

The four numbers

InputWhat it isUsual value
BaselineWhere you are nowFrom your own data
Minimum detectable effectThe smallest change worth acting onA business decision, not a statistical one
αFalse-positive rate you accept0.05
PowerChance of catching a real effect0.80, or 0.90 when missing one is costly

The second row is the one people get wrong, usually by entering the lift they are hoping for rather than the smallest one that would change what they do. Hope inflates the effect, which shrinks the sample, which produces a test too small to detect what actually happens. Ask instead: below what improvement would we not bother shipping this? That is the number.

Small effects are quadratically expensive

Required sample size scales with 1/effect², which has consequences people rarely predict:

Baseline → targetRelative liftPer variant
5% → 7.5%50%~1,500
5% → 6%20%~8,200
5% → 5.5%10%~31,200
5% → 5.25%5%~122,000

Each halving of the effect roughly quadruples the traffic. If your site sees 2,000 visitors a week per variant, the second row is a month and the fourth is well over a year — by which point seasonality and site changes have contaminated the comparison anyway. Knowing that in advance is far better than discovering it in week six.

What an underpowered test does to you

It does not simply fail to find things. A test with 30% power that does come back significant has, almost by necessity, overestimated the effect — only an unusually large sample fluctuation could have cleared the threshold. So the finding both replicates poorly and looks more impressive than the truth.

That is the mechanism behind a good deal of the replication crisis, and it applies just as much to a marketing team shipping a change that measured a 30% lift and delivers 4%. The power table in the results is there so you can see what you are buying before you commit, rather than inferring it afterwards.

The sample size formula

Every sample size formula on this page is the same idea rearranged: the number of observations per group needed for an effect of a given size to clear the significance threshold a given fraction of the time. Two quantiles of the normal distribution do the work — zα, set by your significance level, and zβ, set by your power. At α = 0.05 two-tailed and 80% power those are 1.960 and 0.842.

For two proportions, the standard normal-approximation formula, with the pooled variance in the α term and the unpooled variance in the power term:

n = (zα√(2p̄(1−p̄)) + zβ√(p₁(1−p₁) + p₂(1−p₂)))² / (p₂ − p₁)²

For two means, the equivalent expressed through Cohen's d — the difference in means divided by the pooled standard deviation:

n = 2(zα + zβ)² / d²

The second form is the one to look at if you want to understand the behaviour rather than just apply it, because everything expensive about testing is visible in it. The effect sits in the denominator and it is squared, which is the whole of the quadratic cost above: halve d and n quadruples. Power and α enter only through the two z terms, so tightening them is comparatively cheap — moving from 80% to 90% power raises zβ from 0.842 to 1.282 and the required n by about a third, not by a factor of four. Each formula gives the size of one group; double it for the total across both.

Both are approximations that assume equal group sizes and a normal sampling distribution. They are what most textbooks and most A/B tools use, and they get optimistic when the required n is small or the rates are extreme — the calculator says so when the expected number of conversions per group drops below about 10. Once the test has run, the A/B test calculator gives you the p-value and the interval on the lift you actually observed.

Frequently asked questions

How many visitors do I need for an A/B test?

It depends on your baseline rate and the smallest lift you care about. Detecting a move from 5% to 6% — a 20% relative lift — needs about 8,200 visitors per variant at 80% power and α = 0.05. Halving that to a 10% relative lift takes roughly 31,000 per variant. Small effects are expensive, and the cost rises as the square of the precision you want.

What is statistical power?

The probability of detecting an effect that is genuinely there. Power of 80% means that if your hypothesised effect is real, you have an 80% chance of ending up with a significant result — and a 20% chance of missing it. 80% is convention rather than law; 90% is common in clinical work, where missing a real effect is costly.

What is the sample size formula?

For comparing two means it is n = 2(zα + zβ)² / d², where d is Cohen’s d and the two z terms are normal quantiles set by your significance level and your power — 1.960 and 0.842 at α = 0.05 two-tailed with 80% power. For two proportions the same structure applies with the pooled variance in the α term: n = (zα√(2p̄(1−p̄)) + zβ√(p₁(1−p₁) + p₂(1−p₂)))² / (p₂ − p₁)². Both give the size of one group, so double the result for the total.

Why does detecting a smaller effect cost so much more?

Because required sample size scales with the inverse square of the effect. Halving the effect you want to detect quadruples the observations you need; detecting a third as much takes nine times as many. This is the single most surprising thing about test planning, and the reason "let us just run it and see" tends to produce underpowered tests that find nothing.

Can I stop my test early if it turns significant?

No — this is the most expensive mistake in testing. A fixed-horizon p-value promises a 5% false-positive rate for one pre-planned test. Checking daily for two weeks and stopping at the first significant reading means running fourteen tests and reporting the one that worked, which pushes the real false-positive rate past 30%. If you must monitor continuously, you need a sequential design built for it.

What if I cannot get that many observations?

Then run a smaller test knowing what it buys you. The table in the results shows the power you get at various fractions of the ideal sample. A test with 40% power is not worthless, but it is worth knowing in advance that it will miss the effect more often than it finds it — and that any significant result it does produce is likely to overstate the effect size.

Should I use a one-tailed test to reduce the sample size?

Only if a result in the opposite direction would genuinely be treated the same as no difference — and in most A/B testing it would not, because shipping a change that quietly hurts conversion is precisely the outcome you are trying to avoid. One-tailed testing does reduce the required sample, but it buys that by giving up the ability to detect harm.