A/B test significance calculator
Enter visitors and conversions for each variant. This is a two-proportion Z-test, so it gives you the p-value from your two proportions, the confidence interval on the actual lift, and a straight answer about what the result does and does not justify.
Enter your test results to begin
You get the p-value, the effect size, and a confidence interval together — because a p-value on its own does not tell you whether the result matters.
Results
P-value
Estimate
Effect size
What this means
What it does not mean
Report it (APA)
Getting a p-value from two proportions
Underneath the A/B framing this is a two-proportion Z-test, and it works for any pair of proportions: conversion and click-through rates, defect rates, pass rates, cure rates, yes/no survey answers. If your data are counts of successes out of a total, this is the right calculator whether or not there is an experiment involved.
Enter counts, not percentages. If a rate is all you have, multiply it back out — a 4.2% conversion rate on 12,000 visitors is 504 conversions. Rounding a reported percentage costs you very little here; typing 4.2 into the conversions box instead of 504 costs you the whole answer.
The statistic is:
z = (p₁ − p₂) / √( p̄ (1 − p̄) (1/n₁ + 1/n₂) )
where p̄ is the pooled proportion — every success over every trial. The pooling is deliberate: under the null hypothesis the two rates are identical, so the best estimate of that shared rate uses both samples at once. The resulting z goes to the standard normal distribution, which means you can check any result here by hand against the Z-score to p-value calculator.
Two cases this is not for. If you are testing one proportion against a fixed target rather than two against each other, you want a one-sample proportion test, which this calculator does not run — but it is a short step by hand. With p̂ your observed rate, p₀ the target and n your sample size, the statistic is z = (p̂ − p₀) / √(p₀(1 − p₀) / n); put that z into the Z-score to p-value calculator for the p-value. And if the expected number of successes in any cell falls below about 5, the normal approximation stops being trustworthy — use Fisher’s exact test instead, and see the warning the calculator raises.
The peeking problem
Almost every A/B testing mistake is a version of the same one: watching the dashboard and stopping the moment the result turns green. It feels efficient. It is the fastest way to ship changes that do nothing.
A p-value below 0.05 promises a 5% false-positive rate for a single, pre-planned test. If you check the result every day for two weeks and stop at the first significant reading, you have effectively run fourteen tests and reported only the one that worked. The real false-positive rate climbs past 30%.
The fix is not complicated: decide your sample size before you start, run the test to completion, and look once. If you genuinely need to monitor continuously, you need a sequential testing method designed for it — not a fixed-horizon p-value checked repeatedly.
A note on how this is calculated
The test statistic uses the pooled standard error, which is correct under the null hypothesis where both rates are assumed equal. The confidence interval uses the unpooled standard error, which is correct for estimating a difference you are not assuming to be zero.
Using one for both is a widespread bug in A/B calculators, and it produces a confidence interval that disagrees with its own p-value — an interval that excludes zero next to a non-significant result, or the reverse. If you have ever seen that and assumed it was a rounding artefact, it was not.
Significance is not the same as impact
With enough traffic, almost any difference becomes statistically significant. A conversion lift from 5.00% to 5.04% will clear p < 0.05 given a few million visitors, and it is almost certainly not worth the engineering cost of shipping it.
This is why the confidence interval sits next to the p-value in the results rather than buried underneath it. The interval tells you the range of lifts your data support. If it runs from +0.1% to +6%, you have established that the variant is probably better and learned almost nothing about by how much. Decide what lift would actually justify the change before running the test, and check whether the interval clears it.
Before you start a test
- Fix the sample size in advance. Not "until it looks significant."
- Decide the minimum lift worth shipping. Significance without a threshold is a number without a decision.
- Run for whole weeks. Tuesday traffic does not behave like Sunday traffic; a partial week bakes in a bias.
- Test one change at a time, or accept that you cannot attribute the result.
- Count each visitor once. Repeated measurements of the same person break the independence assumption the test relies on.
Frequently asked questions
How do I calculate a p-value from two proportions?
Use a two-proportion Z-test, which is what this calculator runs. Enter the number of trials and the number of successes for each group — visitors and conversions, or whatever the equivalent is in your data. The statistic is z = (p₁ − p₂) / √(p̄(1 − p̄)(1/n₁ + 1/n₂)), where p̄ is the pooled proportion across both samples, and the p-value is the corresponding area of the standard normal distribution. Enter counts rather than percentages: a 4.2% rate on 12,000 visitors means 504 conversions.
Can I use this for proportions that have nothing to do with A/B testing?
Yes. The underlying test compares any two proportions — defect rates on two production lines, pass rates for two cohorts, response rates to two mailings, cure rates in two arms of a trial. The A/B labelling is just the most common use. The only requirements are that each observation counts once and falls into exactly one of two categories, and that the two groups are independent of each other.
How do I know if my A/B test is significant?
Enter the visitors and conversions for each variant. If the p-value falls below your significance level — usually 0.05 — the difference is statistically significant. But check the confidence interval before you act: if it spans a range from "barely worth it" to "enormous," you have detected an effect without measuring it well enough to plan around.
Can I stop my test as soon as it turns significant?
No, and this is the single most expensive mistake in A/B testing. Checking repeatedly and stopping at the first significant result inflates your false-positive rate from 5% to 30% or more, because you are effectively running many tests and reporting only the one that worked. Decide your sample size before you start, then run to it.
Should I use a one-tailed or two-tailed test for A/B testing?
Two-tailed. You genuinely care if your variant performs worse, not only if it performs better — shipping a change that quietly hurts conversion is precisely the outcome you are trying to avoid. One-tailed testing halves your p-value without adding any evidence, which is why it is popular with vendors and distrusted by statisticians.
What sample size do I need?
It depends on your baseline conversion rate and the smallest lift worth detecting. Detecting a relative 10% lift on a 5% baseline needs roughly 30,000 visitors per variant at 80% power. Smaller effects need dramatically more traffic — halving the effect you want to detect quadruples the sample required.
Why does the calculator warn me about small counts?
The two-proportion Z-test relies on a normal approximation that needs a reasonable number of expected events in every cell — roughly 5 or more. With very low conversion counts the approximation breaks down and the p-value becomes unreliable. In that situation Fisher's exact test gives a trustworthy answer.
My test is not significant. Does that mean the variants are identical?
No. It means you have not gathered enough evidence to distinguish them. Look at the confidence interval on the difference: if it runs from −2% to +3%, you can reasonably conclude any real effect is small. If it runs from −8% to +12%, your test was simply underpowered and the honest answer is that you still do not know.