Enter visitors and conversions for versions A and B to see whether the difference is statistically significant, with the uplift and p-value.
Results are estimates for planning. Check the figures before you make a financial or health decision.
How the A/B test calculator works
It runs a two-proportion z-test. It compares the conversion rate of A with the conversion rate of B, using a pooled standard error:
z = (rate B − rate A) ÷ √(p(1 − p) × (1/visitors A + 1/visitors B))
where p is the combined conversion rate. The z-score is turned into a two-sided p-value. If the p-value is below 1 minus your confidence level (for example 0.05 at 95 per cent), the result is called statistically significant.
Worked example
Using the default values in the calculator above:
| Input | Value |
|---|---|
| Visitors in A (control) | 5,000 |
| Conversions in A | 150 |
| Visitors in B (variant) | 5,000 |
| Conversions in B | 190 |
| Confidence level | 95% |
| Result | Value |
|---|---|
| Result at 95% confidence | B beats A |
| Conversion rate A | 3% |
| Conversion rate B | 3.8% |
| Absolute difference | +0.8 percentage points |
| Relative uplift of B | +26.7% |
| z-score | 2.21 |
| p-value (two-sided) | 0.0273 |
| Confidence interval for the difference | 0.09 to 1.51 points |
| Result at 95% confidence | Statistically significant |
How to read the result
A significant result means the difference is unlikely to be down to chance alone. A result that is not significant does not prove the versions are equal; it means you do not have enough evidence yet. The confidence interval shows the range of plausible differences, so a wide interval that includes zero says the test needs more data.
Tips and common mistakes
Decide your sample size before you start and run the test for the full time. Stopping the moment you see a good result raises the chance of a false win.
Run tests for at least one full business cycle, usually a week or more, so weekday and weekend behaviour is covered.
Test one change at a time and do not test lots of metrics and cherry-pick the one that wins. Each extra comparison raises the risk of a false positive.
Frequently asked questions
What does statistically significant mean?
It means the observed difference is unlikely to have happened by chance, given your chosen confidence level. It does not tell you the difference is large or important.
What confidence level should I use?
95 per cent is the most common. Use 99 per cent when a wrong decision would be costly, and 90 per cent for lower-risk tests.
What is a p-value?
The probability of seeing a difference at least this large if the two versions were really identical.
How many visitors do I need?
It depends on your current rate and the smallest improvement you care about. Small improvements need many thousands of visitors per version to detect.