Significance level (α) explained
Alpha is the false-positive rate you are willing to accept. Choosing it well means thinking about what a wrong answer costs — which is why 0.05 is not always the right number.
6 min read · Last reviewed 7 August 2026
What alpha controls
The significance level α is the probability of rejecting a true null hypothesis — a false positive, or Type I error. Setting α = 0.05 means accepting that, when there is genuinely no effect, you will claim one 5% of the time.
Crucially, α is a decision you make before seeing data. The p-value is what your data produce; α is the standard you hold them to. Choosing α after looking at your p-value removes any meaning it had.
The two errors trade off
| No real effect | Real effect exists | |
|---|---|---|
| You claim an effect | Type I error (α) | Correct |
| You claim no effect | Correct | Type II error (β) |
Lowering α reduces false positives and, holding everything else constant, increases false negatives. You cannot minimise both at once with a fixed sample; the only way to reduce both is to collect more data.
This is what makes α a genuine decision rather than a convention to look up. Which error is worse in your situation?
Choosing a threshold
Use a stricter α (0.01 or lower) when a false positive is expensive:
- Clinical decisions where a wrong conclusion causes harm
- Irreversible or costly commitments
- Any setting where you are running many tests
- Claims that will be widely acted on before replication
Use a looser α (0.10) when a false negative is the bigger problem:
- Exploratory screening intended to be followed up
- Pilot studies deciding whether to invest in a full trial
- Safety monitoring, where missing a signal is the real risk
How different fields set it
| Field | Typical α | Reason |
|---|---|---|
| Particle physics | ~3 × 10⁻⁷ | Enormous numbers of comparisons; discoveries are permanent claims |
| Genomics | ~5 × 10⁻⁸ | Millions of simultaneous tests |
| Clinical trials | 0.05, often 0.025 one-sided | Regulated, with patient safety at stake |
| Psychology, social science | 0.05 | Convention |
| A/B testing | 0.05, sometimes 0.10 | Changes are cheap and reversible |
| Pilot studies | 0.10 | Screening, not concluding |
Alpha and multiple tests
α controls the error rate for one test. Run several and your overall false-positive rate compounds: for k independent tests it is 1 − (1 − α)ᵏ. At α = 0.05 that is 23% for five tests, 40% for ten, and 64% for twenty.
The standard fixes are a Bonferroni correction (divide α by the number of tests — simple and conservative) or a false discovery rate procedure such as Benjamini–Hochberg, which controls the expected proportion of false positives among your discoveries and is much less brutal when you have many tests.
Remember that multiplicity is broader than it looks. Testing several outcomes, several subgroups, or several model specifications all count. So does checking an experiment repeatedly and stopping when it turns significant.
Alpha is not the whole story
A threshold turns a continuous measure of evidence into a binary decision, and something is always lost in that conversion. p = 0.049 and p = 0.051 are the same evidence; only one clears a line at 0.05. If your result has landed on one side of that line and you want to know what it licenses you to claim, start with what "p < 0.05" actually means.
Set your α in advance and use it to make the decision you need to make — but report the exact p-value alongside the effect size and confidence interval, so your readers can weigh the evidence themselves rather than inheriting your threshold.