What "p < 0.05" actually means
This is the number people act on and the number people misread. Here is exactly what a p-value below 0.05 licenses you to say — and the four things it is constantly taken to mean but does not.
8 min read · Last reviewed 7 August 2026
The short answer
A p-value below 0.05 means: if there were genuinely no effect, data at least as extreme as yours would show up less than 5% of the time. Your data are awkward to explain as coincidence, so by convention you reject the null hypothesis and call the result statistically significant.
That is the whole claim. It is a statement about how surprising your data would be in a world with no effect — not about how likely that world is, and not about how big your effect turned out to be.
The four things it does not mean
Each of these is a different sentence from the one above, and each is wrong.
| The misreading | Why it is wrong |
|---|---|
| "There is a 95% chance my hypothesis is true" | The p-value is computed assuming the null is true. It cannot then tell you the probability that it is. That would need a prior. |
| "There is only a 5% chance this is a fluke" | Same error, reworded. 5% is how often the data would look like this given no effect, not how often a significant finding is wrong. |
| "The effect is large" | Significance scales with sample size. A 0.04% conversion lift clears p < 0.05 with enough traffic and is worth nothing. |
| "The result will replicate" | A single p just under 0.05 is weak evidence. Studies powered at 80% that get p = 0.04 replicate far less often than people expect. |
The second row deserves a number, because it is the misreading that costs the most. Suppose one hypothesis in ten that people bother testing is actually true, and studies have 80% power. Test 1,000 of them: 100 are real and 80 come out significant; 900 are dead ends and 45 come out significant anyway (5% of 900). So 125 findings clear p < 0.05, and 45 of them — 36% — are false. The 5% was never the chance your significant result is wrong.
What to do when your p-value is below 0.05
Five things, in order, before you write it up or ship it:
- Look at the effect size. Cohen's d, a difference in means, a lift, an odds ratio — whatever the natural measure is. This is the question you actually care about, and the p-value has not answered it.
- Look at the confidence interval. If it runs from "trivial" to "enormous," you have detected something without measuring it. That is a different finding from a tight interval around a useful value, even at an identical p.
- Check you only ran one test. If you tried several outcomes, several subgroups, or several model specifications, your real false-positive rate is not 5%. With ten tests at α = 0.05 it is about 40%. Correct for it, or say plainly that the analysis was exploratory.
- Check you did not stop early. Watching a test and stopping the moment it turns significant inflates the error rate dramatically. The sample size should have been fixed in advance.
- Report the exact value, not the threshold. "p = .032" carries information; "p < .05" throws it away and hides the difference between .049 and .001.
And if it is above 0.05?
You have failed to reject the null hypothesis. That is emphatically not the same as showing there is no effect — it means your data were not surprising enough, which can happen either because nothing is there or because you did not collect enough data to see it. Note the wording: the outcome is "fail to reject" rather than "accept," and the difference is not pedantry.
The confidence interval tells you which. An interval from −0.05 to +0.06 genuinely supports "no meaningful effect." An interval from −4 to +9 means the study could not tell, and reporting it as "no significant difference" is misleading. If you need to establish that an effect is absent, that requires an equivalence test, not a non-significant p.
Why 0.05, and does it have to be?
There is no mathematical reason. Ronald Fisher proposed it in the 1920s as a convenient rule of thumb — one chance in twenty, a round number that suited the printed tables of the day — and expected researchers to pick a threshold suited to their situation. It calcified into a universal standard anyway.
Other fields chose differently once the cost of a false positive changed. Particle physics requires roughly 3 × 10⁻⁷ before it will call something a discovery. Genomics runs at about 5 × 10⁻⁸ because it tests millions of hypotheses at once. Neither is being fussy; both are answering the same question Fisher intended you to answer, and getting a different number.
So no, it does not have to be 0.05 — but it does have to be chosen before you see your data. Picking the threshold after the p-value arrives is how a threshold stops meaning anything. Our guide to choosing alpha works through how to set it deliberately.
p = 0.049 and p = 0.051
These are the same evidence. Nothing in nature changes between them; only a line drawn by convention falls in between. Treating one as a discovery and the other as a null result is the single strangest habit the 0.05 threshold has produced.
A threshold is useful when you must make a binary decision — ship or do not ship, approve or do not approve. It is not useful as a description of evidence, and a p-value is a continuous measure of evidence. Both things can be true: use α to decide, and report the exact p, the effect size and the interval so your reader can weigh it for themselves.
One small technicality, since people ask: the convention is strictly less than. p = 0.05 exactly does not clear α = 0.05. In practice a result that lands that precisely on the line should be reported as what it is — borderline — rather than pushed to one side.
How to report it
Give the test statistic, the degrees of freedom, the exact p-value, and the effect size —
with the confidence interval on the difference itself alongside them. In APA style the first
part looks like t(28) = 2.14, p = .041, d = 0.78. Note the house rules on the
p-value: no leading zero, three decimal places, and p < .001 rather than a
smaller number.
Every calculator on this site emits that string for you, and shades the tail area so you can see the p-value rather than just read it. Start from the p-value calculator if you have a test statistic, or the t-test calculator if you have raw data.