How to interpret p-values correctly
Most people can calculate a p-value. Far fewer can say what it means without making one of five specific errors. Here is each one, and the correct phrasing to replace it.
9 min read · Last reviewed 7 August 2026
Why this is so hard
Statistics education is heavily weighted toward computation. Students learn to produce a p-value long before — sometimes instead of — learning what it represents. The result is a widespread ability to calculate a number correctly and describe it wrongly.
The errors below are not obscure. They appear routinely in published papers, press releases, and analytics dashboards.
Error 1: "There is a 5% chance the null hypothesis is true"
Why it is wrong: the p-value is computed by assuming the null hypothesis is true. A number derived under an assumption cannot simultaneously measure the probability of that assumption.
The two quantities differ dramatically. P(data | null) is what you have. P(null | data) is what you want, and getting there requires knowing how plausible your hypothesis was beforehand. For a hypothesis with a 50/50 prior, observing p = 0.05 leaves the probability the null is true at roughly 29% — nearly six times the figure people assume.
Say instead: "If there were no effect, we would see data this extreme about 5% of the time."
Error 2: "p = 0.001 means a stronger effect than p = 0.04"
Why it is wrong: p-values reflect both effect size and sample size, and you cannot separate them from the p-value alone.
A concrete case. Study A finds a 2-point improvement with n = 10,000 and reports p = 0.001. Study B finds a 15-point improvement with n = 30 and reports p = 0.04. Study B found an effect more than seven times larger. Study A had more data.
Say instead: compare effect sizes and their confidence intervals. That is what they are for.
Error 3: "Not significant means there is no effect"
Why it is wrong: absence of evidence is not evidence of absence. A non-significant result means your data were not surprising enough under the null — which could be because there is no effect, or because your study was too small to detect one.
The confidence interval resolves the ambiguity instantly. If your interval on a difference runs from −0.3 to +0.4 points, you have good evidence any real effect is small. If it runs from −12 to +18 points, you have learned essentially nothing. Both can produce p = 0.6.
Say instead: "We did not find sufficient evidence of an effect," followed by the interval so readers can judge what was ruled out.
Error 4: Treating 0.05 as a bright line
Why it is wrong: p = 0.049 and p = 0.051 represent virtually identical evidence. Declaring one a discovery and the other a null result is an artefact of the threshold, not a feature of the data.
This dichotomy drives real distortions: results just below the line get published and results just above it disappear, so the literature systematically overstates effects.
Say instead: report the exact p-value and treat it as a continuous measure of evidence. Our guide to what "p < 0.05" actually means works through the threshold in detail, including what to check before you act on a result that clears it.
Error 5: Ignoring how many tests you ran
Why it is wrong: α = 0.05 controls the false-positive rate for one pre-planned test. Run twenty independent tests and the probability of at least one false positive is about 64%.
This applies far more widely than people realise: testing several outcomes, several subgroups, several model specifications, or checking an A/B test daily until it turns significant. Each is a form of multiplicity, and the last one is why the practical false discovery rate in published research runs near 30%.
Say instead: pre-register your primary analysis, or apply a correction and report how many comparisons you made.
Worked example: a significant result
A trial compares a new treatment against a control. Mean improvement is 4.2 points higher in
the treatment arm, with t(198) = 2.31, p = .022, 95% CI [0.61, 7.79].
What you can say: the data are inconsistent with no treatment effect at the 5% level. The best estimate is a 4.2-point improvement, and the data are compatible with anything from 0.6 to 7.8.
What you cannot say: that there is a 97.8% chance the treatment works; that the effect is definitely 4.2 points; or that a clinically meaningful benefit has been established — the interval's lower end of 0.6 may be far too small to matter to a patient.
Worked example: a non-significant result
The same trial, with 40 patients instead of 200: a 4.2-point difference with
t(38) = 1.03, p = .31, 95% CI [−4.05, 12.45].
The estimated effect is identical. Only the precision changed. The interval now spans everything from a meaningful harm to a large benefit.
What you can say: this study was too small to determine whether the treatment helps. The data remain compatible with a substantial benefit.
What you cannot say: that the treatment does not work, or that the two groups are equivalent. Reporting only "p = .31, not significant" would leave readers with exactly the wrong impression — which is why intervals belong in every report.
A checklist before you report
- Have I reported the exact p-value rather than only a threshold?
- Have I reported an effect size in meaningful units?
- Have I reported a confidence interval?
- Have I said how many tests I ran?
- Did I decide the analysis before seeing the data?
- If the result is null, have I said whether the study could have detected an effect?
Our t-test calculator reports the effect size and interval alongside the p-value by default, so the first three come for free.