Skip to main content

How to interpret p-values correctly

Most people can calculate a p-value. Far fewer can say what it means without making one of five specific errors. Here is each one, and the correct phrasing to replace it.

9 min read · Last reviewed 7 August 2026

Why this is so hard

Statistics education is heavily weighted toward computation. Students learn to produce a p-value long before — sometimes instead of — learning what it represents. The result is a widespread ability to calculate a number correctly and describe it wrongly.

The errors below are not obscure. They appear routinely in published papers, press releases, and analytics dashboards.

Error 1: "There is a 5% chance the null hypothesis is true"

Why it is wrong: the p-value is computed by assuming the null hypothesis is true. A number derived under an assumption cannot simultaneously measure the probability of that assumption.

The two quantities differ dramatically. P(data | null) is what you have. P(null | data) is what you want, and getting there requires knowing how plausible your hypothesis was beforehand. For a hypothesis with a 50/50 prior, observing p = 0.05 leaves the probability the null is true at roughly 29% — nearly six times the figure people assume.

Say instead: "If there were no effect, we would see data this extreme about 5% of the time."

Error 2: "p = 0.001 means a stronger effect than p = 0.04"

Why it is wrong: p-values reflect both effect size and sample size, and you cannot separate them from the p-value alone.

A concrete case. Study A finds a 2-point improvement with n = 10,000 and reports p = 0.001. Study B finds a 15-point improvement with n = 30 and reports p = 0.04. Study B found an effect more than seven times larger. Study A had more data.

Say instead: compare effect sizes and their confidence intervals. That is what they are for.

Error 3: "Not significant means there is no effect"

Why it is wrong: absence of evidence is not evidence of absence. A non-significant result means your data were not surprising enough under the null — which could be because there is no effect, or because your study was too small to detect one.

The confidence interval resolves the ambiguity instantly. If your interval on a difference runs from −0.3 to +0.4 points, you have good evidence any real effect is small. If it runs from −12 to +18 points, you have learned essentially nothing. Both can produce p = 0.6.

Say instead: "We did not find sufficient evidence of an effect," followed by the interval so readers can judge what was ruled out.

Error 4: Treating 0.05 as a bright line

Why it is wrong: p = 0.049 and p = 0.051 represent virtually identical evidence. Declaring one a discovery and the other a null result is an artefact of the threshold, not a feature of the data.

This dichotomy drives real distortions: results just below the line get published and results just above it disappear, so the literature systematically overstates effects.

Say instead: report the exact p-value and treat it as a continuous measure of evidence. Our guide to what "p < 0.05" actually means works through the threshold in detail, including what to check before you act on a result that clears it.

Error 5: Ignoring how many tests you ran

Why it is wrong: α = 0.05 controls the false-positive rate for one pre-planned test. Run twenty independent tests and the probability of at least one false positive is about 64%.

This applies far more widely than people realise: testing several outcomes, several subgroups, several model specifications, or checking an A/B test daily until it turns significant. Each is a form of multiplicity, and the last one is why the practical false discovery rate in published research runs near 30%.

Say instead: pre-register your primary analysis, or apply a correction and report how many comparisons you made.

Worked example: a significant result

A trial compares a new treatment against a control. Mean improvement is 4.2 points higher in the treatment arm, with t(198) = 2.31, p = .022, 95% CI [0.61, 7.79].

What you can say: the data are inconsistent with no treatment effect at the 5% level. The best estimate is a 4.2-point improvement, and the data are compatible with anything from 0.6 to 7.8.

What you cannot say: that there is a 97.8% chance the treatment works; that the effect is definitely 4.2 points; or that a clinically meaningful benefit has been established — the interval's lower end of 0.6 may be far too small to matter to a patient.

Worked example: a non-significant result

The same trial, with 40 patients instead of 200: a 4.2-point difference with t(38) = 1.03, p = .31, 95% CI [−4.05, 12.45].

The estimated effect is identical. Only the precision changed. The interval now spans everything from a meaningful harm to a large benefit.

What you can say: this study was too small to determine whether the treatment helps. The data remain compatible with a substantial benefit.

What you cannot say: that the treatment does not work, or that the two groups are equivalent. Reporting only "p = .31, not significant" would leave readers with exactly the wrong impression — which is why intervals belong in every report.

A checklist before you report

  • Have I reported the exact p-value rather than only a threshold?
  • Have I reported an effect size in meaningful units?
  • Have I reported a confidence interval?
  • Have I said how many tests I ran?
  • Did I decide the analysis before seeing the data?
  • If the result is null, have I said whether the study could have detected an effect?

Our t-test calculator reports the effect size and interval alongside the p-value by default, so the first three come for free.

Keep reading

Ready to run the numbers?

Our calculator shows the shaded distribution, the exact p-value, and a plain-English reading of what it supports.

Open the P-Value Calculator

Frequently asked questions

If p = 0.05, what is the actual chance the null hypothesis is true?

It depends on how plausible the hypothesis was before you collected data, but it is far higher than 5%. For a hypothesis with even odds beforehand, observing p = 0.05 leaves roughly a 29% probability that the null is true. This gap between what people assume and what is true is the single largest source of overconfidence in research.

How should I report a non-significant result?

State the estimate, the confidence interval, and the p-value, then say what the interval rules out. "No significant difference (p = .31)" is close to uninformative. "The difference was 4.2 points, 95% CI [−4.1, 12.5], p = .31 — this study could not distinguish a meaningful benefit from no effect" is a genuine finding.

What is p-hacking?

Adjusting your analysis until the p-value crosses your threshold: trying multiple outcomes, dropping inconvenient observations, adding participants until it turns significant, or switching from two-tailed to one-tailed after the fact. Each choice is individually defensible, which is what makes the practice so easy to fall into unintentionally.

Should I stop using p-values?

No — used correctly they answer a genuine question. The problem is not the p-value but the habit of reporting it alone and reading it as something it is not. Report it with an effect size and an interval, treat it as continuous evidence, and it does its job.