What is a p-value?
A p-value is one of the most used and most misunderstood numbers in science. Here is what it actually measures — and why the intuitive reading of it is backwards.
7 min read · Last reviewed 7 August 2026
The one-sentence definition
A p-value is the probability of getting data at least as extreme as yours, assuming the null hypothesis is true.
Every word in that sentence is doing work, and the phrase that matters most is assuming the null hypothesis is true. The p-value is computed inside a hypothetical world where there is no effect. It measures how unusual your data would look in that world. It does not, and cannot, tell you whether you are living in it.
A courtroom analogy
Imagine a trial. The null hypothesis is the presumption of innocence: the defendant did not do it. The evidence is your data.
The p-value asks: if this person were genuinely innocent, how likely is it that evidence this incriminating would turn up anyway? A very small p-value means the evidence would be a remarkable coincidence under innocence — which is grounds for rejecting the presumption.
Notice what the question is not. It is not "how likely is it that the defendant is innocent?" That would require knowing how often people in this situation are guilty in general — a prior — and the p-value has no access to it. This is exactly the trap people fall into when reading p-values as the probability the null hypothesis is true.
The analogy also explains why "not significant" is not "innocent." A verdict of not guilty means the evidence was insufficient, not that innocence was proven. A non-significant p-value means precisely the same thing.
The p-value is an area
The most concrete way to understand a p-value is geometrically. Take the distribution your test statistic would follow if the null hypothesis were true — a bell curve, for a Z or t test. Mark where your observed statistic landed. The p-value is the area of the curve beyond that point.
That is all it is. If your statistic sits near the middle, most of the curve is on the "at least this extreme" side, and the p-value is large — your result is unremarkable. If it sits far out in the tail, only a sliver of area lies beyond it, and the p-value is small.
Our p-value calculator shades this area live as you type, which is the fastest way to build the intuition. Move your statistic outward and watch the shaded region shrink; that shrinking region is your p-value.
Four things a p-value is not
1. It is not the probability that the null hypothesis is true
This is the big one. The calculation begins by assuming the null hypothesis, so it cannot also evaluate it. A p-value of 0.05 does not mean a 5% chance there is no real effect. Once you account for how plausible the hypothesis was to begin with, the actual probability the null is true after observing p = 0.05 is frequently 30% or higher.
2. It is not the probability that your result was due to chance
A subtler version of the same error. The p-value is computed under the assumption of chance alone; it is a conditional probability, not a probability about causes.
3. It is not a measure of effect size
A p-value of 0.001 does not describe a bigger effect than one of 0.04 — it may simply reflect a bigger sample. With enough data, a difference far too small to matter will produce a spectacularly small p-value. This is why effect sizes and confidence intervals belong next to every p-value you report.
4. It is not the probability of replication
p = 0.05 does not mean a 95% chance of replicating. Replication probability depends on the true effect size and the power of the replication study, neither of which the p-value knows anything about.
Where 0.05 came from
The 0.05 threshold has no mathematical justification. Ronald Fisher proposed it in the 1920s as a convenient rule of thumb, remarking that it was up to the researcher to decide what counted as convincing. It stuck, and it hardened into a rule he never intended.
Different fields set the line elsewhere for good reasons. Particle physics requires roughly 0.0000003 — the "five sigma" standard — because it runs enormous numbers of comparisons and the cost of a false discovery is high. Exploratory work sometimes accepts 0.10. The right threshold depends on what a false positive costs you, and it should be chosen before you look at your data.
How to use p-values well
- Report the exact value, not just whether it cleared a threshold. "p = .043" carries more information than "p < .05".
- Always pair it with an effect size and a confidence interval. The p-value says an effect is detectable; only the interval says how big it might be.
- Decide your threshold in advance, along with your sample size and your analysis plan.
- Treat it as continuous evidence, not a switch. p = 0.049 and p = 0.051 are the same result.
- Remember multiplicity. Twenty tests at α = 0.05 give you roughly a 64% chance of at least one false positive.