One-tailed vs two-tailed tests
Switching from two-tailed to one-tailed halves your p-value without adding any evidence. That is why the choice must be made before you see your data — and why the default should be two-tailed.
6 min read · Last reviewed 7 August 2026
What the difference actually is
A two-tailed test asks whether your result differs from the null in either direction. The rejection region is split between both tails of the distribution, so at α = 0.05 you have 2.5% in each.
A one-tailed test asks whether your result differs in one specified direction. The entire 5% sits in a single tail, which makes it easier to reach significance in that direction — and impossible in the other, however extreme the result.
In concrete terms: with a two-tailed test at α = 0.05 you need |z| > 1.96. With a one-tailed test you need only z > 1.645. A z-score of 1.8 is not significant two-tailed and is significant one-tailed. Same data, opposite conclusion.
Why this is so easy to abuse
Because a one-tailed test exactly halves the p-value, switching to one after seeing your data is one of the simplest forms of p-hacking. A result at p = 0.08 becomes p = 0.04 with no new information at all — purely by redescribing the hypothesis you claim to have had.
The tell is that the justification arrives after the result. If you would have run a two-tailed test had the effect gone the other way, then you were always running a two-tailed test, and reporting otherwise misstates your false-positive rate.
When a one-tailed test is legitimate
All of the following must hold:
- The direction was specified before collecting data, ideally in a pre-registration.
- An effect in the other direction would be genuinely uninteresting — not merely unexpected or disappointing, but of no consequence to your decision.
- You would report the same conclusion regardless. If a large effect the wrong way would change your mind, you needed both tails.
That second condition is the strict one, and it disqualifies most cases. A non-inferiority trial genuinely satisfies it: you only care whether the new treatment is worse by more than a margin. Most research does not.
A/B testing: use two-tailed
A/B testing tools frequently default to one-tailed, and the reasoning offered is that you only want to know whether the variant wins. That reasoning is wrong on its own terms.
You care very much if your variant performs worse. Shipping a change that quietly reduces conversion is precisely the outcome the test exists to prevent. If you would roll back on a significant negative result, you are conducting a two-tailed test and should report it as one. Our A/B test calculator defaults to two-tailed for this reason.
Chi-square and F tests are a separate case
Chi-square and F tests are always right-tailed, and this is not the same choice at all. Both statistics accumulate squared deviations, so any departure from the null — in any direction — makes them larger. All the evidence against the null lives in the upper tail by construction.
This is why our chi-square and F-test calculators lock the direction rather than offering a two-tailed option. A doubled chi-square p-value is not a conservative choice; it is simply not a meaningful number.
The practical rule
Use two-tailed unless you can satisfy all three conditions above, and decided so in advance. The cost is a slightly larger p-value. The benefit is a result you can defend, and a false-positive rate that is what you claim it is.
If you are genuinely unsure, that uncertainty is itself the answer: a one-tailed test requires a directional prediction confident enough to stake the analysis on.