How to find a p-value in Python
SciPy is the reference implementation most other tools are checked against, including this one. Here is every function that produces a p-value, and the method call that keeps it accurate in the far tail.
8 min read · Last reviewed 7 August 2026
The short answer
Every distribution in scipy.stats exposes a survival function,
sf, which is the upper-tail area. That is the p-value for a right-tailed test.
from scipy import stats
stats.norm.sf(2.5) # z, one-tailed
#> 0.006209665325776134
stats.t.sf(2.31, 27) # t, one-tailed
#> 0.014384284149660044
2 * stats.t.sf(abs(2.31), 27) # t, two-tailed
#> 0.028768568299320087
stats.chi2.sf(7.815, 3) # chi-square, right-tailed
#> 0.04999390297488388
stats.f.sf(4.26, 2, 27) # F, right-tailed
#> 0.02466186381480589
Chi-square and F stay right-tailed. Doubling them does not produce a two-tailed test, it produces a number with no meaning.
Use sf, never 1 - cdf
These two lines look interchangeable and are not:
stats.norm.sf(9)
#> 1.1285884059538324e-19
1 - stats.norm.cdf(9)
#> 0.0
cdf(9) is so close to 1 that double precision stores it as 1. The
subtraction then yields exactly zero, and a result that should be
1.13 × 10⁻¹⁹ is reported as impossible. sf computes the tail
without ever forming that difference.
This is the single most common source of wrong p-values in analysis code, and it is the reason this site hand-writes its distribution engine instead of subtracting from one. The methodology page publishes the comparisons against these exact SciPy values.
Tests on raw data
If you have observations rather than a statistic, the test functions return both:
from scipy import stats
# Welch's t-test — you must ask for it
stats.ttest_ind(a, b, equal_var=False)
# Student's t-test (SciPy's default)
stats.ttest_ind(a, b)
# Paired
stats.ttest_rel(before, after)
# One sample against a target
stats.ttest_1samp(x, popmean=100)
# Chi-square on a contingency table
stats.chi2_contingency([[31, 19], [22, 28]])
# Pearson correlation
stats.pearsonr(x, y)
# One-way ANOVA
stats.f_oneway(group1, group2, group3)
Note the default that catches people out: ttest_ind assumes
equal variances unless you pass equal_var=False. R does the
opposite. If a SciPy result disagrees with an R result on the same data, this is usually
why. Welch is the safer choice and it is what this calculator runs.
chi2_contingency also applies Yates' continuity correction to 2 × 2 tables by
default. Pass correction=False to match an uncorrected calculator.
Reading the result objects
Modern SciPy returns named tuples, so you can pull fields by name rather than position:
res = stats.ttest_ind(a, b, equal_var=False)
res.pvalue
res.statistic
res.df
# Confidence interval on the difference
res.confidence_interval(confidence_level=0.95)
Most test functions also take an alternative argument —
'two-sided', 'less' or 'greater' — which is cleaner
than halving a two-tailed p-value by hand and gets the direction right for you.
When SciPy is not enough
SciPy covers the standard hypothesis tests. For regression output with a full coefficient
table, use statsmodels, which reports p-values per coefficient:
import statsmodels.api as sm
model = sm.OLS(y, sm.add_constant(X)).fit()
model.pvalues
model.summary()
Formatting for a paper
A raw float is not a reportable p-value. APA style drops the leading zero and caps at three
decimals, so 0.028768568 becomes p = .029 and anything below
.001 is reported as p < .001 rather than as
p = .000, which is never true.
def apa(p):
return "p < .001" if p < 0.001 else f"p = {p:.3f}".replace("0.", ".")
The APA reporting guide covers the rest, and every calculator here emits a ready-made APA string alongside the result.