Methodology & accuracy
Every algorithm this site uses, why it was chosen, and the evidence that it is correct. If you are going to put a number from a free web tool into a paper, you are entitled to know exactly where it came from.
Last reviewed: 9 August 2026 · 83 automated tests against SciPy 1.17
The short version
- All five distributions are implemented from published numerical methods, with no third-party statistics library.
- Tail probabilities use a dedicated survival function, never
1 − CDF. - Results are validated against SciPy 1.17 to a relative error below 1 × 10⁻¹³, by 83 automated tests across 5 suites.
- Every cell of every published reference table — 508 of them — is checked individually, not sampled.
- Everything runs in your browser. No input is ever transmitted, logged, or stored.
The p-value formula, test by test
There is no single p-value formula. A p-value is always the same idea — the area of the null distribution beyond your test statistic — but the statistic, and therefore the formula, changes with the test. These are the ones implemented here.
| Test | Statistic | p-value |
|---|---|---|
| One-sample t | t = (x̄ − μ₀) / (s/√n), df = n − 1 | 2 · sft,df(|t|) |
| Two-sample t (Welch) | t = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂), Welch–Satterthwaite df | 2 · sft,df(|t|) |
| Paired t | t = d̄ / (sd/√n) on the within-pair differences | 2 · sft,n−1(|t|) |
| One-sample Z | z = (x̄ − μ₀) / (σ/√n) | 2 · sfN(|z|) |
| Two-proportion Z | z = (p̂₁ − p̂₂) / √(p̄(1−p̄)(1/n₁ + 1/n₂)) | 2 · sfN(|z|) |
| Chi-square, independence | χ² = Σ (O − E)² / E, df = (r−1)(c−1) | sfχ²,df(χ²) — right tail only |
| One-way ANOVA | F = MSbetween / MSwithin, df = (k−1, N−k) | sfF(F) — right tail only |
| Pearson correlation | t = r√(df/(1−r²)), df = n − 2 | 2 · sft,df(|t|) |
The sf in every right-hand column is a survival function evaluated directly in the upper tail, which is the part most implementations get wrong — see below. For a one-tailed test, drop the factor of 2 and use the tail matching your hypothesis.
Why no statistics library
The obvious choice for a JavaScript statistics tool is jStat. We did not use it, for two reasons. It carries long-standing open accuracy issues in its gamma and chi-square cumulative distribution functions, which are precisely the functions a p-value calculator depends on. And it ships roughly 40 KB to every visitor to provide five functions.
Implementing the underlying special functions directly is about three hundred lines of well-documented numerical analysis. It gives us exact control over the tail behaviour that matters most, and it means the entire calculator is a few kilobytes of JavaScript.
The precision problem that defines this tool
This is the single most important implementation decision on the site, and it is why this calculator gives different answers from several popular ones in the far tail.
A cumulative distribution function Φ(x) returns the probability of a value at or below x. To get an upper-tail p-value, the naive approach computes 1 − Φ(x). That works fine in the middle of the distribution and fails completely in the tail.
At z = 9, Φ(9) is approximately 0.999999999999999999999. Double-precision floating point stores about 16 significant decimal digits, so that value rounds to exactly 1.0. Subtracting gives 1 − 1 = 0, and the calculator reports p = 0 — a number that is never a valid p-value.
Instead, every distribution here exposes a survival function computed directly in the upper tail:
- Normal: via the complementary error function
erfc, evaluated with a Chebyshev expansion that is accurate in the tail by construction — never as1 − erf(x). - Student's t: by exploiting symmetry,
sf(x) = cdf(−x), which sidesteps the subtraction entirely. - Chi-square: via the regularised upper incomplete gamma function Q(a, x), computed by its own continued fraction rather than as 1 − P(a, x).
- F: via the symmetry of the incomplete beta function,
I_(df₂/(df₁x+df₂))(df₂/2, df₁/2).
The result: z = 9 correctly returns 1.13 × 10⁻¹⁹, and the calculator stays accurate down to about 10⁻³⁰⁰, the limit of double-precision arithmetic itself.
Algorithms
| Function | Method |
|---|---|
| log Γ(x) | Lanczos approximation, g = 7, n = 9, with reflection for x < 0.5 |
| erfc(x) | Chebyshev expansion, 24 coefficients (Numerical Recipes §6.2.2) |
| P(a, x), Q(a, x) | Series for x < a + 1; continued fraction via Lentz's method otherwise |
| I_x(a, b) | Continued fraction via Lentz's method, with the symmetry identity past the mode |
| Φ⁻¹(p) | Acklam's rational approximation, refined by two Halley iterations |
| Other quantiles | Bisection on the CDF to a relative tolerance of 1e-14 |
Everything that can overflow is computed in log space. Γ(1000) is roughly 10²⁵⁶⁴ and cannot be represented at all, but log Γ(1000) ≈ 5905.22 is unremarkable — so the F-distribution density, which involves a ratio of large gamma functions, is assembled from logarithms and exponentiated only once at the end. This is why the calculator handles df in the hundreds without returning NaN, which several competitors do. For a worked example of where this shows up in practice, see how these results compare with GraphPad QuickCalcs.
Validation
Every figure below is produced by this calculator and compared against SciPy 1.17. The reference script is committed to the repository so anyone can regenerate the values, and the comparisons run as an automated test suite — if any of them drift, the build fails.
| Function | This calculator | SciPy 1.17 | Rel. error |
|---|---|---|---|
| Φ(1.96) | 0.9750021048517796 | 0.9750021048517795 | 1.1e-16 |
| sf(8) normal | 6.220960574271822e-16 | 6.22096057427174e-16 | 1.3e-14 |
| t(28).sf(2.14) | 0.02060458039399649 | 0.020604580393996423 | 3.2e-15 |
| t(17.3).sf(2.5) | 0.0113737500879005 | 0.011373750087900572 | 6.4e-15 |
| χ²(10).sf(25) | 0.0053455054871340765 | 0.005345505487134069 | 1.5e-15 |
| χ²(2).sf(100) | 1.9287498479639188e-22 | 1.9287498479639183e-22 | 2.4e-16 |
| F(3,16).sf(5.29) | 0.010015837244915275 | 0.010015837244915286 | 1.0e-15 |
| lgamma(1000) | 5905.220423209181 | 5905.220423209181 | 0 |
| erfc(6) | 2.151973671249893e-17 | 2.151973671249891e-17 | 1.0e-15 |
Tests that start from data
The table above validates the distribution functions. The calculators that take raw data add arithmetic on top of them, so those are validated separately against SciPy and statsmodels — f_oneway for ANOVA, chi2_contingency with correction=False for contingency tables, and proportion_confint(method='wilson') for proportion intervals.
| Quantity | This calculator | Reference | Rel. error |
|---|---|---|---|
| ANOVA F, 3 groups | 26.718975705843693 | 26.718975705843736 | 1.6e-15 |
| ANOVA p | 4.093470054585312e-6 | 4.093470054585261e-6 | 1.3e-14 |
| η² (eta squared) | 0.7480330882352939 | 0.748033088235294 | 1.5e-16 |
| χ² on a 2 × 2 table | 1.0204081632653061 | 1.0204081632653061 | 0 |
| Cramér's V | 0.10204081632653061 | 0.10204081632653061 | 0 |
| One-sample Z, p | 0.045500263896358174 | 0.0455002638963582 | 6.1e-16 |
| Wilson CI lower, 45/300 | 0.11403246871185316 | 0.11403246871185314 | 1.2e-16 |
| Wilson CI upper, 45/300 | 0.194817611145329 | 0.19481761114532903 | 1.4e-16 |
| Cohen's d | 0.5028037869188999 | 0.5028037869188999 | 0 |
| Hedges' g | 0.4960396104132645 | 0.4960396104132645 | 0 |
The published tables are validated cell by cell
The reference tables on the p-value tables page and the critical values on the critical value calculator are generated from these same functions at build time rather than transcribed from a textbook. Printed statistical tables are copied from other printed tables, and the transcription errors that creep in tend to outlive the book they started in.
Generating them only helps if the generator is right across the whole grid, so the suite walks every one of the 508 published cells — every degree of freedom, every α, every row — against stored SciPy values. Not a sample. If one cell were wrong, the build would fail and this page would be making a false claim.
Alongside the reference comparisons, the suite checks identities that must hold if the implementations are mutually consistent:
t(df)² = F(1, df)— agrees to 1 × 10⁻¹¹z² = χ²(1)— agrees to 1 × 10⁻¹²cdf(x) + sf(x) = 1wherever both are representable — agrees to 1 × 10⁻¹⁴Φ(Φ⁻¹(p)) = pacross eight orders of magnitude — agrees to 1 × 10⁻¹¹
These matter because a reference comparison only checks the values someone thought to test, while an identity has to hold everywhere. An error in either the t or the F implementation breaks the first one immediately.
Statistical choices
Several decisions here differ from what other calculators do, deliberately.
Chi-square and F are locked to right-tailed
Both statistics accumulate squared deviations, so every departure from the null hypothesis pushes them upward and all the evidence sits in the upper tail. A "two-tailed" chi-square p-value — obtained by doubling — is simply not a meaningful quantity. Rather than compute one on request, the calculator disables the option and explains why.
Welch's t-test is the default
The pooled-variance Student's t-test assumes both groups share a population variance. When that fails, it produces p-values that are too small — the direction that manufactures false positives. Welch's test drops the assumption at almost no cost in power when the variances happen to be equal. R has defaulted to it for years, and so do we. The pooled version stays available for coursework that requires it.
Pooled and unpooled standard errors in proportion tests
The two-proportion test statistic uses the pooled standard error, which is correct under a null hypothesis that assumes the rates are equal. The confidence interval uses the unpooled standard error, which is correct for estimating a difference not assumed to be zero. Using one for both is a common bug that yields an interval contradicting its own p-value.
Proportion intervals use Wilson, not Wald
The interval nearly every textbook teaches for a proportion is Wald, p̂ ± z√(p̂(1−p̂)/n). It is simple and it is unreliable: its actual coverage falls well below the nominal level when p̂ approaches 0 or 1, it can return limits below 0 or above 1, and at 0 successes out of 20 it collapses to zero width — asserting perfect certainty from twenty observations.
The confidence interval calculator inverts the score test instead, giving the Wilson interval. It stays inside [0, 1] by construction, holds its coverage at extreme proportions, and returns roughly 0% to 16% for 0 out of 20. This has been the recommendation in the statistical literature for decades; most calculators still ship Wald because it is the one people were taught.
An omnibus test is not given an effect estimate
A one-way ANOVA establishes that at least one group mean differs. It cannot say which, or by how much, and no post-hoc-free summary of "the difference" exists. The same applies to a contingency table larger than 2 × 2.
So in both cases the estimate panel reports no value and says why, rather than printing something quotable that would not survive scrutiny. What is shown instead is the material you need to go further: the per-group means and standard deviations for ANOVA, the full expected-count table for chi-square.
Effect sizes and intervals are shown, not hidden
The American Statistical Association's 2016 statement on p-values is explicit that a p-value alone is an inadequate basis for a scientific conclusion. Every data-driven calculator here reports an effect size and a confidence interval next to the p-value, at equal visual weight, because the interval usually answers the question you actually have.
Known limits
Where this tool stops, stated plainly:
- Double precision. p-values below roughly 10⁻³⁰⁰ are reported as a bound rather than a value. This is a limit of f64 arithmetic, not of the algorithms.
- Very large degrees of freedom. Above about 10⁶ the t-distribution is numerically indistinguishable from the normal. That is a mathematical fact, not an error, but the quantile search will be slower.
- Fisher's exact test is not included. When expected counts are too small for the chi-square approximation, the calculator warns you and names the right alternative rather than pretending.
- No multiple-comparison correction. If you are running many tests, your effective false-positive rate is far above your nominal α, and no single-test calculator can fix that for you.
- No non-parametric tests. Mann-Whitney, Wilcoxon and Kruskal-Wallis are not implemented. Where the data suggest one is needed, the pages say so.
- No post-hoc comparisons. A significant ANOVA needs Tukey's HSD or an equivalent to locate the difference, and that is not built here.
- A paired t-test needs raw data. The summary-statistics mode covers one- and two-sample tests only. The standard deviation of within-pair differences cannot be recovered from two group SDs — it depends on the pair correlation, which summary statistics do not record. A calculator that offers it anyway is running an independent-samples test under the wrong label.
- Sample sizes are normal approximations. The sample size calculator uses the standard closed-form formulas, which get optimistic when the required n is small or the rates are extreme. It flags the cases where that bites.
- ANOVA assumes equal variances. Welch's ANOVA is not implemented; the calculator warns when the group standard deviations diverge enough for it to matter.
References
- Press, W. H. et al. Numerical Recipes: The Art of Scientific Computing, 3rd ed., Cambridge University Press, 2007 — §6.2 incomplete gamma, §6.4 incomplete beta, §6.2.2 erfc.
- Lanczos, C. "A Precision Approximation of the Gamma Function," SIAM Journal on Numerical Analysis, 1964.
- Lentz, W. J. "Generating Bessel functions in Mie scattering calculations using continued fractions," Applied Optics, 1976.
- Acklam, P. J. "An algorithm for computing the inverse normal cumulative distribution function," 2003.
- Welch, B. L. "The generalization of Student's problem when several different population variances are involved," Biometrika, 1947.
- Wasserstein, R. L. & Lazar, N. A. "The ASA Statement on p-Values: Context, Process, and Purpose," The American Statistician, 2016.
- Greenland, S. et al. "Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations," European Journal of Epidemiology, 2016.
Found an error?
If any result here disagrees with R, SciPy, SAS, or Stata by more than floating-point noise, that is a bug and we want to know. Send the exact inputs and both outputs — a reproducible discrepancy will be fixed and recorded on this page.