Skip to main content

Chi-square to p-value calculator

Enter your χ² statistic and degrees of freedom for an exact right-tailed p-value. Chi-square tests are always right-tailed, so the calculator sets that for you.

Hypothesis direction

Testing for a difference in either direction

Results update as you type — there is no submit button.

Enter a test statistic to begin

Your p-value, a plain-English reading, and a shaded distribution chart appear here instantly.

Where this statistic comes from

You will typically be holding one of these:

  • A chi-square test of independence on a contingency table
  • A goodness-of-fit test against expected frequencies
  • McNemar's test on paired categorical data
  • A likelihood-ratio test comparing two nested models

Working out the degrees of freedom

Chi-square degrees of freedom depend on the test, not on your sample size:

  • Test of independence: (rows − 1) × (columns − 1). A 2 × 2 table gives df = 1; a 3 × 4 table gives df = 6.
  • Goodness of fit: number of categories − 1, minus one more for every parameter you estimated from the data.
  • Likelihood-ratio test: the difference in the number of parameters between the two models.

Sample size never enters the df calculation. If you find yourself using n, you have the wrong formula.

If you are holding the raw table rather than a χ² value, none of this needs working out by hand — the chi-square calculator takes a 2 × 2 up to 5 × 5 contingency table, derives the statistic and its degrees of freedom, and prints the expected counts it used.

Why chi-square is always right-tailed

A chi-square statistic accumulates squared deviations between what you observed and what the null hypothesis predicts. Every discrepancy makes it larger, in any direction — so all the evidence against the null sits in the upper tail. There is no such thing as a meaningfully small χ², and doubling the tail for a "two-tailed" test is simply wrong.

Some calculators will happily hand you a two-tailed chi-square p-value anyway. This one locks the direction to right-tailed and says why, rather than returning a number you should not use.

What your chi-square p-value means

The p-value answers one narrow question: if the variables really were unrelated — or your data really did follow the distribution you proposed — how often would a table this uneven turn up by chance alone? A χ² of 6.0 on 1 degree of freedom gives p = 0.0143, so a table that lopsided would appear roughly 14 times in every 1,000 tries if nothing were going on.

Below your significance level, conventionally 0.05, the association is called statistically significant: chance alone is a strained explanation for what you observed. Above it, you have not shown the variables are independent — you have failed to show they are not, which is a weaker claim and a statement about your evidence rather than about the world. With small tables and modest samples that happens constantly.

What the p-value never tells you is how strong the association is, because χ² scales with sample size. A 55/45 versus 45/55 split gives χ²(1) = 1.02, p = 0.31 at n = 100 — and the identical split at n = 2,000 gives χ²(1) = 20.0, p < 0.001. Same pattern, same practical importance, opposite verdicts.

So report a measure of strength next to the p-value: Cramér's V for a table of any size, or the phi coefficient for a 2 × 2. And look at the cell percentages, which usually describe the finding better than either number does.

The expected-count assumption

The chi-square distribution is an approximation to the true sampling distribution of the statistic, and it degrades when expected cell counts get small. The usual rule: every expected count should be at least 5, and none below 1.

Note that this is about expected counts, not observed ones — a cell with zero observations is fine if its expected count is comfortable. When the assumption fails, use Fisher's exact test for a 2 × 2 table, or combine sparse categories if that is defensible on subject-matter grounds.

Frequently asked questions

What does the chi-square p-value mean?

It is the probability of getting a table at least as uneven as yours if the null hypothesis were true — if the two variables were independent, or if your counts really followed the expected distribution. A p of 0.02 means a table this lopsided would arise about twice in every hundred tries by chance alone. It is not the probability that the null hypothesis is true, and it is not a measure of how strong the association is.

Is my chi-square result significant?

Compare the p-value to your significance level, usually 0.05. Below it, the result is conventionally called statistically significant. Above it, you have not established independence — only that your data were not surprising enough to rule chance out. Because chi-square grows with sample size, always report a strength measure such as Cramér's V alongside it: a significant χ² in a very large table can accompany a difference too small to act on.

What chi-square value is significant?

It depends on degrees of freedom. At α = 0.05 the critical value is 3.841 for df = 1, 5.991 for df = 2, 7.815 for df = 3, and 11.070 for df = 5. Your χ² must exceed the critical value for its df to be significant.

Is a chi-square test one-tailed or two-tailed?

Always one-tailed, specifically right-tailed. The statistic sums squared deviations, so any departure from the null makes it larger. All the evidence against the null hypothesis lives in the upper tail, and a doubled two-tailed p-value would be incorrect.

What does a large chi-square value mean?

It means your observed counts are far from what the null hypothesis predicts. In a test of independence, a large χ² is evidence that the two variables are associated. In a goodness-of-fit test, it is evidence your data do not follow the distribution you proposed.

When is a chi-square p-value unreliable?

When the expected counts behind it are too small. The chi-square distribution is an approximation to the true sampling distribution of the statistic, and it degrades once any expected cell count falls below about 5 — at which point the p-value cannot be trusted however precisely it was computed. Note that this is about expected counts, not observed ones. For a 2 × 2 table the fix is Fisher's exact test, which is exact at any sample size; for larger tables, merging sparse categories where that is defensible on subject-matter grounds.