T-test calculator
Paste two columns straight from a spreadsheet, or work from the mean and standard deviation if that is all you have. You get the p-value, Cohen’s d, and a confidence interval together — because the p-value alone cannot tell you whether the difference is worth acting on.
Paste your data to begin
You get the p-value, the effect size, and a confidence interval together — because a p-value on its own does not tell you whether the result matters.
Results
P-value
Estimate
Effect size
What this means
What it does not mean
Report it (APA)
Which t-test do you need?
| Your situation | Test |
|---|---|
| Two separate groups | Independent (two-sample) t-test — the default here |
| Same subjects, twice | Paired t-test — before/after, matched pairs |
| One group vs. a fixed value | One-sample t-test — against a target, a norm, or zero |
| Three or more groups | One-way ANOVA |
| Rates or proportions | Two-proportion Z-test |
If you only have the mean and standard deviation
Plenty of people arrive here holding summary statistics rather than raw numbers — a table in a published paper, a figure from a report, a line in a colleague’s email. Switch the input above from Raw data to Mean and SD, enter n, the mean and the standard deviation for each group, and you get the same p-value, Cohen’s d and confidence interval the underlying values would have given you.
Nothing is lost statistically, because a t-statistic is built entirely out of those three numbers. For two independent groups, using Welch’s formula:
t = (mean₁ − mean₂) / √(sd₁²/n₁ + sd₂²/n₂)
Worked example. Group 1: mean 52.3, SD 8.1, n = 30. Group 2: mean 47.9, SD 9.4, n = 28. The numerator is 4.4. The denominator is √(65.61/30 + 88.36/28) = √(2.187 + 3.156) = 2.311. So t = 1.904, the Welch–Satterthwaite degrees of freedom come to 53.5, and p = 0.0624 — a difference of 4.4 points that misses significance at 0.05. Type those six numbers in above and that is exactly what you will see.
For one sample against a fixed value the formula is t = (mean − μ) / (sd / √n) with df = n − 1; choose “One sample vs. a value” and enter your μ₀.
The one test summary statistics cannot support is the paired t-test, so that option disappears in this mode. A paired test works on the difference within each pair, and the standard deviation of those differences is not recoverable from two group SDs — it depends on how strongly the pairs are correlated, which summary statistics simply do not record. A calculator that lets you do it anyway is running an independent-samples test and labelling it paired.
One thing to check before you start: does your source report the standard deviation or the standard error? The two are routinely confused and they differ by a factor of √n, so mistaking one for the other at n = 30 changes your t-statistic more than fivefold. If the number came off an error bar, read the caption. To convert, sd = se × √n.
Raw data is still the better input when you have it. The summary under each box reports the n, mean and SD actually parsed, which is a fast way to check your figures against their source — and only the raw values let you see the skew and outliers that would make a t-test the wrong tool in the first place.
Pasting your data
The boxes accept whatever a spreadsheet gives you: values separated by commas, tabs, semicolons, spaces, or newlines. Copying a column straight out of Excel, Google Sheets, or a CSV works without any cleanup.
Anything that is not a number is reported rather than silently dropped. That matters — quietly skipping a stray N/A changes your sample size and therefore your p-value, and you would never know. The running summary under each box shows the n, mean, and standard deviation actually being used, so you can confirm the parse before trusting the result.
Why Welch’s test is the default
The classic Student’s t-test assumes both groups share the same population variance. That assumption is rarely tested, frequently violated, and when it fails the test gives p-values that are too small — the direction that produces false positives.
Welch’s t-test drops the assumption. It adjusts the degrees of freedom via the Welch–Satterthwaite equation, which is why you will usually see a fractional df such as 21.34 in the results. That is correct and expected. When the variances genuinely are equal, Welch’s costs you almost no power, which makes it the better default in essentially every case. R has used it as the default for years.
The equal-variance option is still available for coursework that requires the pooled version. If you choose it and the variances differ by more than a factor of four, the calculator says so.
Read the interval, not just the p-value
The confidence interval is the most useful number on the results panel and the one most often ignored. It tells you the range of differences your data are compatible with, in the original units of your measurement.
If you want the effect on a standardised scale instead, the effect size calculator reports Cohen's d and Hedges' g with an interval of their own.
A significant result whose interval runs from 0.1 to 12.0 is not the same finding as one running from 5.8 to 6.2, even at an identical p-value. And a non-significant result whose interval spans −0.05 to 0.06 genuinely suggests no meaningful effect, while one spanning −4 to +9 means your study simply could not tell. Reporting only "p > .05" throws that distinction away.
Frequently asked questions
Can I calculate a p-value from just the mean and standard deviation?
Yes, as long as you also have the sample sizes. Switch the input above to "Mean and SD" and enter n, the mean and the standard deviation for each group — you get the same p-value, Cohen's d and confidence interval the raw values would have produced, because a t-statistic is built entirely out of those three numbers. The one test this cannot support is a paired t-test, which needs the differences within each pair. The other thing you lose is the ability to see skew and outliers, which is exactly where a t-test is most likely to mislead you.
Is the number I have the standard deviation or the standard error?
They differ by a factor of √n — se = sd / √n — and they are constantly mixed up in published tables and chart captions. Using an SE where the formula wants an SD makes your t-statistic far too large and your p-value far too small. If a figure came from an error bar, check the caption for which one it plots. To convert back, sd = se × √n.
What is the difference between a paired and an independent t-test?
Use a paired t-test when each value in one group corresponds to a specific value in the other — the same person measured before and after, or matched pairs. Use an independent test when the two groups contain different subjects with no natural pairing. Pairing removes between-subject variation, so a paired test is considerably more powerful when it applies. Using it when the data are not genuinely paired is a serious error.
Should I use Welch's t-test or Student's pooled t-test?
Welch's, in almost all cases. It does not assume the two groups have equal variances, and when they happen to be equal it loses almost nothing in power. Statisticians have recommended it as the default for decades, and R uses it by default. This calculator uses Welch's unless you explicitly tick the equal-variance box.
How much data do I need for a t-test?
Technically two values per group, but that gives almost no power to detect anything. As a practical floor, aim for at least 15 to 20 per group for a moderate effect. If your groups are small and clearly non-normal, a Mann-Whitney U test is a better choice than a t-test.
What is Cohen's d and why does it matter?
Cohen's d expresses the difference between two means in standard-deviation units, which makes it comparable across studies and independent of sample size. Conventionally 0.2 is small, 0.5 medium, and 0.8 large. It matters because a p-value only tells you an effect is detectable — with 100,000 observations, a d of 0.01 will be highly significant and completely irrelevant.
Does my data need to be normally distributed?
The t-test assumes the sampling distribution of the mean is roughly normal, which is a weaker requirement than the data themselves being normal. Thanks to the central limit theorem, samples above about 30 per group are quite robust to non-normality. Heavy skew or strong outliers in small samples are the real problem — inspect your data before trusting the result.