Skip to main content

Effect size calculator

Cohen’s d and Hedges’ g for two groups, with a confidence interval on the estimate — because an effect size measured from twenty observations is a good deal less certain than the two decimal places suggest.

What you have

Enter two groups to begin

An effect size expresses a difference in standard deviations, which makes it comparable across studies and independent of how much data you collected.

Why a p-value is not enough

A p-value answers "could this be chance?" An effect size answers "how big is it?" They are different questions and only one of them is usually the question you had.

The gap opens up because p depends on sample size and d does not. Two groups differing by 0.3 standard deviations produce p = 0.35 with 20 per group and p < 0.001 with 500 per group. The effect is identical in both. Only the evidence changed.

That is also why meta-analyses run on effect sizes rather than p-values. d is expressed in standard deviations, so studies that measured the same construct on different scales become directly comparable — which a collection of p-values never can be.

Cohen's d and Hedges' g

They estimate the same thing. d divides the difference in means by the pooled standard deviation; g multiplies d by a correction factor that removes the upward bias d carries in small samples.

Per groupCorrection factorEffect on d = 0.80
50.9030.72
100.9580.77
200.9800.78
500.9920.79

So: report g when groups are small, and it stops mattering above about 20 per group. Many software packages label g as "corrected d" or "bias-corrected d", which is the same thing.

Cohen's benchmarks, and their limits

0.2 small, 0.5 medium, 0.8 large. Cohen proposed these for behavioural research and was openly uncomfortable doing so — he wanted researchers to calibrate against their own literature, and offered the numbers as a stopgap for fields that had none.

Fifty years later they get applied to everything, which is how a d of 0.25 in an education intervention gets dismissed as "small" when it would represent a substantial policy win, and a d of 0.9 in a lab task where the manipulation is obvious gets celebrated as "large" when it is unremarkable. The useful benchmark is the distribution of effects other studies in your area report.

A more concrete reading: d = 0.5 means the average member of one group scores above about 69% of the other group. d = 0.8 puts them above 79%. Even a "large" effect leaves the two distributions overlapping heavily, which is worth remembering before an effect size gets described as dramatic.

The interval on d

An effect size is an estimate, and at typical sample sizes it is not a precise one. Twelve per group giving d = 0.8 comes with an interval running from −0.03 to 1.63 — the same data are compatible with no effect and with an enormous one.

This matters most in the small-study literature, where an eye-catching d gets quoted without its interval, fails to replicate, and the failure gets treated as a puzzle. Usually there was no puzzle: the original estimate was consistent with almost anything, and the replication landed somewhere else inside that range.

The calculator reports the interval next to the point estimate for that reason. If it crosses zero, say so when you write it up — that is a different finding from an interval that does not.

Frequently asked questions

How do you calculate Cohen’s d?

Divide the difference between the two group means by the pooled standard deviation: d = (mean₁ − mean₂) / s_pooled, where s_pooled = √(((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)). Paste your raw data above, or switch to "Mean and SD" and enter n, mean and standard deviation for each group.

What is a large effect size?

Cohen suggested 0.2 as small, 0.5 as medium and 0.8 as large — and said explicitly that he was offering them reluctantly, for behavioural research, in the absence of anything better. They are not universal. In education, an effect of 0.2 can justify a policy change; in a drug trial for a mild condition, 0.5 might not be worth the side effects. The benchmark that matters is what other work in your field finds.

What is the difference between Cohen’s d and Hedges’ g?

They estimate the same quantity, but d is biased upward in small samples and g applies the standard correction for it. The correction is negligible above about 20 per group and matters below that: with 5 per group it shrinks the estimate by about 10%. Report g when your groups are small, and note that many packages label g as "corrected d".

Why does an effect size need a confidence interval?

Because it is an estimate like any other, and a surprisingly imprecise one at small sample sizes. Twelve per group giving d = 0.8 typically produces an interval spanning roughly 0 to 1.6 — the data are compatible with no effect and with a very large one. Quoting the 0.8 alone implies a precision the study does not have.

Can I have a significant p-value and a tiny effect size?

Yes, and it is routine in large datasets. p depends on both the size of the effect and the amount of data; d depends only on the effect. With 100,000 observations per group, a d of 0.01 will clear p < 0.001 and mean nothing in practice. The reverse also happens: a substantial d in a small study can miss significance entirely, which is a statement about the study rather than the effect.

Which effect size should I use for other tests?

Cohen's d for a difference between two means. Cohen's h for a difference between two proportions — the A/B test calculator reports it. Cramér's V for a contingency table, on the chi-square calculator. Eta squared for ANOVA. r² for a correlation. They all answer the same question in the units of the test you ran: how much of what you measured does the grouping actually account for?