Statistical significance is the least interesting thing about the results. You should describe the results in terms of measures of magnitude – not just does treatment affect people, but how much does it affect them.
Effect size is a quantitative measure of the magnitude of the experimental effect. The larger the effect size the stronger the relationship between two variables.
You can look at the effect size when comparing any two groups to see how substantially different they are.
Typically, research studies will comprise an experimental group and a control group. The experimental group may be an intervention or treatment which is expected to affect a specific outcome.
For example, we might want to know the effect of therapy on treating depression. The effect size value will show whether the therapy has had a small, medium, or large effect on depression.
Key Takeaways
- What it measures: Effect size shows how large a difference or relationship is. A p-value only tells you whether the result is statistically significant, not how big it is.
- Two common measures: Cohen’s d compares two group means, such as a treatment group and a control group. Pearson’s r describes how strongly two variables are related.
- Cohen’s benchmarks: For d, 0.2 counts as small, 0.5 as medium, and 0.8 as large. For r, the same bands sit at roughly 0.1, 0.3 and 0.5.
- Independent of sample size: A small effect can be statistically significant in a large sample. A large effect can also miss significance in a small sample.
- A modern caveat: Recent research argues Cohen’s benchmarks undervalue small effects, particularly in personality and social psychology, where they can still be practically meaningful.
Calculate and interpret effect sizes
Effect sizes measure either the strength of association between variables or the size of the difference between group means.
Cohen’s d
Cohen’s d is an appropriate effect size for comparing two means, often accompanying the reporting of t-test and ANOVA results. It is also widely used in meta-analysis.
To calculate the standardized mean difference, subtract the mean of one group from the other (M1 – M2). Then divide the result by the standard deviation (SD) of the population the groups were sampled from.

A d of 1 means the two groups differ by 1 standard deviation. A d of 2 means they differ by 2, and so on. Standard deviations are equivalent to z-scores (1 standard deviation = 1 z-score).

Cohen suggested that d = 0.2 be considered a “small” effect size, 0.5 a “medium” effect size, and 0.8 a “large” effect size. If the difference between two groups’ means is less than 0.2 standard deviations, the difference is negligible, even if it is statistically significant.
Pearson r correlation
This measure of effect size shows how strongly two variables are related. Pearson’s r ranges from -1 (a perfect negative correlation) to +1 (a perfect positive correlation).

According to Cohen (1988, 1992), the effect size is small when r is around 0.1, medium around 0.3, and large above 0.5.

Squaring r gives the coefficient of determination (r²), the percentage of variance the two variables share. An r of 0.30 explains about 9% of the variance, an r of 0.50 about 25%, and an r of 0.70 roughly half. Even a “large” correlation still leaves much of the variation unexplained.
Why report effect sizes?
The p -value is not enough
A lower p -value is sometimes interpreted as meaning there is a stronger relationship between two variables. However, statistical significance just means the null hypothesis is unlikely to be true (below 5%).
Therefore, a significant p -value tells us that an intervention works, whereas an effect size tells us how much it works.
Because effect size and statistical significance measure different things, they can point in different directions. A tiny effect can turn out statistically significant if the sample is very large.
A genuinely large effect, on the other hand, can fail to reach significance in a small sample. This is why the American Psychological Association’s Publication Manual now requires researchers to report an effect size alongside every p-value.
Emphasizing effect size promotes a more scientific approach. Unlike significance tests, effect size is independent of sample size.
To compare the results of studies done in different settings
Unlike a p -value, effect sizes can be used to compare results from studies in different settings, which is why they matter for meta-analysis.
Critical Evaluation
Cohen’s small, medium and large bands give researchers a shared yardstick for comparing results across studies. But Cohen himself warned these were only “operational definitions”: a rough guide for planning research, not fixed rules.
He said the labels are relative to each other and to the field in which they are used.
Contemporary Research
Funder and Ozer (2019) reviewed the evidence behind Cohen’s benchmarks by comparing them against large-sample studies in personality and social psychology. They found that r ≈ 0.10, labeled “small” by Cohen, is close to the typical effect actually reported in this field.
On their analysis, r ≈ 0.20 compares to the size of many established medical treatments. An r of 0.30, which Cohen calls “medium,” is unusually large outside tightly controlled experiments.
A separate analysis of published personality research found the same pattern. Real correlations typically range from about r = 0.11 to r = 0.29. Cohen’s “medium” band therefore sits at the upper end of what studies actually find (Gignac & Szodorai, 2016).
The practical lesson is that a “small” correlation is not necessarily a small finding. Many researchers now report effect size alongside a confidence interval. They also compare it to the typical effect size in their own field, rather than to Cohen’s original bands alone.
Further Information
- What a p -value Tells You About Statistical Significance
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum.
- Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155–159.
- Ferguson, C. J. (2009). An effect size primer: A guide for clinicians and researchers. Professional Psychology: Research and Practice, 40(5), 532–538.
- Normal Distribution (Bell Curve)
- Z-Score: Definition, Calculation and Interpretation
- Statistics for Psychology
- Statistics for Psychology Book Download