What Does Effect Size Tell You?

Statistical significance is the least interesting thing about the results. You should describe the results in terms of measures of magnitude – not just does treatment affect people, but how much does it affect them.

Effect size is a quantitative measure of the magnitude of the experimental effect. The larger the effect size the stronger the relationship between two variables.

You can look at the effect size when comparing any two groups to see how substantially different they are.

Typically, research studies will comprise an experimental group and a control group. The experimental group may be an intervention or treatment which is expected to affect a specific outcome. The control group gets none.

For example, we might want to know the effect of therapy on treating depression. The effect size value will show whether the therapy has had a small, medium, or large effect on depression.

Key Takeaways

  • What it measures: Effect size shows how large a difference or relationship is. A p-value only tells you whether the result is statistically significant, not how big it is.
  • Two common measures: Cohen’s d compares two group means, such as a treatment group and a control group. Pearson’s r describes how strongly two variables are related.
  • Cohen’s benchmarks: For d, 0.2 counts as small, 0.5 as medium, and 0.8 as large. For r, the same bands sit at roughly 0.1, 0.3 and 0.5.
  • Sample-Size Independence: A small effect can be statistically significant in a large sample. A large effect can also miss significance in a small sample.
  • A modern caveat: Recent research argues Cohen’s benchmarks undervalue small effects, particularly in personality and social psychology, where they can still be practically meaningful.

Calculate and interpret effect sizes

Effect sizes measure either the strength of association between variables or the size of the difference between group means.

Cohen’s d

Cohen’s d is an appropriate effect size for comparing two means, often accompanying the reporting of t-test and ANOVA results. It is also widely used in meta-analysis.

To calculate the standardized mean difference, subtract the mean of one group from the other (M1 – M2). Then divide the result by the standard deviation (SD) of the population the groups were sampled from.

This standardized difference is Cohen’s d.

effect size formula for cohen

A d of 1 means the two groups differ by 1 standard deviation. A d of 2 means they differ by 2, and so on. The scale has no fixed upper bound. Standard deviations are equivalent to z-scores (1 standard deviation = 1 z-score).

Cohen's d effect size illustration

Cohen suggested that d = 0.2 be considered a “small” effect size, 0.5 a “medium” effect size, and 0.8 a “large” effect size. If the difference between two groups’ means is less than 0.2 standard deviations, the difference is negligible, even if it is statistically significant.

Pearson r correlation

This measure of effect size shows how strongly two variables are related. Pearson’s r ranges from -1 (a perfect negative correlation) to +1 (a perfect positive correlation).

Pearson r

According to Cohen (1988, 1992), the effect size is small when r is around 0.1. Medium sits around 0.3, and large above 0.5.

small medium and large effect sizes r

Squaring r gives the coefficient of determination (r²), the percentage of variance the two variables share. An r of 0.30 explains about 9% of the variance, an r of 0.50 about 25%, and an r of 0.70 roughly half. Even a “large” correlation still leaves much of the variation unexplained.

Why square the number? Because r is not on a linear scale: an r of 0.60 is not “twice as strong” as an r of 0.30. Doubling the coefficient more than doubles the variance explained, which is why psychologists compare r² values rather than raw r.

Why report effect sizes?

The p -value is not enough

A lower p -value is sometimes interpreted as meaning there is a stronger relationship between two variables. However, statistical significance just means the null hypothesis is unlikely to be true, using the standard 5% threshold.

Therefore, a significant p -value tells us that an intervention works, whereas an effect size tells us how much it works. The two are not the same.

Because effect size and statistical significance measure different things, they can point in different directions. A correlation as small as r = 0.10, just 1% of the variance, can still turn out statistically significant if the sample is large enough. Sample size explains why.

A genuinely large effect, such as an r of 0.40 (about 16% of the variance), can fail to reach significance in a small sample. This is why the American Psychological Association’s Publication Manual now requires researchers to report an effect size alongside every p-value.

Emphasizing effect size promotes a more scientific approach. Unlike significance tests, effect size is independent of sample size.

Comparing Results Across Studies

Unlike a p -value, effect sizes can be used to compare results from studies in different settings, which is why they matter for meta-analysis.

Effect sizes travel between studies because they are standardized. The scale removes each study’s own units, so an r or d means the same thing whatever the sample size. A p-value carries no such guarantee, because it reflects each study’s sample size as much as the size of the effect.

Critical Evaluation

Cohen’s small, medium and large bands give researchers a shared yardstick for comparing results across studies. But Cohen himself warned these were only “operational definitions”: a rough guide for planning research, not fixed rules.

He said the labels are relative to each other and to the field in which they are used.

Contemporary Research

Funder and Ozer (2019) reviewed the evidence behind Cohen’s benchmarks by comparing them against large-sample studies in personality and social psychology.

Even r ≈ 0.05 can matter.

A tiny per-encounter effect accumulates into something real across many repeated exposures or a whole population. They found that r ≈ 0.10, labeled “small” by Cohen, is close to the typical effect actually reported in this field.

A daily habit shows the idea in action. A per-day link between exercise and mood might sit around r = 0.10. That barely shows up on any single day, but repeated across months and years, it adds up to something real.

On their analysis, r ≈ 0.20 compares to the size of many established medical treatments.

Bigger is not always better.

An r of 0.30, which Cohen calls “medium,” is unusually large outside tightly controlled experiments. A reported r at Cohen’s “large” band (0.50) deserves scrutiny.

It often signals method variance: same-source measurement or a restricted range, rather than a genuinely large effect.

So how big is “big,” really?

A separate analysis of published personality research found the same pattern. Real correlations typically range from about r = 0.11 to r = 0.29, with a median around r = 0.19.

Cohen’s “medium” band therefore sits in the top quarter of what studies actually find (Gignac & Szodorai, 2016).

Correlation (r) Cohen’s (1988) label Funder & Ozer’s (2019) label
≈ 0.05 Below “small” Can matter if it repeats often enough
≈ 0.10 Small Small, but the field’s typical effect
≈ 0.20 Between “small” and “medium” Medium, similar to many medical treatments
≈ 0.30 Medium Large, and rare outside controlled experiments
≥ 0.50 Large Worth scrutiny for method variance

The lesson is simple: “small” does not mean unimportant.

Many researchers now report effect size alongside a confidence interval. They also compare it to the typical effect size in their own field, rather than to Cohen’s original bands alone.

Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.


Saul McLeod, PhD

Chartered Psychologist (CPsychol)

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.