Effect Size & Confidence Intervals

Which effect size for which test, how to interpret one honestly, and the precise wording of a confidence interval that examiners listen for.

The short answer

An effect size says how large a difference or relationship is, in a form that does not depend on your sample size. A confidence interval says how precisely you have estimated it. Together they answer the question a p-value cannot: does this matter, and how sure can we be? Most journals and examiners now expect both as standard, and a result reported with significance alone reads as dated.

Why significance is not enough

A p-value confounds two different things: how big the effect is, and how many people you measured. That single fact explains most of the trouble it causes.

With a large enough sample, a difference far too small to matter to anyone will reach significance. With a small sample, an effect large enough to change practice may not. So “significant” and “important” are separate questions, and a results chapter that only answers the first has left the substantive question untouched.

Effect sizes separate the two. Because they are expressed independently of sample size, they can be compared across studies — which is also why meta-analysis is possible at all. Confidence intervals then add the precision of your estimate, which is what turns a point estimate into something a reader can reason about.

Which effect size for which analysis

AnalysisEffect sizeWhat it expresses
T-test, two groupsCohen’s d, or Hedges’ gThe difference in standard deviation units
ANOVAEta squared, partial eta squared, omega squaredProportion of variance associated with the factor
Correlationr, and r squaredStrength of association; shared variance
RegressionR squared, standardised coefficients, Cohen’s f squaredVariance explained; contribution of each predictor
Logistic regressionOdds ratioMultiplicative change in the odds of the outcome
Chi-squareCramér’s V, or phi for a two-by-two tableStrength of association between categorical variables
Non-parametric comparisonsRank-biserial correlation, or r derived from the test statisticStrength of the difference in ranks

Three details worth getting right.

Hedges’ g is a version of Cohen’s d corrected for the upward bias that appears in small samples. With modest group sizes it is the better choice, and reporting it signals you know why.

Partial eta squared and eta squared are not interchangeable. Partial eta squared removes other factors from the denominator, so values from designs with different numbers of factors are not comparable — and partial values in a multi-factor design can sum to more than one, which surprises people. Say which you are reporting. Omega squared is less biased than either and is worth considering.

Odds ratios are not risk ratios. They diverge as the outcome becomes more common, so describing an odds ratio as “twice as likely” is only defensible when the outcome is rare.

Interpreting an effect size honestly

The widely quoted labels — small, medium, large — are conventions, and their author was explicit that they were a fallback for situations where no better benchmark existed. Treating them as fixed standards is exactly the misuse he warned against.

The reason matters: what counts as a large effect differs enormously between fields, and between outcomes within a field. A small effect on mortality can be far more important than a large effect on a laboratory task. An intervention producing a small improvement across an entire population may matter more than a large improvement in a handful of people.

So interpret against something real:

  • Benchmarks from your own field, cited. Many disciplines have published guidance on typical effect magnitudes, and using it is far stronger than generic labels.
  • Comparable studies with similar measures and populations.
  • Practical consequence. What does an effect of this size mean in the units people actually care about — days, points, cases, pounds? Converting back to original units is often the most persuasive move available.

Where you do fall back on conventional labels, say that is what you are doing and cite the source. A single sentence acknowledging that the benchmarks are general conventions rather than field-specific standards will forestall the obvious question.

Confidence intervals, worded correctly

This is a point on which examiners in quantitative disciplines listen closely to your phrasing, so it is worth having the precise version ready.

A 95% confidence interval is a statement about the procedure: intervals constructed this way would contain the true population value in 95% of repeated samples. It is not a 95% probability that the true value lies inside this particular interval. That distinction is subtle, it is examinable, and the loose version is what most people say by default.

Safe phrasing in a thesis: the 95% confidence interval was 0.12 to 0.48, followed by what that range means substantively — that the data are consistent with effects across that range, and what the smallest of them would mean in practice.

What intervals are genuinely useful for:

  • Judging precision. A wide interval says your estimate is imprecise, whatever the p-value indicates.
  • Interpreting null results. If the interval excludes all effects large enough to matter, that is evidence the effect is small. If it spans from trivial to substantial, the study simply could not resolve the question. This is what to report instead of post hoc power.
  • Comparing studies. Overlapping intervals across studies say more about consistency than comparing two p-values, which is not a meaningful comparison.

Report intervals around effect sizes, not only around means. Most modern software provides them, and they are increasingly expected.

Reporting them together

A complete report of a quantitative result contains four things: the test and its statistic, the exact p-value, the effect size, and an interval. Each answers a different question, and dropping any one leaves a gap a reader has to fill by guessing.

Two habits improve the chapter noticeably. First, interpret in the text rather than leaving the numbers to speak: a sentence saying what an effect of this size means in practical terms is worth more than another decimal place. Second, be consistent — use the same effect size family throughout for the same kind of analysis, so results can be compared across your own chapters.

Follow your discipline’s reporting standard for the exact format, including how many decimals and whether to report exact p-values throughout. Deviating from it is noticed, and conforming costs nothing.

Questions researchers ask

What is an effect size and why does it matter?

It expresses how large a difference or relationship is, independently of sample size — which is exactly what a p-value cannot do, because significance depends on both the size of the effect and the number of participants. Effect sizes let you say whether a result matters as well as whether it is unlikely under the null hypothesis, and because they are sample-size independent they can be compared across studies.

Which effect size should I report?

It follows from the analysis: Cohen’s d or Hedges’ g for t-tests, eta squared or omega squared for ANOVA, r and r squared for correlation, R squared and standardised coefficients for regression, odds ratios for logistic regression, and Cramér’s V for chi-square. With small samples prefer Hedges’ g, which corrects the upward bias in d. Always state which variant you are reporting, since partial eta squared and eta squared are not comparable.

Are Cohen’s small, medium and large benchmarks reliable?

They are conventions, and their author explicitly presented them as a last resort for fields with no better information. What counts as large varies enormously by discipline and outcome — a small effect on mortality can matter more than a large effect on a laboratory task. Interpret against benchmarks from your own field where they exist, against comparable studies, or by converting the effect back into units people care about. If you do use the generic labels, say so and cite them.

What does a 95% confidence interval actually mean?

It is a statement about the procedure: intervals built this way would contain the true population value in 95% of repeated samples. It is not a 95% probability that the true value lies within this particular interval, and that distinction is genuinely examinable in quantitative disciplines. In writing, report the interval and then interpret the range substantively rather than attaching a probability statement to it.

How do confidence intervals help with non-significant results?

They tell you what the study could and could not rule out, which a p-value does not. If the interval excludes every effect large enough to matter in your field, you have evidence that any real effect is small — a genuine finding. If it stretches from trivial to substantial, your study was unable to resolve the question, and saying so is the honest conclusion. This is what to report instead of post hoc power, which adds nothing.

Do I need to report effect sizes for every test?

For every substantive result, yes — most journals and examiners now expect it, and a bare significance test reads as dated. Keep the same effect size family across comparable analyses so your own results can be compared with each other, and report an interval alongside where your software provides one. Follow your discipline’s reporting standard for the exact format.

Related guides

Get the results reported to current standards

Missing effect sizes and loosely worded intervals are among the easiest things to fix before submission and among the most commonly raised afterwards. Send your results and a PhD in your field will go through them.

Discuss your analysis