Which effect size for which test, how to interpret one honestly, and the precise wording of a confidence interval that examiners listen for.
An effect size says how large a difference or relationship is, in a form that does not depend on your sample size. A confidence interval says how precisely you have estimated it. Together they answer the question a p-value cannot: does this matter, and how sure can we be? Most journals and examiners now expect both as standard, and a result reported with significance alone reads as dated.
A p-value confounds two different things: how big the effect is, and how many people you measured. That single fact explains most of the trouble it causes.
With a large enough sample, a difference far too small to matter to anyone will reach significance. With a small sample, an effect large enough to change practice may not. So “significant” and “important” are separate questions, and a results chapter that only answers the first has left the substantive question untouched.
Effect sizes separate the two. Because they are expressed independently of sample size, they can be compared across studies — which is also why meta-analysis is possible at all. Confidence intervals then add the precision of your estimate, which is what turns a point estimate into something a reader can reason about.
| Analysis | Effect size | What it expresses |
|---|---|---|
| T-test, two groups | Cohen’s d, or Hedges’ g | The difference in standard deviation units |
| ANOVA | Eta squared, partial eta squared, omega squared | Proportion of variance associated with the factor |
| Correlation | r, and r squared | Strength of association; shared variance |
| Regression | R squared, standardised coefficients, Cohen’s f squared | Variance explained; contribution of each predictor |
| Logistic regression | Odds ratio | Multiplicative change in the odds of the outcome |
| Chi-square | Cramér’s V, or phi for a two-by-two table | Strength of association between categorical variables |
| Non-parametric comparisons | Rank-biserial correlation, or r derived from the test statistic | Strength of the difference in ranks |
Three details worth getting right.
Hedges’ g is a version of Cohen’s d corrected for the upward bias that appears in small samples. With modest group sizes it is the better choice, and reporting it signals you know why.
Partial eta squared and eta squared are not interchangeable. Partial eta squared removes other factors from the denominator, so values from designs with different numbers of factors are not comparable — and partial values in a multi-factor design can sum to more than one, which surprises people. Say which you are reporting. Omega squared is less biased than either and is worth considering.
Odds ratios are not risk ratios. They diverge as the outcome becomes more common, so describing an odds ratio as “twice as likely” is only defensible when the outcome is rare.
The widely quoted labels — small, medium, large — are conventions, and their author was explicit that they were a fallback for situations where no better benchmark existed. Treating them as fixed standards is exactly the misuse he warned against.
The reason matters: what counts as a large effect differs enormously between fields, and between outcomes within a field. A small effect on mortality can be far more important than a large effect on a laboratory task. An intervention producing a small improvement across an entire population may matter more than a large improvement in a handful of people.
So interpret against something real:
Where you do fall back on conventional labels, say that is what you are doing and cite the source. A single sentence acknowledging that the benchmarks are general conventions rather than field-specific standards will forestall the obvious question.
This is a point on which examiners in quantitative disciplines listen closely to your phrasing, so it is worth having the precise version ready.
A 95% confidence interval is a statement about the procedure: intervals constructed this way would contain the true population value in 95% of repeated samples. It is not a 95% probability that the true value lies inside this particular interval. That distinction is subtle, it is examinable, and the loose version is what most people say by default.
Safe phrasing in a thesis: the 95% confidence interval was 0.12 to 0.48, followed by what that range means substantively — that the data are consistent with effects across that range, and what the smallest of them would mean in practice.
What intervals are genuinely useful for:
Report intervals around effect sizes, not only around means. Most modern software provides them, and they are increasingly expected.
A complete report of a quantitative result contains four things: the test and its statistic, the exact p-value, the effect size, and an interval. Each answers a different question, and dropping any one leaves a gap a reader has to fill by guessing.
Two habits improve the chapter noticeably. First, interpret in the text rather than leaving the numbers to speak: a sentence saying what an effect of this size means in practical terms is worth more than another decimal place. Second, be consistent — use the same effect size family throughout for the same kind of analysis, so results can be compared across your own chapters.
Follow your discipline’s reporting standard for the exact format, including how many decimals and whether to report exact p-values throughout. Deviating from it is noticed, and conforming costs nothing.
It expresses how large a difference or relationship is, independently of sample size — which is exactly what a p-value cannot do, because significance depends on both the size of the effect and the number of participants. Effect sizes let you say whether a result matters as well as whether it is unlikely under the null hypothesis, and because they are sample-size independent they can be compared across studies.
It follows from the analysis: Cohen’s d or Hedges’ g for t-tests, eta squared or omega squared for ANOVA, r and r squared for correlation, R squared and standardised coefficients for regression, odds ratios for logistic regression, and Cramér’s V for chi-square. With small samples prefer Hedges’ g, which corrects the upward bias in d. Always state which variant you are reporting, since partial eta squared and eta squared are not comparable.
They are conventions, and their author explicitly presented them as a last resort for fields with no better information. What counts as large varies enormously by discipline and outcome — a small effect on mortality can matter more than a large effect on a laboratory task. Interpret against benchmarks from your own field where they exist, against comparable studies, or by converting the effect back into units people care about. If you do use the generic labels, say so and cite them.
It is a statement about the procedure: intervals built this way would contain the true population value in 95% of repeated samples. It is not a 95% probability that the true value lies within this particular interval, and that distinction is genuinely examinable in quantitative disciplines. In writing, report the interval and then interpret the range substantively rather than attaching a probability statement to it.
They tell you what the study could and could not rule out, which a p-value does not. If the interval excludes every effect large enough to matter in your field, you have evidence that any real effect is small — a genuine finding. If it stretches from trivial to substantial, your study was unable to resolve the question, and saying so is the honest conclusion. This is what to report instead of post hoc power, which adds nothing.
For every substantive result, yes — most journals and examiners now expect it, and a bare significance test reads as dated. Keep the same effect size family across comparable analyses so your own results can be compared with each other, and report an interval alongside where your software provides one. Follow your discipline’s reporting standard for the exact format.
Missing effect sizes and loosely worded intervals are among the easiest things to fix before submission and among the most commonly raised afterwards. Send your results and a PhD in your field will go through them.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.