Sample Size & Power Analysis

How the calculation actually works, why the hard part is deciding what effect would matter, and why post hoc power answers nothing.

The short answer

Power is the probability that your study will detect an effect if one genuinely exists. A power analysis works out how many participants you need for a reasonable chance of that, given the size of effect you care about. Four quantities are locked together — the significance level, power, effect size and sample size — so fixing any three determines the fourth. The arithmetic is the easy part; deciding what size of effect would actually matter in your field is the work.

What power actually means

A study can go wrong in two directions. It can find an effect that is not real, or it can miss one that is. The first risk is what your significance level controls. The second is what power addresses, and it is the one doctoral projects routinely ignore until it is too late.

An underpowered study is a genuine problem rather than a technicality. It is likely to return a non-significant result whether or not the effect exists, which means a null finding tells you almost nothing — you cannot distinguish “no effect” from “not enough participants to see it”. Worse, when an underpowered study does reach significance, the effect it reports is usually inflated, because only unusually large sample estimates could clear the threshold.

The four quantities are connected: significance level, power, effect size and sample size. Fix any three and the fourth follows. In practice you set the significance level by convention, choose the power you want, decide the smallest effect worth detecting, and solve for the sample size. Conventional power targets circulate widely and differ between fields, so state and cite the level you are aiming for rather than treating any figure as standard.

The hard part: what effect would matter

Everything in a power calculation is straightforward except this input, and it is the one examiners ask about.

You are not guessing what effect you expect to find. You are deciding the smallest effect that would be worth detecting — the point below which a difference would be too trivial to change anything in your field. That is a substantive judgement about your discipline, not a statistical one.

Reasonable ways to arrive at it, roughly in order of strength:

  • Practical significance. What size of difference would actually change a decision, a policy or a practice? This is the strongest justification because it is grounded in what your research is for.
  • Meta-analyses in your field. The average effect across studies of similar phenomena, cited.
  • Comparable published studies using similar measures in similar populations, with the caveat that published effects tend to be larger than true effects, since larger effects are more likely to be published.
  • Field conventions, cited, as a last resort. Generic small, medium and large labels were offered precisely as a fallback for situations with no better information, and using them where better information exists is weak.

One approach to avoid: running a small pilot and using its observed effect size to power the main study. Effect estimates from small pilots are extremely imprecise, and building a sample size calculation on one gives the appearance of rigour without the substance. Pilots are for testing procedures, instruments and timings — which is genuinely valuable, and a different job.

Running the calculation

G*Power is free, widely accepted, and covers most designs a doctoral project will use. R packages and commercial software do the same job; for complex designs such as multilevel models or structural equation models, simulation-based approaches are often the only realistic route.

Whichever tool you use, the analysis has to be tied to the specific test you will run. The sample size for an independent-samples t-test is not the sample size for a factorial ANOVA with an interaction, and an interaction typically needs considerably more participants than a main effect. Running a generic calculation and applying it to a different analysis is a common shortcut that does not survive questioning.

Then add to the number, for reasons that are entirely predictable:

  • Attrition. Longitudinal and intervention studies lose participants, and the calculation gives you the number you need to analyse, not to recruit. Estimate the loss from comparable studies and recruit above the target accordingly.
  • Missing data. Listwise deletion in a model with several variables can remove more cases than people expect.
  • Clustering. If participants are grouped inside schools, wards or teams, they carry less independent information than their count implies, and the required sample rises — sometimes substantially. This needs building into the calculation rather than noticing afterwards.

In the methodology chapter, report the test, the effect size you targeted and where it came from, the significance level, the power, the resulting number, and how attrition was accounted for. The reasoning is what is being assessed, not the figure.

Why post hoc power tells you nothing

This one is worth knowing precisely, because supervisors and reviewers still occasionally request it.

After a non-significant result, it is tempting to compute the power the study had, using the effect size you actually observed. Software will produce a number. That number is not informative, and the reason is straightforward: computed this way, power is a direct mathematical function of the p-value you already have. A non-significant result will always produce low observed power, and a significant result will always produce high observed power. It cannot tell you anything the p-value did not, and it certainly cannot tell you whether your study was well designed.

The methodological literature is clear about this, and the practice has been criticised for decades. If you are asked for it, the constructive response is to offer something that does answer the underlying question: a confidence interval around your observed effect. That shows the range of effects your data are consistent with, and if that interval excludes effects large enough to matter, you have genuine evidence of absence rather than merely absence of evidence. If it includes large effects, your study simply could not resolve the question — which is the honest conclusion.

Power analysis belongs before data collection, where it can change what you do. Afterwards, intervals do the work.

When you cannot recruit the number you need

This is a common position in doctoral research, particularly with clinical populations, senior professionals or small organisations. The calculation says four hundred; access realistically gives you ninety.

Pretending otherwise is the worst option. Better ones, depending on circumstances:

  • Change the question to one the achievable sample can answer. A study that estimates a relationship precisely in a narrower group beats an underpowered attempt at a broader claim.
  • Simplify the analysis. Fewer predictors, fewer groups and no interaction terms all reduce what you need.
  • Use a repeated-measures design where possible. Each participant acting as their own control removes between-person variation and can markedly reduce the sample required.
  • Reframe the study as estimation rather than testing. Report effect sizes with confidence intervals and be explicit that the study was not powered for hypothesis testing. This is increasingly accepted and is far more defensible than a significance test the design could never support.
  • State it plainly as a limitation, with the power the study did have for the effect you considered meaningful, calculated in advance.

Examiners deal with constrained samples constantly. What they respond badly to is a candidate who did not notice, or who noticed and covered it.

Questions researchers ask

How do I calculate the sample size for my study?

Decide three things in advance: your significance level, the power you want, and the smallest effect that would be meaningful in your field. Then run the calculation for the specific test you intend to use — G*Power handles most doctoral designs, and simulation is usually needed for multilevel or structural equation models. Add for expected attrition, missing data and any clustering in your design, since the calculation gives the number you need to analyse rather than to recruit.

What effect size should I use in a power analysis?

The smallest effect worth detecting, not the effect you expect to find. Ideally that comes from what would be practically meaningful in your field — the difference that would change a decision or a practice. Failing that, use meta-analyses or comparable published studies, bearing in mind that published effects tend to run larger than true ones. Generic small, medium and large conventions are a last resort, and were intended as exactly that.

Can I use a pilot study to estimate my effect size?

It is best avoided. Effect estimates from small pilots are very imprecise, so a sample size built on one has the appearance of rigour without the substance, and can send you badly wrong in either direction. Pilots are genuinely valuable for testing procedures, instruments, timings and recruitment — use them for that, and take your effect size from published evidence or from what would be practically meaningful.

Should I report post hoc power?

No. Power computed after the fact from your observed effect size is a direct mathematical function of the p-value, so a non-significant result always yields low observed power and a significant one always yields high power. It adds nothing to what you already know and has been criticised in the literature for decades. If you need to address whether your study could have detected an effect, report a confidence interval around your observed effect instead.

My results were not significant. Was my study underpowered?

Possibly, and post hoc power will not tell you. Look at the confidence interval around your observed effect: if it excludes effects large enough to matter in your field, you have real evidence that any effect is small, which is a genuine finding. If it includes large effects, your study could not resolve the question, and saying so directly is the honest conclusion. Either way, report the power the study had for the effect you specified before collection.

What if I cannot recruit enough participants?

Say so, and adapt. Options include narrowing the question to one the achievable sample can answer, simplifying the analysis by using fewer predictors or dropping interaction terms, switching to a repeated-measures design where each participant acts as their own control, or reframing the study as estimation — reporting effect sizes with confidence intervals rather than significance tests. Examiners are used to constrained samples; what they respond badly to is the problem going unacknowledged.

Related guides

Run the calculation before you collect, not after

A sample size settled in advance is a paragraph in your methodology chapter; one discovered afterwards is a limitation you cannot fix. Describe your design and intended analysis, and a PhD in your field will work it through with you.

Discuss your design