How the calculation actually works, why the hard part is deciding what effect would matter, and why post hoc power answers nothing.
Power is the probability that your study will detect an effect if one genuinely exists. A power analysis works out how many participants you need for a reasonable chance of that, given the size of effect you care about. Four quantities are locked together — the significance level, power, effect size and sample size — so fixing any three determines the fourth. The arithmetic is the easy part; deciding what size of effect would actually matter in your field is the work.
A study can go wrong in two directions. It can find an effect that is not real, or it can miss one that is. The first risk is what your significance level controls. The second is what power addresses, and it is the one doctoral projects routinely ignore until it is too late.
An underpowered study is a genuine problem rather than a technicality. It is likely to return a non-significant result whether or not the effect exists, which means a null finding tells you almost nothing — you cannot distinguish “no effect” from “not enough participants to see it”. Worse, when an underpowered study does reach significance, the effect it reports is usually inflated, because only unusually large sample estimates could clear the threshold.
The four quantities are connected: significance level, power, effect size and sample size. Fix any three and the fourth follows. In practice you set the significance level by convention, choose the power you want, decide the smallest effect worth detecting, and solve for the sample size. Conventional power targets circulate widely and differ between fields, so state and cite the level you are aiming for rather than treating any figure as standard.
Everything in a power calculation is straightforward except this input, and it is the one examiners ask about.
You are not guessing what effect you expect to find. You are deciding the smallest effect that would be worth detecting — the point below which a difference would be too trivial to change anything in your field. That is a substantive judgement about your discipline, not a statistical one.
Reasonable ways to arrive at it, roughly in order of strength:
One approach to avoid: running a small pilot and using its observed effect size to power the main study. Effect estimates from small pilots are extremely imprecise, and building a sample size calculation on one gives the appearance of rigour without the substance. Pilots are for testing procedures, instruments and timings — which is genuinely valuable, and a different job.
G*Power is free, widely accepted, and covers most designs a doctoral project will use. R packages and commercial software do the same job; for complex designs such as multilevel models or structural equation models, simulation-based approaches are often the only realistic route.
Whichever tool you use, the analysis has to be tied to the specific test you will run. The sample size for an independent-samples t-test is not the sample size for a factorial ANOVA with an interaction, and an interaction typically needs considerably more participants than a main effect. Running a generic calculation and applying it to a different analysis is a common shortcut that does not survive questioning.
Then add to the number, for reasons that are entirely predictable:
In the methodology chapter, report the test, the effect size you targeted and where it came from, the significance level, the power, the resulting number, and how attrition was accounted for. The reasoning is what is being assessed, not the figure.
This one is worth knowing precisely, because supervisors and reviewers still occasionally request it.
After a non-significant result, it is tempting to compute the power the study had, using the effect size you actually observed. Software will produce a number. That number is not informative, and the reason is straightforward: computed this way, power is a direct mathematical function of the p-value you already have. A non-significant result will always produce low observed power, and a significant result will always produce high observed power. It cannot tell you anything the p-value did not, and it certainly cannot tell you whether your study was well designed.
The methodological literature is clear about this, and the practice has been criticised for decades. If you are asked for it, the constructive response is to offer something that does answer the underlying question: a confidence interval around your observed effect. That shows the range of effects your data are consistent with, and if that interval excludes effects large enough to matter, you have genuine evidence of absence rather than merely absence of evidence. If it includes large effects, your study simply could not resolve the question — which is the honest conclusion.
Power analysis belongs before data collection, where it can change what you do. Afterwards, intervals do the work.
This is a common position in doctoral research, particularly with clinical populations, senior professionals or small organisations. The calculation says four hundred; access realistically gives you ninety.
Pretending otherwise is the worst option. Better ones, depending on circumstances:
Examiners deal with constrained samples constantly. What they respond badly to is a candidate who did not notice, or who noticed and covered it.
Decide three things in advance: your significance level, the power you want, and the smallest effect that would be meaningful in your field. Then run the calculation for the specific test you intend to use — G*Power handles most doctoral designs, and simulation is usually needed for multilevel or structural equation models. Add for expected attrition, missing data and any clustering in your design, since the calculation gives the number you need to analyse rather than to recruit.
The smallest effect worth detecting, not the effect you expect to find. Ideally that comes from what would be practically meaningful in your field — the difference that would change a decision or a practice. Failing that, use meta-analyses or comparable published studies, bearing in mind that published effects tend to run larger than true ones. Generic small, medium and large conventions are a last resort, and were intended as exactly that.
It is best avoided. Effect estimates from small pilots are very imprecise, so a sample size built on one has the appearance of rigour without the substance, and can send you badly wrong in either direction. Pilots are genuinely valuable for testing procedures, instruments, timings and recruitment — use them for that, and take your effect size from published evidence or from what would be practically meaningful.
No. Power computed after the fact from your observed effect size is a direct mathematical function of the p-value, so a non-significant result always yields low observed power and a significant one always yields high power. It adds nothing to what you already know and has been criticised in the literature for decades. If you need to address whether your study could have detected an effect, report a confidence interval around your observed effect instead.
Possibly, and post hoc power will not tell you. Look at the confidence interval around your observed effect: if it excludes effects large enough to matter in your field, you have real evidence that any effect is small, which is a genuine finding. If it includes large effects, your study could not resolve the question, and saying so directly is the honest conclusion. Either way, report the power the study had for the effect you specified before collection.
Say so, and adapt. Options include narrowing the question to one the achievable sample can answer, simplifying the analysis by using fewer predictors or dropping interaction terms, switching to a repeated-measures design where each participant acts as their own control, or reframing the study as estimation — reporting effect sizes with confidence intervals rather than significance tests. Examiners are used to constrained samples; what they respond badly to is the problem going unacknowledged.
A sample size settled in advance is a paragraph in your methodology chapter; one discovered afterwards is a limitation you cannot fix. Describe your design and intended analysis, and a PhD in your field will work it through with you.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.