Which sampling logic applies to your study, how to choose a method, and how to defend the sample you could actually get.
Sampling decides what you can claim about anyone you did not study, which makes it one of the few design decisions that cannot be repaired later. Two logics exist: probability sampling, where everyone has a known chance of selection and statistical generalisation becomes possible, and non-probability sampling, where cases are chosen for what they can tell you. Most doctoral research uses the second — and that is fine, provided the claims are adjusted to match.
Everything else follows from this distinction, and confusing the two is where sampling chapters come apart.
Probability sampling means every member of the population has a known, non-zero chance of being selected. That is what licenses statistical inference — the machinery of standard errors, confidence intervals and generalisation to a population assumes it. Achieving it requires a sampling frame: an actual list of the population you are drawing from.
Non-probability sampling means selection is by judgement, availability or referral. It cannot support statistical generalisation, because there is no known relationship between your sample and any wider population. What it can support is depth, access to hard-to-reach groups, and generalisation to theory rather than to people.
The practical position for most doctoral research: you will use non-probability sampling, because a sampling frame for your population does not exist or is not available to you. That is normal and defensible. What is not defensible is using non-probability sampling and then writing as though you had a random sample — describing findings as representative, or generalising to a national population from people recruited through one organisation and a social media post.
Inferential statistics still run on convenience samples, and the software gives no warning. The honest response is to run the analysis, report it, and state precisely which population your claims extend to — which may be the sample itself.
| Method | How it works | Use when |
|---|---|---|
| Simple random | Every member has an equal chance | You have a complete list and the population is fairly homogeneous |
| Systematic | Every nth case from an ordered list | Practical alternative to simple random; watch for periodicity in the list order |
| Stratified | Population divided into groups, sampled within each | You need adequate numbers in subgroups, or want to reduce sampling error |
| Cluster | Whole groups sampled, then all or some members within them | The population is geographically dispersed and listing everyone is impractical |
Two notes that matter for analysis. Stratified sampling with unequal proportions requires weighting if you want to describe the whole population, and reporting unweighted results from a disproportionately stratified sample is an error. Cluster sampling breaks independence — people within a cluster resemble each other, so the analysis has to account for the clustering rather than treating every response as independent. Both are analysis consequences of a sampling decision, which is why the two have to be planned together.
| Method | Selects for | Use when |
|---|---|---|
| Purposive | People meeting criteria your question requires | The default for most qualitative doctoral work |
| Maximum variation | Deliberate spread across relevant differences | You want patterns holding despite variety |
| Homogeneous | Participants closely sharing the experience | Phenomenological work, where shared experience is the point |
| Critical case | A case chosen because it is decisive | “If it does not hold here, it holds nowhere” |
| Theoretical | Whatever the developing theory needs next | Grounded theory, and only there |
| Snowball | Referral through existing participants | Hidden or hard-to-reach populations |
| Quota | Fixed numbers per category, filled non-randomly | You need category coverage without a sampling frame |
| Convenience | Who was available | Rarely defensible alone — state it plainly if used |
Two cautions. Snowball sampling inherits your starting point: people refer people like themselves, so a network-based sample can be far narrower than it appears. Starting from several unconnected points helps, and the limitation should be named. Quota sampling looks like stratified sampling and is not — the categories are filled by whoever is available, so the resulting sample does not support the inference stratification would.
The answer differs entirely by tradition, and the two are frequently conflated.
In quantitative work the number is calculated in advance from three things: the smallest effect that would be meaningful in your field, the power you want, and the specific analysis you intend to run. Rules of thumb about cases per variable circulate widely, disagree with each other, and are much weaker than a calculation tied to your own design. Then add for attrition, missing data and any clustering.
In qualitative work there is no calculation. The number depends on how narrow your question is, how similar participants are to each other, how rich each account is, and which design you are using. Numbers quoted as typical are conventions, not requirements.
On saturation: it means further data collection stopped changing the developing analysis, which presupposes you were analysing as you collected. If all your interviews were done first and coded afterwards, you cannot claim it — and claiming it anyway invites an easy question. Two honest alternatives: state that the sample was determined by what the design and question required, or report that you continued until new material was adding refinement rather than anything new, and show the evidence.
Almost every doctoral sample is a compromise. The difference between a strong and a weak sampling section is not the sample — it is whether the compromise is acknowledged and reasoned about.
What a defensible section does:
The reliable principle: a modest claim argued carefully beats an ambitious claim the sample never supported. Examiners are markedly more forgiving of the first.
In probability sampling every member of the population has a known chance of selection, which is what licenses statistical generalisation — and it requires an actual list of the population to sample from. In non-probability sampling cases are chosen by judgement, availability or referral, which supports depth and access but not statistical generalisation. Most doctoral research uses the second, which is fine as long as the claims are adjusted to match.
It follows from your question and what access you have. Questions about prevalence or population-level differences need probability sampling and a sampling frame. Questions about meaning, process or experience need purposive selection — choosing people who can inform the question. Hard-to-reach populations often require snowball sampling, with its limitation named. If you genuinely only have convenience access, say so plainly rather than dressing it up.
In quantitative work, calculate it in advance from the smallest meaningful effect, the power you want and the analysis you plan, then add for attrition and any clustering. In qualitative work there is no calculation — justify it by what the design and question required, and be straightforward about practical constraints. Either way, explain what determined the number rather than presenting it as given.
You can run the analysis, and the software will not object — but inferential statistics assume a sampling process your recruitment did not follow, so the population your results generalise to is unclear. The honest approach is to run it, report it, and state precisely which population your claims extend to, which may be no wider than the sample itself. Describing convenience-sampled findings as representative is the error to avoid.
It means further data collection stopped changing your developing analysis, which presupposes you were analysing while collecting. If you conducted all interviews first and coded afterwards, you cannot claim it — there was no developing analysis to stop changing. Honest alternatives are to justify the sample by what the design required, or to report that new material was adding refinement rather than anything new, with evidence such as a record of when each code first appeared.
Very likely, provided your claims match it. Validity here is about the fit between sample and conclusion, not sample size in the abstract. State the method and criteria, name the constraint honestly, describe who declined, and draw the boundary of your claim explicitly — what these findings describe and what they do not establish. A modest claim argued carefully is far stronger than an ambitious one the sample never supported.
Sampling is one of the few decisions that cannot be repaired afterwards — only the claim can. Describe who you can realistically reach and a PhD in your field will tell you what it will support.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.