How the design decides what you are allowed to conclude, why measurement matters more than the model, and where quantitative theses come apart.
Quantitative research answers questions about how much, how many, how often, whether groups differ and whether one thing predicts another. The single thing worth understanding is that your design — not your statistics — decides what you may conclude. A cross-sectional survey analysed with the most sophisticated model available still cannot demonstrate causation. Statistics quantify a relationship; the design is what licenses the claim you make about it.
This is the point on which quantitative doctorates most often come unstuck, and it is worth stating in its bluntest form: no analysis can produce a stronger claim than the design supports.
If you measured everything once, at the same time, in one group, you have a cross-sectional design. It can establish that variables are associated. It cannot establish direction, because nothing in the data says which came first, and it cannot rule out that some third thing produced both. Fitting a regression does not change this. The word “predictor” in statistics means a variable on the right-hand side of an equation — it is not a claim about cause, though it reads like one, which is why examiners watch for it.
To claim causation you need a design that earns it: random allocation to conditions, or a natural experiment with a credible comparison, or measurement over time with a defensible argument about what could otherwise explain the change.
The practical consequence is that the strength of your thesis was largely fixed at the design stage. Deciding what claim you want to make, and then choosing the design that can support it, is the order that works. Doing it the other way round produces studies where the conclusion has to be quietly narrowed at the very end.
| Design | What it does | What it supports | Main threat |
|---|---|---|---|
| Descriptive | Measures and reports without comparison | Statements about prevalence or distribution in the sampled population | Sampling — who was left out |
| Correlational | Measures variables and their association | Association, and its strength | Confounding, and reversed direction |
| Cross-sectional survey | One measurement occasion | Association, group comparison at a moment | Any claim involving time |
| Quasi-experimental | Compares conditions without random allocation | Cautious causal claims with a defended comparison | Groups differing before the intervention |
| Experimental | Randomly allocates to conditions | Causal claims within the study conditions | Whether it transfers outside the setting |
| Longitudinal | Repeated measurement over time | Change, order of events, stronger causal argument | Attrition — and who drops out is rarely random |
If you are choosing between two of these, the deciding question is not which is more impressive but which one your conclusion actually requires.
Quantitative research turns concepts into numbers, and that translation is where most of the damage is done. A model fitted to badly measured constructs produces precise numbers about nothing in particular, and no amount of analytical sophistication repairs it.
“Job satisfaction”, “trust”, “engagement” and “resilience” are constructs: they are not directly observable, and they only become measurable through a definition you commit to. That definition is a theoretical claim, not an administrative step, and it decides which items belong. Two studies measuring “engagement” with different instruments may not be studying the same thing at all — which is worth remembering before comparing your results to theirs.
Reliability asks whether the instrument gives consistent results. Validity asks whether it measures what you say it measures. An instrument can be highly reliable and entirely invalid: a scale that consistently weighs six kilos heavy is reliable and wrong.
Reporting an internal-consistency coefficient and stopping there is common and insufficient. Validity is argued through several distinct questions an examiner may raise separately: whether the items cover the whole construct, whether the measure behaves as theory predicts it should, whether it agrees with other measures of the same thing, and whether it stays distinguishable from measures of different things.
Using a published, validated scale is usually the right decision and saves an enormous amount of work. Two cautions. First, validation is population-specific: a scale validated with US undergraduates has not been validated for mid-career professionals in another country, and saying so honestly is stronger than ignoring it. Second, changing items — even wording, even dropping one — means the published validity evidence no longer straightforwardly applies to your version.
Two separate questions get confused here: who is in the sample, and how many.
Who decides whether you can generalise at all. Probability sampling — where everyone in the population has a known chance of selection — is what supports inference to that population. Most doctoral research does not achieve it, and uses convenience or purposive recruitment instead. That is normal and workable, provided the thesis says so plainly and narrows its claims accordingly rather than writing as though a random sample had been obtained.
How many should be argued in advance from three things: the smallest effect that would be meaningful in your field, the probability you want of detecting it if it exists, and the analysis you intend to run. That calculation belongs in the methodology chapter, with its inputs stated so a reader can check them.
Rules of thumb circulate widely — so many cases per predictor, minimum totals for particular models — and they vary considerably between sources. They are rough guidance, not standards, and quoting one as though it settled the matter invites the obvious follow-up question. A calculation tied to your own design is always more defensible.
Where the number was determined by what was achievable, say so. Examiners are used to constraint; they are much less tolerant of a justification reverse-engineered after the fact.
Research that answers questions through measurement and numerical analysis — how much, how many, how often, whether groups differ, whether variables are related. It works with larger samples than qualitative research, aims at findings that extend beyond the individuals studied, and is judged largely on the quality of its measurement and the fit between its design and its claims.
Not on its own. A survey measuring everything at one point in time can establish that variables are associated, but not which came first or whether something else produced both. Statistical models fitted to that data quantify the association more precisely; they do not upgrade it into a causal claim. Causal conclusions need random allocation, a credible natural experiment, or measurement over time with alternative explanations argued away.
Work forwards from three things decided in advance: the smallest effect that would be meaningful in your field, the probability you want of detecting it if it is real, and the specific analysis you intend to run. Free tools such as G*Power handle the arithmetic once those are settled. State the inputs in your methodology chapter so a reader can check them — the reasoning is what is being assessed, not the number.
Reliability is consistency: does the instrument give the same answer under the same conditions? Validity is accuracy of meaning: does it measure the thing you claim it measures? They are independent. A measure can be perfectly consistent and consistently wrong, which is why reporting a reliability coefficient alone does not establish that your variables mean what you say they do.
Use a validated instrument where one exists that fits your construct and population — it saves substantial work and gives you evidence to cite. Writing your own is legitimate when nothing suitable exists, but then developing and validating the instrument becomes part of the doctorate: item generation, expert review, piloting, and evidence about its structure and behaviour. That is a serious undertaking and should be planned for, not discovered halfway through.
Most quantitative problems are fixed cheaply at the design stage and expensively afterwards. Describe your question, your instruments and your intended analysis, and a PhD in your field will tell you whether it supports the claim you want to make.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.