Validity & Reliability

The four kinds of validity examiners distinguish, the threats that undermine each, and why consistency is not accuracy.

The short answer

Reliability is consistency — whether a measure gives the same result under the same conditions. Validity is whether you are measuring and concluding what you claim. They are independent: a measure can be perfectly consistent and consistently wrong, which is why reporting a reliability coefficient does not establish that your variables mean what you say. Validity is not one thing either, and examiners distinguish several kinds that fail in different ways.

Reliability: consistency, and its several forms

Reliability asks whether a measurement is repeatable. It comes in forms that answer different questions, and reporting the wrong one for your design is a small error that signals a larger confusion.

  • Internal consistency — do the items in a scale hang together? The most commonly reported, usually as a coefficient across items. Note that this says nothing about whether the scale measures one thing; a scale capturing two constructs can still look consistent.
  • Test-retest — does the same person score similarly on two occasions? Relevant for stable traits, meaningless for states that genuinely change. The interval matters and should be justified: too short and people remember their answers, too long and real change is mistaken for unreliability.
  • Inter-rater — do two observers or coders agree? Essential wherever judgement is involved in producing the data, and reported with a chance-corrected statistic rather than raw percentage agreement.
  • Parallel forms — do two versions of an instrument produce equivalent results? Relevant where alternate versions exist.

The point worth holding onto: reliability is necessary and never sufficient. A bathroom scale that reads six kilos heavy every time is perfectly reliable. Consistency tells you nothing about accuracy.

The four kinds of validity

“Is it valid?” is four questions, and confusing them is why validity sections often ramble.

TypeThe question it answersEstablished by
ConstructAm I measuring the thing I say I am measuring?Item coverage, theoretical behaviour, agreement with related measures, distinction from unrelated ones
InternalIs the relationship I observed really caused by what I think?Design — randomisation, control, ruling out alternatives
ExternalDoes this hold beyond my sample and setting?Sampling, and the range of contexts studied
Statistical conclusionAre my statistical inferences sound?Adequate power, met assumptions, correct tests, honest handling of multiple comparisons

They trade off against each other, which is worth saying explicitly in a limitations section. A tightly controlled laboratory experiment maximises internal validity and often has poor external validity. A large naturalistic survey has better external validity and cannot establish causation. Recognising the trade-off you made, and why, reads as command of your design.

Threats to internal validity

These are the alternative explanations for your finding, and being able to name the ones relevant to your design is what makes a limitations section convincing rather than apologetic.

  • History — something happened between measurements, unrelated to your study, that explains the change.
  • Maturation — participants changed simply through time passing: growing older, more experienced, more tired.
  • Selection — the groups differed before anything was done to them. The main threat in any non-randomised comparison.
  • Attrition — who dropped out was not random. If the people who struggled left, your remaining sample looks better than the intervention made it.
  • Testing — taking a measure changed people, through practice or through prompting them to think about the topic.
  • Instrumentation — the measure itself changed: a revised questionnaire, an observer who became more skilled, a coder who drifted.
  • Regression to the mean — participants selected for extreme scores will tend to score nearer average next time regardless of any intervention. A genuinely common false positive in studies selecting the worst-performing cases.
  • Diffusion — the control group learned about or received elements of the intervention.

Randomisation addresses many of these at once, which is why experiments carry the causal claims they do. Where randomisation was impossible, the argument has to be made threat by threat: which apply here, and what makes each implausible in your case.

External validity, and the claim you can make

External validity concerns whether findings extend beyond your study — to other people, settings, times or measures.

The main determinant is sampling. Findings from a convenience sample of students in one institution describe those students; extending to a wider population requires either probability sampling or a theoretical argument about why the mechanism should generalise. The second is legitimate and needs stating as an argument rather than assumed.

Other limits worth naming:

  • Setting. Effects found under controlled conditions may not survive real environments, where people are distracted and constrained.
  • Time. Findings tied to a particular moment — a policy period, an economic condition, a public health situation — may not hold later, and saying so is honest rather than weak.
  • Measures. A relationship established with one instrument may not appear with another measuring the same construct differently.

The strongest move is the same one that works elsewhere: state the boundary of the claim explicitly, and say what would be needed to extend it. A modest claim argued carefully survives examination far better than an ambitious one the design never supported.

Writing the section

Validity sections fail in a predictable way: they define the terms, assert that validity was ensured, and move on. Definitions belong in a textbook, and assertion is not evidence.

A section that works does four things:

  1. Names the relevant threats for your specific design. Not all of them — the ones that actually apply.
  2. Says what you did about each. Randomisation, matching, a control group, a validated instrument, blinding, pre-testing.
  3. Reports the evidence. Reliability coefficients for your own sample rather than quoted from the original validation study, and validity evidence for your measures.
  4. States what remains unaddressed, and what that means for the claims.

Two frequent errors worth avoiding. Quoting the reliability from the paper that developed an instrument is not evidence about your data — reliability is a property of a dataset, not of a questionnaire. And using validity and reliability language for qualitative work is a mismatch: that tradition has its own criteria, and applying the wrong vocabulary signals confusion about both.

Questions researchers ask

What is the difference between validity and reliability?

Reliability is consistency — whether a measure gives the same result under the same conditions. Validity is whether you are measuring and concluding what you claim to be. They are independent, and a measure can be perfectly reliable and completely invalid: a scale that reads six kilos heavy every time is consistent and wrong. Reporting a reliability coefficient therefore does not establish that your variables mean what you say they do.

What are the different types of validity?

Construct validity asks whether you are measuring the thing you name. Internal validity asks whether the relationship you observed is really caused by what you think. External validity asks whether it holds beyond your sample and setting. Statistical conclusion validity asks whether your statistical inferences are sound. They trade off against each other — tight experimental control buys internal validity at the cost of external, and naming that trade-off reads as command of your design.

What are threats to internal validity?

Alternative explanations for your finding: history, maturation, selection differences between groups, non-random attrition, effects of being tested, changes in the instrument or observer, regression to the mean, and diffusion between conditions. Randomisation addresses many at once, which is why experiments support causal claims. Without it, you have to argue threat by threat — which apply, and why each is implausible in your case.

Can I quote the reliability from the original validation study?

You can cite it as background, but it is not evidence about your data. Reliability is a property of a dataset rather than of a questionnaire, so it has to be calculated and reported for your own sample. This matters particularly when an instrument validated in one population is used in another, where it may behave quite differently.

Do validity and reliability apply to qualitative research?

Not in this form. Qualitative work is judged by its own criteria — credibility, transferability, dependability and confirmability — which address the same underlying concern on terms suited to interpretive research. Applying reliability language to qualitative work assumes repetition should produce identical results, which does not hold where the researcher is the instrument, and using the wrong vocabulary signals confusion about both traditions.

How do I write the validity and reliability section?

Name the threats that actually apply to your design rather than defining all of them, say what you did about each, report the evidence from your own data, and state honestly what remains unaddressed and what it means for your claims. A section that defines the terms and asserts validity was ensured convinces nobody — the test is whether a reader could tell from it what you specifically did.

Related guides

Check which threats actually apply

A limitations section naming the right threats and answering them reads as command; one listing all of them reads as a textbook. Describe your design and a PhD in your field will tell you which matter.

Discuss your design