Structural Equation Modelling

Measurement model before structural model, mediation done the current way, and the causal claim cross-sectional SEM cannot support.

The short answer

Structural equation modelling estimates a system of relationships at once, and can do so between latent constructs measured by multiple items rather than between single observed variables. Its advantage is that it accounts for measurement error and tests a whole theoretical model together. Its limitation is one people constantly overlook: fitting a model with arrows in it does not demonstrate that the arrows point the right way. On cross-sectional data, SEM tests whether a pattern of associations is consistent with your model — not that your model is true.

Two models, and the order matters

A full SEM has two parts, and confusing them is the source of many uninterpretable results.

The measurement model links observed items to the latent constructs they are taken to reflect. It is a confirmatory factor analysis: are these items measuring what you say they measure?

The structural model specifies the relationships between the constructs — the arrows your theory predicts.

The established practice is a two-step approach: establish that the measurement model fits before estimating the structural model. The reason is practical. If the whole thing is estimated at once and fit is poor, you cannot tell whether your theory is wrong or your measures are, and those call for entirely different responses. Sorting the measurement out first removes that ambiguity.

Where all your variables are directly observed rather than latent — single measured indicators throughout — what you are running is path analysis. It is a legitimate technique, but it does not account for measurement error, and calling it SEM overstates what was done.

Covariance-based or PLS

Two estimation approaches exist, they have different purposes, and choosing one because your department’s software supports it is not a justification.

Covariance-based SEM — the approach implemented in AMOS, Mplus and lavaan — tries to reproduce the observed covariance matrix and gives you the global fit indices used to judge whether the model is consistent with the data. It suits theory testing where an established model is being confirmed or compared against alternatives, and it generally wants larger samples and reasonably well-behaved data.

Partial least squares SEM — the approach in SmartPLS — is variance-based and oriented towards prediction and explaining variance in target constructs. It is more tolerant of smaller samples and non-normal data, and it handles formatively measured constructs, which covariance-based SEM struggles with. It does not produce equivalent global fit statistics, so it is assessed differently.

There is genuine and ongoing methodological debate about when PLS is appropriate, and it has attracted substantive criticism as well as strong advocacy. That debate is a reason to justify your choice explicitly with citations, rather than a reason to avoid either method. What examiners react badly to is a choice with no stated rationale — particularly PLS chosen solely because the sample was small, which is a weak argument on its own.

Mediation, done the current way

Mediation asks whether the relationship between two variables operates through a third. It is among the most common analyses in doctoral research and among the most commonly done in an outdated way.

What has changed

The classic causal-steps procedure — establish a total effect, then a path to the mediator, then the mediator to the outcome, then check whether the direct effect shrinks — is still widely taught and has been superseded in the methodological literature.

Two criticisms drove the change. The requirement of a significant total effect first is now understood to be unnecessary: mediation can be present when the total effect is not significant, particularly where two indirect paths work in opposite directions and cancel out. And the approach tests a series of conditions rather than testing the indirect effect itself, which is the quantity the hypothesis is actually about.

Current practice tests the indirect effect directly, with a bootstrapped confidence interval. Bootstrapping resamples your data many times to build an empirical distribution for the indirect effect, which matters because that quantity is a product of coefficients and is typically not normally distributed — the reason older tests assuming normality performed poorly. If the interval excludes zero, you have evidence of an indirect effect.

The vocabulary has moved too. Rather than “full” and “partial” mediation, which depend heavily on statistical power, it is more defensible to report the size of the indirect effect with its interval and describe the pattern.

The limitation to state honestly

Mediation assumes a causal sequence: X affects M, which affects Y. Cross-sectional data cannot establish that ordering. Everything was measured at once, and the same associations would be produced by several different orderings of the same three variables.

This is the objection examiners raise most reliably, and the answer is not to avoid the analysis but to be straightforward about it. State that the model is theoretically specified rather than empirically demonstrated, acknowledge that the design cannot rule out alternative orderings, and note what would be needed — longitudinal measurement with the mediator measured between predictor and outcome, or experimental manipulation. Candidates who raise this themselves fare considerably better than those led to it.

Moderation

Moderation asks whether the strength or direction of a relationship depends on a third variable — not how an effect happens, which is mediation, but when or for whom it happens.

It is tested by adding an interaction term, the product of predictor and moderator, to a model containing both. A significant interaction indicates moderation.

Practical points that decide whether the analysis is defensible:

  • Centre or standardise continuous variables before creating the product. This reduces the multicollinearity that otherwise arises between the product and its components, and makes the lower-order coefficients interpretable at a meaningful value rather than at a zero that may not exist in your data.
  • Both constituent variables stay in the model. An interaction term without its main effects is not interpretable.
  • A significant interaction is the beginning, not the result. Probe it: plot simple slopes at representative values of the moderator, and report the relationship at each. The Johnson–Neyman technique goes further, identifying the range of the moderator over which the effect is significant, which is often more informative than testing at arbitrary points.
  • Interactions need power. They are typically harder to detect than main effects, so a non-significant interaction in a modest sample is weak evidence of absence.

Where mediation and moderation are combined — an indirect effect that varies by a moderator, or a moderated path within a mediation model — the analysis is best specified explicitly as a conditional process model, with the specific effects of interest tested rather than inferred from a set of separate results.

Reporting an SEM

A complete report lets a reader reconstruct what you did and judge whether it holds. Include:

  • The theoretical model and where each hypothesised path comes from, stated before the results.
  • Software and estimator used, and how missing data was handled — full information maximum likelihood is common in covariance-based work and should be named if used.
  • The measurement model first: loadings, reliability, convergent and discriminant validity evidence.
  • Fit indices with the criteria you are applying and citations for them.
  • Structural paths with standardised coefficients, standard errors and significance, and the variance explained in each endogenous construct.
  • For mediation, the indirect effect with a bootstrapped confidence interval and the number of resamples.
  • Any modifications made after seeing the results, labelled as post hoc rather than presented as planned.

Two habits worth keeping. Report models you tested and rejected, not only the one that worked — concealed model-fishing is difficult to defend if it emerges. And remember that a well-fitting model is one plausible account among several; equivalent models with different arrow directions often fit identically, and saying so yourself is stronger than being told.

Questions researchers ask

What is structural equation modelling?

A method for estimating a system of relationships simultaneously, including relationships between latent constructs measured by several items. It combines a measurement model, linking items to constructs, with a structural model specifying the relationships between those constructs. Its advantages are accounting for measurement error and testing a whole theoretical model at once; its limitation is that fitting a model does not demonstrate the direction of its arrows.

What is the difference between path analysis and SEM?

Path analysis models relationships between directly observed variables. Full SEM includes latent constructs measured by multiple items, which allows measurement error to be modelled rather than ignored. Path analysis is a legitimate technique, but describing it as SEM overstates what was done — and an examiner will ask where the latent variables are.

Should I use AMOS or SmartPLS?

It depends on your purpose. Covariance-based SEM in AMOS suits theory testing, gives global fit indices, and generally wants larger samples and reasonably well-behaved data. PLS-SEM in SmartPLS is oriented towards prediction and explained variance, tolerates smaller samples and non-normality, and handles formative constructs. There is genuine methodological debate about when PLS is appropriate, so justify your choice with citations rather than by software availability — and note that a small sample alone is a weak argument for PLS.

How should I test mediation?

Test the indirect effect directly with a bootstrapped confidence interval, reporting the number of resamples. The older causal-steps procedure has been superseded: requiring a significant total effect first is unnecessary, since mediation can exist where opposing indirect paths cancel out, and testing a sequence of conditions does not test the indirect effect itself. Bootstrapping matters because the indirect effect is a product of coefficients and is typically not normally distributed.

Can I test mediation with cross-sectional data?

You can run it, and you should be careful about what you claim. Mediation assumes a causal sequence that data collected at a single point cannot establish — the same associations are consistent with several orderings of the three variables. The defensible approach is to state that the ordering is theoretically specified rather than empirically demonstrated, acknowledge that alternatives cannot be ruled out, and say what would be needed. Raising this yourself is far better than being led to it in the viva.

What is the difference between mediation and moderation?

Mediation explains how an effect occurs: X influences M, which influences Y. Moderation explains when or for whom it occurs: the strength or direction of the X–Y relationship depends on the level of a third variable. Mediation is tested through the indirect effect; moderation through an interaction term, followed by probing the interaction with simple slopes rather than stopping at its significance.

Do I need to centre variables before testing an interaction?

It is standard practice for continuous variables, for two reasons. It reduces the multicollinearity that arises between a product term and the variables composing it, and it makes the lower-order coefficients interpretable at a meaningful value rather than at zero, which is often outside the range of the data. It does not change the interaction coefficient itself, but it makes everything around it easier to interpret.

Related guides

Have the model specified before you collect

SEM problems usually originate in the design — a mediation hypothesis a cross-sectional survey cannot support, or constructs measured too thinly. Describe your model and a PhD in your field will tell you what it will and will not carry.

Discuss your model