AMOS For SEM

The two-step sequence, what the fit indices do and do not settle, and why modification indices are the fastest route to an indefensible model.

The short answer

AMOS runs covariance-based structural equation modelling, estimating a model by trying to reproduce the observed covariances among your variables. It suits theory testing, produces the global fit indices used to judge whether a model is consistent with the data, and generally wants larger samples than variance-based alternatives. Its main hazard is a feature rather than a limitation: modification indices will always suggest changes that improve fit, and following them without a theoretical reason turns confirmation into undisclosed exploration.

The two-step sequence

Established practice is to establish the measurement model before estimating the structural model, and doing it the other way round is what produces results nobody can interpret.

Step one: confirmatory factor analysis. Specify which items load on which latent constructs and test whether that measurement structure fits. Assess loadings, reliability, convergent validity and discriminant validity here, before any structural path is estimated.

Step two: the structural model. Only once measurement holds do you estimate the relationships between constructs that your theory predicts.

The reason for the order is diagnostic. Estimate everything at once, get poor fit, and you cannot tell whether your theory is wrong or your measures are — and those require completely different responses. Separating the steps removes that ambiguity, and it is what an examiner will expect to see described.

A note on what you are drawing: AMOS uses a graphical interface where the model is built as a path diagram. That makes specification unusually visible, which is helpful, but it also makes it easy to draw a model that is not identified — one where there is not enough information in the data to estimate every parameter. If estimation fails or produces implausible values, identification is the first thing to check, before anything else.

Reporting fit honestly

Several indices are reported together because each is sensitive to different things, and no single one settles the question.

The chi-square test is reported and rarely decisive: it is sensitive to sample size and will usually be significant in reasonably large samples even for good models. Comparative indices such as CFI and TLI judge your model against a baseline. RMSEA penalises complexity and should be reported with its confidence interval. SRMR summarises the difference between observed and predicted correlations.

On the thresholds everyone quotes: they derive from simulation work, their authors did not present them as universal rules, and later research has shown they behave differently depending on sample size, model complexity and estimation method. State the criteria you are applying, cite their source, and treat borderline fit as something to discuss rather than to disguise.

Three reporting habits matter here:

  • Report all the indices you examined, not the subset that looked best. Selective reporting is difficult to defend if noticed, and it usually is.
  • Good fit does not mean the model is correct. It means the model is consistent with the data — and other models, including ones with arrows pointing the opposite way, may fit equally well. Saying this yourself is stronger than being told.
  • Report the estimator and how missing data was handled. Full information maximum likelihood is common and should be named if used.

Why modification indices are a trap

This deserves the most attention, because the software actively invites the mistake.

When fit is poor, AMOS will list modifications that would improve it — typically correlating error terms, or adding paths. Applying enough of them will make almost any model fit. The problem is what that fit then means.

Each modification is derived from the specific quirks of your sample. A model modified until it fits is a model shaped by your data rather than tested against it, so the resulting fit statistics no longer indicate that a theory was supported — they indicate that a model was adjusted until it matched. That is exploration, and it is legitimate work, but it has to be labelled as such.

Correlated error terms are the most common case. They are sometimes justifiable: items with near-identical wording, or a shared method such as reverse-scoring, can plausibly share variance for reasons unrelated to the construct. That is a substantive argument, made in advance or at least articulated explicitly. “The modification index was large” is not an argument.

If you do modify:

  • Only make changes you can justify theoretically, and state the justification.
  • Make them one at a time, re-estimating between each, since each modification changes everything else.
  • Report the original model and the modified one, and describe the modifications as post hoc.
  • Say that the modified model requires confirmation on independent data, because it does.

A candidate who reports a model with mediocre fit and discusses why is in a far stronger position than one presenting excellent fit achieved through undisclosed modification.

Mediation in AMOS

The current approach is to test the indirect effect directly using bootstrapping, which AMOS provides. Request bootstrapped confidence intervals for indirect effects, report the number of resamples, and interpret an interval excluding zero as evidence of an indirect effect.

This has superseded the older causal-steps procedure that many supervisors still teach, for two reasons: requiring a significant total effect first is unnecessary, since mediation can exist where opposing indirect paths cancel out, and testing a sequence of conditions does not test the indirect effect itself, which is what the hypothesis concerns.

Two further points. With more than one mediator, request specific indirect effects rather than only the total — otherwise you know something is mediating without knowing which path carries it. And the limitation stands regardless of software: on cross-sectional data the causal ordering is theoretically specified, not empirically demonstrated, and stating that yourself is much better than being led to it.

Problems that stop estimation

AMOS will sometimes refuse to produce a solution or return values that cannot be right. The usual causes, in the order worth checking:

  • Identification. Not enough information to estimate every parameter, usually from too few indicators per construct or a structure with no fixed reference point for a latent variable’s scale.
  • Heywood cases — negative error variances, or standardised loadings above one. These are impossible values and signal a real problem: a small sample, a misspecified model, an outlier, or a construct with too few indicators. Constraining the offending variance to zero makes the message disappear without addressing the cause, and is not a fix.
  • Missing data handling. Some options are unavailable with incomplete data, which is one reason full information maximum likelihood is commonly used.
  • Multicollinearity between constructs, which destabilises estimates and often shows up as implausible path coefficients.
  • Sample size. Covariance-based estimation with a complex model and a modest sample produces unstable results even when it does converge. Rules of thumb here vary widely between sources, so cite whichever guidance you follow rather than treating a number as settled.

An error message is information about your model, not an obstacle to be silenced. The temptation to apply whatever change makes it disappear is exactly how an uninterpretable model reaches a thesis.

Questions researchers ask

What is the two-step approach in AMOS?

Establishing the measurement model before estimating the structural model. First run a confirmatory factor analysis testing whether your items load on their intended constructs, assessing loadings, reliability and validity. Only when that holds do you estimate the paths between constructs. Doing both at once means that if fit is poor you cannot tell whether the theory or the measurement is at fault, and those need entirely different responses.

Should I use modification indices to improve my model fit?

Only where you can justify the change theoretically, and always reported as post hoc. Modification indices are derived from quirks of your particular sample, so a model adjusted until it fits has been shaped by the data rather than tested against it — the fit statistics then show that a model was adjusted, not that a theory was supported. Correlated errors can be defensible for items with near-identical wording or a shared method, but the justification has to be substantive rather than statistical.

What fit indices should I report in AMOS?

Report several, since each is sensitive to different things: chi-square, a comparative index such as CFI and TLI, RMSEA with its confidence interval, and SRMR. Report all of those you examined rather than the ones that looked best. State the criteria you are applying and cite them, since the widely quoted cut-offs came from simulation work and behave differently by sample size, model complexity and estimation method.

Does good model fit mean my theory is correct?

No — it means your model is consistent with the data. Other models can fit the same data equally well, including ones with paths pointing in the opposite direction, which is why fit alone never establishes causal direction. Acknowledging this yourself, and noting that equivalent models exist, is considerably stronger than having it raised in the viva.

How do I test mediation in AMOS?

Request bootstrapped confidence intervals for the indirect effect, report the number of resamples used, and treat an interval excluding zero as evidence of an indirect effect. With more than one mediator, request the specific indirect effects rather than only the total, or you will know that mediation is occurring without knowing which path carries it. This approach has superseded the older causal-steps procedure.

What is a Heywood case and how do I fix it?

A negative error variance or a standardised loading above one — values that are impossible and therefore signal a real problem rather than a display quirk. Usual causes are a small sample, a misspecified model, an influential outlier, or a construct with too few indicators. Constraining the offending variance to zero removes the warning without addressing any of those, so it is not a fix. Diagnose the cause instead.

Related guides

Before you start modifying to reach fit

A model adjusted until it fits is difficult to defend, and the alternatives are easier than they look. Send your model and output, and a PhD in your field will tell you what the fit is actually indicating.

Discuss your model