The two-step sequence, what the fit indices do and do not settle, and why modification indices are the fastest route to an indefensible model.
AMOS runs covariance-based structural equation modelling, estimating a model by trying to reproduce the observed covariances among your variables. It suits theory testing, produces the global fit indices used to judge whether a model is consistent with the data, and generally wants larger samples than variance-based alternatives. Its main hazard is a feature rather than a limitation: modification indices will always suggest changes that improve fit, and following them without a theoretical reason turns confirmation into undisclosed exploration.
Established practice is to establish the measurement model before estimating the structural model, and doing it the other way round is what produces results nobody can interpret.
Step one: confirmatory factor analysis. Specify which items load on which latent constructs and test whether that measurement structure fits. Assess loadings, reliability, convergent validity and discriminant validity here, before any structural path is estimated.
Step two: the structural model. Only once measurement holds do you estimate the relationships between constructs that your theory predicts.
The reason for the order is diagnostic. Estimate everything at once, get poor fit, and you cannot tell whether your theory is wrong or your measures are — and those require completely different responses. Separating the steps removes that ambiguity, and it is what an examiner will expect to see described.
A note on what you are drawing: AMOS uses a graphical interface where the model is built as a path diagram. That makes specification unusually visible, which is helpful, but it also makes it easy to draw a model that is not identified — one where there is not enough information in the data to estimate every parameter. If estimation fails or produces implausible values, identification is the first thing to check, before anything else.
Several indices are reported together because each is sensitive to different things, and no single one settles the question.
The chi-square test is reported and rarely decisive: it is sensitive to sample size and will usually be significant in reasonably large samples even for good models. Comparative indices such as CFI and TLI judge your model against a baseline. RMSEA penalises complexity and should be reported with its confidence interval. SRMR summarises the difference between observed and predicted correlations.
On the thresholds everyone quotes: they derive from simulation work, their authors did not present them as universal rules, and later research has shown they behave differently depending on sample size, model complexity and estimation method. State the criteria you are applying, cite their source, and treat borderline fit as something to discuss rather than to disguise.
Three reporting habits matter here:
This deserves the most attention, because the software actively invites the mistake.
When fit is poor, AMOS will list modifications that would improve it — typically correlating error terms, or adding paths. Applying enough of them will make almost any model fit. The problem is what that fit then means.
Each modification is derived from the specific quirks of your sample. A model modified until it fits is a model shaped by your data rather than tested against it, so the resulting fit statistics no longer indicate that a theory was supported — they indicate that a model was adjusted until it matched. That is exploration, and it is legitimate work, but it has to be labelled as such.
Correlated error terms are the most common case. They are sometimes justifiable: items with near-identical wording, or a shared method such as reverse-scoring, can plausibly share variance for reasons unrelated to the construct. That is a substantive argument, made in advance or at least articulated explicitly. “The modification index was large” is not an argument.
If you do modify:
A candidate who reports a model with mediocre fit and discusses why is in a far stronger position than one presenting excellent fit achieved through undisclosed modification.
The current approach is to test the indirect effect directly using bootstrapping, which AMOS provides. Request bootstrapped confidence intervals for indirect effects, report the number of resamples, and interpret an interval excluding zero as evidence of an indirect effect.
This has superseded the older causal-steps procedure that many supervisors still teach, for two reasons: requiring a significant total effect first is unnecessary, since mediation can exist where opposing indirect paths cancel out, and testing a sequence of conditions does not test the indirect effect itself, which is what the hypothesis concerns.
Two further points. With more than one mediator, request specific indirect effects rather than only the total — otherwise you know something is mediating without knowing which path carries it. And the limitation stands regardless of software: on cross-sectional data the causal ordering is theoretically specified, not empirically demonstrated, and stating that yourself is much better than being led to it.
AMOS will sometimes refuse to produce a solution or return values that cannot be right. The usual causes, in the order worth checking:
An error message is information about your model, not an obstacle to be silenced. The temptation to apply whatever change makes it disappear is exactly how an uninterpretable model reaches a thesis.
Establishing the measurement model before estimating the structural model. First run a confirmatory factor analysis testing whether your items load on their intended constructs, assessing loadings, reliability and validity. Only when that holds do you estimate the paths between constructs. Doing both at once means that if fit is poor you cannot tell whether the theory or the measurement is at fault, and those need entirely different responses.
Only where you can justify the change theoretically, and always reported as post hoc. Modification indices are derived from quirks of your particular sample, so a model adjusted until it fits has been shaped by the data rather than tested against it — the fit statistics then show that a model was adjusted, not that a theory was supported. Correlated errors can be defensible for items with near-identical wording or a shared method, but the justification has to be substantive rather than statistical.
Report several, since each is sensitive to different things: chi-square, a comparative index such as CFI and TLI, RMSEA with its confidence interval, and SRMR. Report all of those you examined rather than the ones that looked best. State the criteria you are applying and cite them, since the widely quoted cut-offs came from simulation work and behave differently by sample size, model complexity and estimation method.
No — it means your model is consistent with the data. Other models can fit the same data equally well, including ones with paths pointing in the opposite direction, which is why fit alone never establishes causal direction. Acknowledging this yourself, and noting that equivalent models exist, is considerably stronger than having it raised in the viva.
Request bootstrapped confidence intervals for the indirect effect, report the number of resamples used, and treat an interval excluding zero as evidence of an indirect effect. With more than one mediator, request the specific indirect effects rather than only the total, or you will know that mediation is occurring without knowing which path carries it. This approach has superseded the older causal-steps procedure.
A negative error variance or a standardised loading above one — values that are impossible and therefore signal a real problem rather than a display quirk. Usual causes are a small sample, a misspecified model, an influential outlier, or a construct with too few indicators. Constraining the offending variance to zero removes the warning without addressing any of those, so it is not a fix. Diagnose the cause instead.
A model adjusted until it fits is difficult to defend, and the alternatives are easier than they look. Send your model and output, and a PhD in your field will tell you what the fit is actually indicating.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.