Factor Analysis & Reliability

Exploratory or confirmatory, how many factors to retain, and what Cronbach’s alpha does and does not tell you.

The short answer

Factor analysis examines whether a set of measured items reflects a smaller number of underlying constructs. Exploratory factor analysis asks what structure is in the data; confirmatory factor analysis tests a structure you specified in advance from theory. Reliability analysis then asks whether the items grouped together are measuring consistently. Two things are widely misreported here: alpha is not evidence that a scale is unidimensional, and the rule of keeping factors with eigenvalues above one is known to retain too many.

Exploratory or confirmatory

The choice follows from how much you already know, and using the wrong one is a design error rather than a preference.

Exploratory factor analysis is for when you do not have a firm structure in mind: a new instrument, an existing one in a new population or language, or a set of items whose organisation is genuinely an open question. The analysis suggests a structure; it does not test one.

Confirmatory factor analysis tests a structure specified in advance — these items load on this factor and not on others. It belongs to the structural equation modelling family and produces fit statistics telling you how well the specified model reproduces the observed relationships between items.

The error to avoid is running an EFA, adopting whatever structure it produced, and then presenting a CFA of that same structure on the same data as confirmation. Nothing has been confirmed: a model derived from a dataset will fit that dataset. If you want both, split the sample or collect a fresh one.

A note on terminology worth getting right: principal components analysis is not factor analysis, though software often lists them together. PCA is a data reduction technique that summarises total variance in components; factor analysis models shared variance to identify latent constructs presumed to cause the item responses. If your argument is that items reflect an underlying construct, a factor extraction method such as principal axis factoring or maximum likelihood matches that argument; PCA does not.

Running an exploratory factor analysis

Four decisions have to be made and reported. Leaving any of them implicit means the analysis cannot be evaluated.

Is the data suitable?

Two checks are conventionally reported. The Kaiser–Meyer–Olkin measure indicates whether the variables share enough common variance to be worth factoring, with higher values better. Bartlett’s test of sphericity should be significant, indicating the correlation matrix is not essentially an identity matrix. Bartlett’s is a weak hurdle and will usually be significant in any reasonable dataset, so it is a minimum rather than evidence of suitability.

Sample size guidance varies widely between sources, from fixed minimums to ratios of participants per item. Those figures conflict, and what actually matters is the strength of the loadings and how many items load well on each factor — strong, clean structures need fewer cases than weak, cross-loading ones. Cite whichever guidance you follow rather than presenting a number as settled.

How many factors to keep

This is the decision that shapes everything after it, and the default in most software is the weakest option.

The Kaiser criterion — retain factors with eigenvalues above one — is the software default and is well known to over-extract, frequently suggesting more factors than are substantively meaningful. Using it alone is difficult to defend.

The scree plot is a visual judgement about where the curve levels off. Useful, and subjective.

Parallel analysis compares your eigenvalues against those from random data of the same size, retaining only factors that outperform chance. It generally performs better than the alternatives in simulation studies and is increasingly the expected standard.

The strongest approach is to use more than one method, check they agree, and then apply the test that outranks all of them: does the solution make theoretical sense, with factors you can name and defend? A statistically defensible factor nobody can interpret is not a finding.

Which rotation

Rotation makes the structure interpretable by moving items towards loading clearly on one factor.

Orthogonal rotation, such as varimax, forces the factors to be uncorrelated. Oblique rotation, such as promax or oblimin, allows them to correlate.

In most social science the constructs plausibly correlate — dimensions of job satisfaction are not independent of each other — so oblique rotation usually matches reality better, and if the factors turn out to be near-uncorrelated it produces much the same result as orthogonal rotation anyway. Choosing varimax by default because it is first in the menu is common and worth avoiding. If you use oblique rotation, report the factor correlations and interpret the pattern matrix.

Which items to drop

Items are usually removed for loading too weakly on any factor, or for loading substantially on more than one. Thresholds for both circulate widely and vary by source and sample size, so state the cut-offs you applied and cite them rather than treating them as universal.

Two disciplines are worth keeping. Remove items one at a time, re-running between each, because removing one changes the solution. And do not remove an item that is theoretically central purely because it behaves awkwardly — that finding may be telling you something about the construct in your population, which is worth reporting rather than deleting.

Confirmatory factor analysis and fit

CFA tests a specified measurement model and reports how well it reproduces the observed covariances. Several indices are conventionally reported together, because each is sensitive to different things.

The chi-square test is reported but rarely relied on: it is sensitive to sample size and will usually be significant in reasonably large samples even for good models. The comparative fit index and Tucker–Lewis index compare your model against a baseline. The root mean square error of approximation penalises complexity and is usually reported with a confidence interval. The standardised root mean square residual summarises the difference between observed and predicted correlations.

An important caution about the cut-off values everyone quotes. The widely used thresholds derive from simulation work, and their authors did not intend them as universal rules; subsequent research has shown they perform differently depending on sample size, model complexity and estimation method. Report the indices, state the criteria you are applying and cite their source, and treat borderline fit as something to discuss rather than to disguise.

Two habits protect you here. Modification indices will always suggest changes that improve fit — adopting them without a theoretical reason turns confirmation into exploration, and should be reported as such if you do it. And good fit shows your model is consistent with the data; it does not show the model is correct, since other models may fit equally well.

Reliability, and the trouble with alpha

Cronbach’s alpha is the most reported statistic in quantitative theses and among the least well understood. It is worth knowing what it actually does, because examiners in quantitative fields increasingly ask.

Alpha estimates internal consistency: how strongly the items in a scale relate to one another. Three limits follow, all well established in the methods literature.

  • It is not evidence of unidimensionality. A scale measuring two distinct things can return a high alpha. Establishing that a scale is unidimensional requires factor analysis, not alpha.
  • It rises with the number of items. A long scale of mediocre items can produce a comfortable alpha, which is why a high value on a twenty-item scale says less than the same value on four items.
  • It assumes tau-equivalence — that all items contribute equally to the construct. Real scales rarely satisfy this, and where they do not, alpha typically underestimates reliability.

Because of that last point, McDonald’s omega is increasingly recommended, since it does not require equal item contributions and is straightforward to obtain in current software. Reporting omega alongside alpha, or in place of it with a citation, is a small change that reads as current.

Two further points on reporting. Thresholds for acceptable alpha are conventions, differ by field and by whether a measure is used for research or individual decisions, so cite the standard you are applying. And where a validated instrument is used, report reliability in your sample rather than quoting the original study — reliability is a property of the data, not of the questionnaire.

Construct validity alongside reliability

Reliability is consistency. Validity is whether the scale measures what you say it measures, and it is argued rather than proved by a single number.

In measurement work two forms are commonly evidenced. Convergent validity asks whether items intended to reflect the same construct hold together, often supported through loadings and average variance extracted. Discriminant validity asks whether constructs that should be distinguishable actually are.

On discriminant validity the field has moved. The long-standing approach comparing the square root of average variance extracted against inter-construct correlations has been shown to perform poorly at detecting problems in common conditions, and the heterotrait–monotrait ratio is now widely preferred and often expected. Reporting the older criterion alone is increasingly likely to draw a question in fields using structural equation modelling.

Questions researchers ask

What is the difference between exploratory and confirmatory factor analysis?

EFA asks what structure exists in a set of items without specifying one in advance, and is appropriate for new instruments or existing ones in a new population. CFA tests a structure you specified from theory, and reports fit statistics indicating how well it reproduces the observed relationships. Running an EFA and then a CFA of the resulting structure on the same data confirms nothing, since a model derived from a dataset will fit that dataset — you need a split sample or fresh data.

How many factors should I retain?

Not by the eigenvalue-over-one rule alone, which is the software default and is well known to retain too many factors. Parallel analysis, which compares your eigenvalues against those from random data of the same size, performs better and is increasingly expected. Use more than one method, check they agree, and apply the decisive test: can you name and defend each factor theoretically? A statistically defensible factor nobody can interpret is not a finding.

Should I use varimax or oblique rotation?

Oblique rotation, such as promax or oblimin, is usually the better default in social science because it allows factors to correlate — and constructs in these fields generally do. If the factors turn out to be near-independent, oblique rotation produces much the same solution as varimax anyway, so little is lost. Choosing varimax because it appears first in the menu is common and hard to justify when asked.

What does Cronbach’s alpha actually tell me?

That the items in a scale relate to one another, and no more than that. It is not evidence that a scale measures one thing — a scale capturing two distinct constructs can still return a high alpha, and unidimensionality is established through factor analysis. It also rises simply with the number of items, and it assumes all items contribute equally, which real scales rarely satisfy. McDonald’s omega does not carry that assumption and is increasingly recommended.

What is an acceptable Cronbach’s alpha?

The commonly quoted thresholds are conventions rather than standards, and they differ by field and by purpose — a measure used for decisions about individuals is held to a higher bar than one used for group-level research. Cite the source of whatever criterion you apply, report reliability for your own sample rather than quoting the original validation study, and consider reporting omega alongside alpha.

What fit indices should I report for CFA?

Report several, since each is sensitive to different things: the chi-square test, a comparative index such as CFI and TLI, RMSEA with its confidence interval, and SRMR. The widely quoted cut-off values come from simulation work and were not intended as universal rules; later research shows they behave differently by sample size, model complexity and estimation method. State the criteria you are applying, cite them, and discuss borderline fit rather than concealing it.

Related guides

Get the measurement model checked first

Structural results built on a measurement model that does not hold cannot be repaired later. Send your items, loadings and reliability output, and a PhD in your field will work through them with you.

Discuss your analysis