The six phases as they are actually meant to be used, why themes do not emerge, and the difference between the versions examiners expect you to name.
Thematic analysis is a method for identifying and interpreting patterns of meaning across a qualitative dataset. It is the most widely used qualitative analysis method in doctoral research, and the most widely misreported. Two things separate a strong thematic analysis from a weak one: themes that make a claim rather than name a topic, and naming which version of thematic analysis you followed, because the main versions differ in what they expect of you.
“Thematic analysis was used” is not a sufficient description of a method, and it is the sentence most likely to attract a question. Several distinct approaches share the name, and they carry different expectations.
Reflexive thematic analysis, developed by Braun and Clarke, treats the researcher’s interpretive work as central rather than as a source of bias to be controlled. Themes are actively constructed by the analyst from coded data, not discovered lying in it. Coding is organic and can evolve throughout; there is no codebook fixed in advance, and agreement between coders is not the measure of quality.
Coding reliability approaches treat coding more like measurement. A codebook is developed early, multiple coders apply it, and agreement between them is calculated and reported. Themes are often determined before the main coding begins.
Codebook approaches sit between the two: a structured codebook, often mapped to research questions, but with interpretive analysis rather than reliability statistics as the goal. Framework analysis is the best-known example.
These are not interchangeable, and mixing their machinery causes trouble. Reporting an inter-coder agreement coefficient inside a reflexive thematic analysis, for instance, is a contradiction: it imports an assumption that there is one correct coding to converge on, which reflexive TA explicitly rejects. Braun and Clarke have written about this misapplication directly, so it is a well-documented objection rather than a matter of taste.
Choose one, cite the specific paper or book, and follow its expectations.
Braun and Clarke set out six phases. They are explicit that these are recursive rather than a linear sequence — you move back and forth between them, and a thesis describing them as six tidy steps completed in order is describing something that did not happen.
Read everything, more than once, before coding anything. If you transcribed the interviews yourself, that counts as part of this. Write notes as you go — not codes, but observations about what strikes you, what surprises you, what seems to be being avoided.
This phase is routinely skipped under time pressure and it is where the analysis actually begins. Coding without knowing the dataset produces codes that fit the first three transcripts and nothing after them.
Work systematically through the data, labelling anything relevant to your question. Code generously at this stage: it is easier to discard a code later than to notice its absence. Give equal attention to each item rather than coding the articulate participants heavily and the terse ones lightly.
Codes can be semantic — capturing what was explicitly said — or latent, capturing assumptions underlying it. Most doctoral analyses use both; saying which you were doing, and where, is part of describing your method.
Now work on the codes rather than the transcripts. Look for clusters that share an underlying idea, and start assembling candidate themes. This is where the analysis stops being organisation and starts being interpretation.
Watch the vocabulary here. Braun and Clarke are pointed about the phrase “themes emerged”, because it implies themes were sitting in the data waiting to be found, and hides the analytic decisions you made. You constructed them. Write it that way.
Test each candidate theme twice: against its own coded extracts, and against the whole dataset. Does the extract set actually cohere? Does the theme hold up when you read the transcripts again with it in mind?
Themes commonly collapse at this stage — two turn out to be one, or one turns out to be two, or a favourite disappears because it rested on a single vivid quotation. That is the phase working, not failing.
Write a short definition of each theme: what it is about, what it is not, and what it contributes to answering your question. If you cannot write that paragraph without listing what people said, the theme is still a category.
Names should carry the claim. “Trust” tells a reader nothing. “Trust is extended to individuals, never to the system” tells them what you found.
The report is part of the analysis, not a record of it. Each theme needs an argument, evidence in the form of extracts chosen to be representative rather than merely quotable, and an explicit link back to the research question.
Include what did not fit. A theme that survived a deliberate search for contradicting cases, with that search described, is far harder to challenge than one where everything agreed.
An inductive analysis builds codes and themes from the data without fitting them to a pre-existing frame. A deductive analysis approaches the data with a theoretical structure and codes in relation to it.
Pure induction is largely a fiction — you bring a research question, a literature and a discipline to the data, and pretending otherwise is not credible. Most doctoral work is genuinely a mixture, and describing that honestly is stronger than claiming a purity examiners will doubt. What matters is being specific: which codes came from the data, which from theory, and what you did when the data would not fit the frame.
Familiarisation with the data; generating initial codes; generating themes from those codes; reviewing themes against the extracts and the whole dataset; defining and naming them; and writing up. Braun and Clarke, who set them out, stress that the phases are recursive rather than linear — you move back and forth, and themes frequently collapse or split during review. A thesis that presents them as six steps completed once in order is describing something that did not happen.
A code labels a segment of data as being about something. A theme is a pattern of shared meaning organised around a central idea, and it makes a claim. “Workload” is a code; “staff absorb unmanageable workload privately because admitting it reads as incompetence” is a theme. A theme is not simply a group of codes with a heading on it.
Better not to. The phrase implies themes were lying in the data waiting to be found, which conceals the interpretive decisions you actually made, and Braun and Clarke have criticised it directly. Since most theses cite them, using the phrase invites an easy question. Write instead that themes were constructed, developed or generated through analysis, and describe how.
It depends which version you are using. In coding reliability approaches, a second coder and a reported agreement statistic are appropriate, because coding is treated as measurement. In reflexive thematic analysis they are not: the approach holds that there is no single correct coding to converge on, so an agreement coefficient contradicts its own assumptions. Discussing coding with a colleague to develop and challenge your thinking is valuable in either case — describe it as that, rather than presenting it as a reliability check.
There is no rule, though a long list usually means codes have been renamed rather than analysed, and two or three very broad themes often means nothing specific is being claimed. Most doctoral chapters land on a small handful, sometimes with sub-themes. The real tests are whether each theme is distinct, whether it holds across the dataset rather than resting on one participant, and whether it does work in answering the research question.
Thematic analysis is interpretive and looks for patterns of meaning; it does not require counting and generally avoids it. Content analysis is more systematic and replicable, works from a coding frame, and often reports frequencies. If the presence or prevalence of categories is genuinely part of your finding, content analysis fits. If the question is what something means, thematic analysis fits.
The gap between competent coding and themes that make a claim is where most qualitative chapters lose marks. Describe your data and where the analysis has reached, and a PhD in your field will work through it with you.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.