What to actually do with transcripts: coding, the step from codes to themes, choosing an approach, and what the software will and will not do for you.
Qualitative data analysis turns raw material — transcripts, notes, documents — into an account that answers your question. The work runs from familiarisation, through coding, to the step that decides whether the chapter is any good: raising codes into themes that make a claim. Most weak qualitative chapters have coded competently and stopped there, producing tidy categories of what was mentioned rather than findings about what it means.
Two decisions taken at the start save weeks later, and both are frequently skipped.
Decide what your data actually is. A verbatim transcript, a cleaned transcript and a summary are three different datasets supporting three different levels of claim. If you plan to say anything about how something was said — hesitation, repair, emphasis — that has to survive transcription. Decide the convention, apply it consistently, and state it in the chapter.
Decide whether your coding starts from the data or from theory. Inductive coding builds categories up from what is there. Deductive coding applies a framework decided in advance. Most doctoral work is a mixture, and that is fine, but the mixture has to be described: which codes came from where, and what happened when the data did not fit the framework.
Then read everything before coding anything. Familiarisation is not a formality — it is where you notice what is surprising, and surprise is usually where the finding is.
Coding is labelling segments of material so they can be gathered and compared. The vocabulary varies between traditions, and using terms from one tradition inside another is a small error that reads as confusion.
First-cycle coding stays close to the data: descriptive labels, or the participants’ own words kept as codes where the phrasing itself matters. Second-cycle coding works on the codes rather than the transcripts — merging, splitting, grouping, discarding what turned out not to be doing any work.
A codebook holds each code’s name, its definition, when to apply it, when not to, and an example. Writing definitions carefully is what keeps coding consistent across two hundred pages of material, and it is also what allows anyone else to follow what you did.
Expect the codebook to change as you go. That is normal, but it means earlier material has to be revisited under the revised scheme — a step often skipped, and visible in the finished chapter when early transcripts are coded differently from late ones.
These belong to grounded theory specifically. Open coding breaks data into concepts; axial coding works out how those concepts relate; selective coding organises everything around a core category. Used properly they come with the rest of the grounded-theory apparatus — collecting and analysing in parallel, sampling driven by the developing theory, constant comparison.
Borrowing this vocabulary for a study that did none of that is one of the most common mislabels in doctoral qualitative work. If you coded a fixed set of transcripts thematically, call it thematic analysis; it is a perfectly respectable thing to have done.
This is the single most important point on the page, and the one that separates a chapter that passes comfortably from one that gets picked apart.
A code labels what a piece of data is about. A theme makes a claim about what it means. Collecting codes into groups and giving each group a heading produces categories, not themes — and categories are what most weak chapters present.
The test is simple: read the theme name aloud. If it is a noun or a topic — “Communication”, “Barriers”, “Support” — it is a category. It could have been written before you collected any data, which means it is telling the reader nothing you found out. If it makes a claim — “Formal channels are treated as unusable, so decisions move through informal ones” — it is a theme, and it could only have come from this data.
Getting from one to the other means asking, of each group of codes: what is going on here, and why does it matter to my question? Themes should also earn their place across the dataset rather than resting on one articulate participant, and they should be distinct enough that a piece of data does not sit equally well under two of them.
The other thing a strong analysis does is look actively for material that does not fit. A theme that survives a deliberate search for contradictions is far more convincing than one where everything agreed, and reporting the case that did not fit — and what you concluded from it — is one of the cheapest ways to strengthen the chapter.
| Approach | What it produces | Choose it when |
|---|---|---|
| Thematic analysis | Patterns of meaning across the dataset | The default for most doctoral work; flexible, and works with most designs |
| Content analysis | Systematic categorisation, sometimes with counts | You need a replicable scheme, or the frequency of categories is genuinely meaningful |
| Framework analysis | A matrix of cases against themes | Applied or policy work; comparison across cases matters |
| Grounded theory | An explanatory theory of a process | You sampled theoretically and analysed as you collected |
| Interpretative phenomenological analysis | Close accounts of how individuals made sense of an experience | Small, homogeneous sample; the individual case matters |
| Narrative analysis | Analysis of whole stories and their construction | Fragmenting accounts into codes would destroy what you are studying |
| Discourse analysis | How language constructs its objects | Your question is about language as action, not as a window onto thought |
Thematic analysis is the right answer far more often than its reputation suggests. It is sometimes treated as the unambitious choice, but a well-executed thematic analysis with clearly argued themes is stronger than a badly executed anything else — and considerably stronger than a study wearing a grounded-theory label it did not earn.
Whichever you choose, name the specific version and cite it. “Thematic analysis” alone is under-specified: the recognised approaches differ in whether themes are expected to emerge or to be actively constructed, and in what the analyst is understood to contribute. Examiners ask which you followed.
NVivo, ATLAS.ti, MAXQDA and Dedoose organise qualitative analysis. They store material, hold codes against segments, let you retrieve everything coded a particular way, and keep a record of what you did. On a large project that is genuinely valuable — retrieving every extract coded two ways across forty transcripts by hand is miserable and error-prone.
What they do not do is analyse. No software decides what a theme is, whether it answers your question, or what a contradiction between two participants means. Interpretation stays with you, and a chapter that leans on the tool to imply rigour has the relationship backwards.
Which is why “the data were analysed using NVivo” is not a description of a method, and reads to an examiner as an evasion. Name the approach, cite the version of it you followed, describe how coding proceeded and how themes were developed — then mention the software as the place it happened.
Coding by hand is entirely defensible for smaller datasets, and for some approaches — particularly those working with whole narratives — it can be preferable, because the tools encourage fragmenting material into segments.
Qualitative rigour is demonstrated by what you did, not asserted in a paragraph. The things that carry weight:
Two practices are often assumed to be requirements and are not. Having a second coder and reporting agreement between coders fits approaches that treat coding as measurement, such as content analysis, but sits awkwardly with approaches where interpretation is understood as constructed rather than discovered — there, discussion between coders to develop the analysis is usually more appropriate than a coefficient. Returning findings to participants can strengthen credibility, but it is not always suitable: participants may disagree with an analysis that is nonetheless well founded. Do either if it fits your approach, and say why.
Read everything first without coding, so you know the dataset. Then code systematically, labelling segments and keeping definitions in a codebook. Then work on the codes rather than the transcripts — grouping, merging and discarding — and develop themes that make claims rather than name topics. Check each theme against the whole dataset, look deliberately for material that contradicts it, and keep memos throughout, because that is where the interpretation actually forms.
A code labels what a piece of data is about; a theme says what it means. “Workload” is a code. “Staff absorb unmanageable workload privately because admitting it is read as incompetence” is a theme. The quick test is whether the name could have been written before you collected the data. If it could, it is a category rather than a finding.
There is no correct number, but very many usually means codes have been renamed rather than analysed, and very few often means themes are too broad to say anything. Most doctoral chapters settle somewhere in the small handful, sometimes with sub-themes beneath them. The real test is whether each theme is distinct, supported across the dataset, and doing work in answering your question.
No. NVivo and similar tools organise coding and make retrieval across large datasets manageable, which is genuinely useful on a big project. They do not interpret anything. Coding by hand is entirely defensible for smaller datasets and is sometimes preferable, particularly where fragmenting accounts into segments would damage what you are studying.
Usually not, and it is a habit worth resisting. In a purposive sample, frequency reflects who you recruited rather than how common something is, so “eight of twelve participants said” implies a precision the design cannot support. Prevalence language such as “most” or “several” is generally more honest. Counting belongs in content analysis, where the design is built for it.
It depends on your approach. Where coding functions as measurement, as in content analysis, having a second coder and reporting agreement is appropriate. Where the analysis treats interpretation as actively constructed, a coefficient is a poor fit — there, discussing coding with a colleague to develop and challenge the analysis is more useful, and should be described as exactly that rather than dressed up as a reliability check.
The gap between competent coding and defensible themes is where most qualitative chapters lose marks. Describe your data and where the analysis has reached, and a PhD in your field will work through it with you.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.