What a code actually is, the types worth knowing, building a codebook that holds, and staying consistent across a long dataset.
Coding is labelling segments of qualitative data so they can be gathered, compared and built into something analytic. A code is a label, not a finding: the analytic work happens afterwards, when codes are compared, grouped and raised into concepts. Most coding problems in doctoral work come from two things — codes defined too loosely to apply consistently, and coding that never moves beyond describing what was said.
A code attaches a short label to a segment of data, marking it as an instance of something. That is all it does. It is a filing decision with analytic intent, and treating it as more than that is where confusion starts.
Three things follow.
Codes are not themes. A code says what a segment is about; a theme makes a claim about what it means. Grouping codes under headings produces categories, not themes, and the step from one to the other is separate analytic work.
Codes are not findings. A list of forty codes is not a result. Nobody should ever see your full code list in a findings chapter, though it belongs in an appendix.
Codes are working tools and will change. Codes get merged, split, renamed and abandoned throughout. Expecting to define them correctly at the start is what makes people over-plan the codebook and then feel they have failed when it moves.
Saldaña’s manual is the standard reference here and catalogues a large number of coding methods. A handful cover most doctoral work, and knowing their names lets you describe your own practice precisely rather than saying “the data were coded”.
| Type | What it labels | Useful when |
|---|---|---|
| Descriptive | The topic of the segment, in a word or short phrase | First pass over varied material; inventorying a large dataset |
| In vivo | The participant’s own words, kept as the code | Participants’ own language matters; avoiding premature translation into your terms |
| Process | Action, usually as a gerund — “managing uncertainty” | The question is about how something unfolds; central in grounded theory |
| Emotion | Feelings recalled or displayed | Experience and its affective content are part of the question |
| Values | Values, attitudes and beliefs expressed | Studying culture, motivation or conflicting commitments |
| Versus | Oppositions and conflicts in the data | Contested settings; competing accounts of the same events |
Most doctoral projects use two or three of these together, often descriptive coding to get oriented and then something more analytic. Say which you used, and where.
One further distinction cuts across all of them. Semantic coding stays with what was explicitly said; latent coding labels what appears to underlie it. Latent coding is more interpretive and needs more justification, and mixing the two without noticing produces a code list where some labels are descriptions and others are arguments.
Coding runs in at least two passes, and conflating them is why some analyses stall.
First cycle works on the data. You move through transcripts labelling segments, coding generously, staying close to the material. The aim is coverage: getting everything relevant marked, without worrying yet about how it fits together.
Second cycle works on the codes. You are no longer reading transcripts primarily; you are comparing, merging codes that turned out to be the same thing, splitting ones doing two jobs, discarding those that never earned their place, and grouping the rest into categories. This is where the dataset starts becoming an analysis.
People who feel stuck after coding have usually completed the first cycle and are waiting for meaning to appear from it. It will not. The second cycle is a different activity and has to be started deliberately.
Expect a substantial reduction between the two. A first cycle producing well over a hundred codes is normal; a second cycle that does not collapse that considerably usually means merging is being avoided.
Whether or not your approach requires a formal codebook, keeping one makes coding consistent and makes the chapter writeable. Approaches vary: coding-reliability work fixes it early, while reflexive approaches let it evolve throughout. Either way it should exist.
For each code, record:
That last item is doing double duty: it keeps you consistent, and it is the audit trail your methodology chapter will need as evidence of dependability. Reconstructing it at the end is close to impossible.
Doctoral coding happens over a long period, often with gaps. Consistency is the practical problem nobody warns you about, and it is visible in a finished thesis when early transcripts are coded differently from late ones.
What helps:
Coding can be extended indefinitely, and it is a comfortable place to hide when the interpretive work ahead feels harder. Two signals that it is time to stop.
New material is producing no new codes. Everything is being absorbed by existing labels. In approaches using theoretical sampling this connects to saturation; in others it simply means coverage is adequate.
You can describe what the dataset is about without consulting it. If you can sketch the shape of your findings from memory, the coding has done its job and the second cycle is where the remaining value is.
Coding is not the analysis. It is what makes the analysis possible, and a thesis whose findings chapter reports the code list has stopped one step short of the work.
Attaching short labels to segments of data so they can be gathered, compared and built into an analysis. A code marks a segment as an instance of something; it is a filing decision with analytic intent rather than a finding in itself. The interpretive work happens afterwards, when codes are compared, merged and raised into concepts or themes.
They are the coding stages in grounded theory as set out by Strauss and Corbin. Open coding breaks data into concepts and categories; axial coding works out how those categories relate; selective coding identifies a core category and integrates everything around it. They belong to that methodology specifically, so using the terms for a study that did not sample theoretically or analyse alongside collection is a mislabel examiners recognise.
First-cycle coding commonly produces well over a hundred, and that is fine — it is coverage, not analysis. What matters is what happens next: a second cycle that does not substantially reduce that number usually means merging is being avoided. There is no target figure, and a thesis is not judged on the size of its code list.
A code labels what a segment is about; a theme makes a claim about what it means and is organised around a central idea. “Workload” is a code. “Staff absorb unmanageable workload privately because admitting it reads as incompetence” is a theme. Grouping codes under a heading produces a category, not a theme — the move between them is separate analytic work, and it is where most qualitative chapters stop too early.
Most doctoral work does both, and describing that honestly is stronger than claiming a purity examiners will doubt — you bring a question, a literature and a discipline to the data, so pure induction is largely a fiction. What matters is being specific about which codes came from the data, which came from theory, and what you did when the data would not fit the framework.
Write exclusion rules as well as inclusion rules in your codebook, since that is where you will later disagree with yourself. Recode earlier material whenever a code changes, work in reasonably long sessions rather than short scattered ones, and re-read your definitions after any break. Recoding a transcript you coded weeks earlier and comparing the two is a useful check regardless of whether your approach uses reliability statistics.
The stall after first-cycle coding is the most common point at which qualitative theses lose months. Describe your dataset and where the analysis has reached, and a PhD in your field will work through the next step with you.
Leave an email or a WhatsApp number — whichever you prefer — and tick how we should reach you. We reach out within 30 minutes.