PhD Data Analysis

Why analysis takes longer than anyone plans for, how to match it to the design you already have, and the difference between output and findings.

The short answer

Data analysis is the point where your design either pays off or presents its bill. The technique matters less than the fit: an analysis can only answer the question your design was built to answer, and no amount of sophistication afterwards will get a cross-sectional survey to demonstrate causation or nine interviews to speak for a population. The work divides into preparing the data, choosing an approach the design supports, running it properly, and then the step most people underestimate — turning output into findings.

Why analysis always takes longer than planned

Almost every doctoral timeline allocates a few weeks to analysis. Almost none of them survive contact with the data. It is worth knowing in advance where the time actually goes, because it is rarely where the plan assumed.

Preparing the data is most of the job. Cleaning, coding, checking, reverse-scoring, handling missing responses, reconciling two spreadsheets that disagree, transcribing and correcting transcripts. This is unglamorous, it is invisible in the finished thesis, and in most projects it takes more time than the analysis itself.

The first result is usually not the last. An assumption fails, a variable behaves oddly, a theme turns out to be two themes. Each of those sends you back. That is not a sign you are doing it badly — it is what analysis is.

Interpretation is a separate piece of work. Producing a table and understanding what it licenses you to claim are different tasks, and the second one is where a thesis is actually made. Plans that end at “run the analysis” have left out the part that takes longest.

The design decides what the analysis can say

The single most useful thing to understand about analysis is that it inherits limits it cannot remove. Everything about the value of your findings was largely fixed before you opened the software.

A cross-sectional survey can establish that two things go together. It cannot establish that one caused the other, no matter which model you fit — and calling a variable “predictor” does not change that, because in statistics prediction is a mathematical relationship, not a claim about cause. A sample recruited through a Facebook group can describe that group well and cannot be generalised to a country. Twelve interviews can produce a rich, defensible account of how those twelve people understood something, and cannot tell you how common that understanding is.

None of these are flaws. They are boundaries, and staying inside them is what makes a thesis defensible. Nearly every analysis that collapses in a viva collapses because the claim wandered outside the design that was supposed to support it.

So the first question in analysis is not “which test?” but “what is this design entitled to conclude?” Once that is settled, the choice of technique narrows sharply on its own.

The three families of analysis

Which family you are in follows from the question and the data, not from preference or confidence with numbers.

Quantitative analysis

You have numbers, and you want to describe them, compare groups, measure relationships, or test whether something predicted by theory holds. The work runs from describing the data, through checking whether the conditions your chosen test depends on are actually met, to the test itself, and then to effect sizes and intervals — which say how much something matters, where a p-value only says how surprising it would be under the null hypothesis.

Qualitative analysis

You have text, talk or observation, and you want to understand meaning, process or experience. The work runs from familiarisation, through coding, to the step people most often skip: raising codes into themes that say something, rather than leaving them as tidy categories of what was mentioned.

Mixed analysis

You have both, and the point is what happens where they meet. Analysed separately and reported in sequence, they are two thin studies in one document. The value is in integration — the numbers showing an effect and the accounts explaining why it occurs, or contradicting it in a way that is itself the finding.

The order the work goes in

Analysis done out of order tends to produce results that have to be thrown away. This sequence holds for most doctoral projects, whichever family you are in.

  1. Get the data into a state you trust. Clean, check, document every decision. Keep the raw file untouched and work on a copy — you will need to retrace steps later, and “I cannot remember what I did to this column” is a genuinely serious problem at write-up.
  2. Look at it before you model it. Distributions, frequencies, a read-through of transcripts. Most of the surprises that would have derailed the analysis are visible at this stage, cheaply.
  3. Check what your chosen approach assumes. Statistical tests carry conditions. Qualitative approaches carry commitments about what counts as evidence. Both are easier to satisfy before you have built on them.
  4. Run the analysis you planned. Then, separately and honestly labelled, anything you did not plan.
  5. Interpret against the question. Not against what you hoped. This is where output becomes findings.
  6. Write it so someone could repeat it. Versions, settings, decisions, exclusions. If a reader cannot reconstruct what you did, the analysis is not finished.

Output is not findings

This is the most common weakness in doctoral results chapters, and it is worth naming plainly.

Output is what the software produced: a coefficient, a table, a list of themes. Findings are what that means for the question you asked. A results chapter that walks through output table by table has described its own analysis rather than answered anything, and an examiner will keep asking “and so?” until it does.

The move from one to the other is a sentence you have to write yourself: this pattern, given this design and these limits, means this about the question. Software cannot produce that sentence, and it is the sentence the thesis is actually for.

A related habit worth building: state what would have counted as the opposite result. If no plausible pattern in your data would have changed your conclusion, the analysis is decoration.

Software is a tool, not a method

SPSS, R, Python, Stata, AMOS, SmartPLS, NVivo and the rest do what you tell them and have no opinion about whether it made sense. This produces two failures that recur constantly.

The first is running something the data cannot support and reporting it because it produced a number. Nothing in the software objects. The objection arrives from an examiner or a reviewer instead, at the point where there is no time left to collect anything else.

The second is treating the tool as the method. “The data were analysed using NVivo” describes where the work happened, not what was done — NVivo organises coding, it does not interpret. In the same way, naming SPSS says nothing about which test was run, why it was appropriate, or whether its conditions held. Examiners read the software sentence as a warning sign for exactly this reason.

Name the method, justify it, then mention the tool.

Questions researchers ask

How long should data analysis take in a PhD?

Longer than the plan says, and the reason is almost always data preparation rather than the analysis itself — cleaning, checking, transcribing and documenting routinely take more time than running anything. It is also normal to go round the loop more than once, because a failed assumption or an unexpected pattern sends you back. Budget for the interpretation stage separately: turning output into defensible findings is its own piece of work.

Which analysis method should I use?

It follows from the question and the design rather than being chosen freely. Questions about magnitude, difference and relationship point to quantitative analysis; questions about meaning, process and experience point to qualitative; questions needing both an effect and an explanation of it point to mixed, provided the strands are genuinely integrated. If two methods still seem equally plausible, the question is usually not yet specific enough.

Can I change my analysis after I have collected the data?

Yes, and it happens often. What matters is being straightforward about it: analysis you planned in advance and analysis you decided on after seeing the data carry different weight, and examiners are far more comfortable with exploratory work that is labelled as exploratory than with the same work presented as if it had been the plan all along.

My results are not significant. Have I failed?

No. A well-designed study that finds no effect has produced a real result, and the honest reporting of it is worth more than a significant finding squeezed out by testing everything until something appeared. What matters is whether the study was capable of detecting an effect worth detecting — which is a question about design and sample size, and is best answered before collection rather than after.

Do I need to know how to code in R or Python?

Only if your analysis needs it. Plenty of excellent doctorates are completed in SPSS, Stata or NVivo without a line of code. R and Python become worth the learning curve when you need something the menus do not offer, when the analysis has to be repeated many times, or when reproducibility matters enough that you want the whole procedure written down as a script.

Related guides

Get the analysis checked before it reaches your supervisor

Describe your design, your data and what you are trying to establish. A PhD in your own field will tell you whether the analysis fits — and what the results are actually entitled to claim.

Discuss your analysis