Questionnaire Design

Writing items that measure what you intended, the faults that quietly ruin data, and why piloting is the cheapest thing you will ever do.

The short answer

A questionnaire is a measuring instrument, and every ambiguous item adds noise you cannot remove later. The decisions that matter are made before distribution: whether to use a validated instrument or write your own, how each item is worded, what response format it uses, and whether you piloted properly. Once responses are in, a badly worded question cannot be fixed — only excluded, or reported with a caveat.

Use an existing instrument if one fits

The first question is not how to write items but whether you need to. If a validated instrument exists for your construct and population, using it is almost always the better decision.

You gain published evidence about its structure and behaviour, comparability with other studies using it, a defensible answer when asked why these items, and weeks of your own time. Writing and validating an instrument is a substantial project — item generation, expert review, cognitive interviewing, piloting, and then factor analysis and reliability evidence — and it becomes part of the doctorate rather than a step within it.

Two cautions when borrowing:

Validation is population-specific. An instrument validated with North American undergraduates has not been validated for mid-career professionals in another country, and saying so honestly is stronger than ignoring it. Report reliability in your own sample rather than quoting the original paper.

Changing items breaks the evidence. Reword an item, drop one, or translate the scale, and the published validity evidence no longer straightforwardly applies to your version. Modifications are sometimes necessary — report exactly what you changed and why, and treat the modified version as needing its own evidence.

Check licensing too. Some instruments are freely available, some require permission, and some are commercial.

Writing items that measure one thing

Where you do write your own, most problems come from a small set of recurring faults. Each is obvious once named and invisible while drafting.

  • Double-barrelled items. “My supervisor is available and supportive” — a respondent who finds them available but unsupportive cannot answer. Watch for “and” and “or”.
  • Leading wording. “How much did you benefit from the training?” presumes benefit. Ask whether they benefited before asking how much.
  • Ambiguous quantifiers. “Regularly”, “often” and “frequently” mean different things to different people. Where you need frequency, use actual intervals.
  • Assumed knowledge or experience. Items only answerable by people in a particular situation need a filter question first.
  • Negation, especially double negation. “I do not feel unsupported” is a puzzle. Negatively worded items are sometimes used deliberately to disrupt automatic responding, but they are frequently misread and can distort a scale’s structure — use them sparingly and knowingly.
  • Jargon and internal terminology. Words that are unambiguous in your field are not outside it.
  • Sensitive items placed early. Income, health and anything embarrassing belong later, once the respondent has committed some time.
  • Recall beyond what people can do. Asking how many times something happened in the last year produces estimates rather than counts. Shorten the window or accept the imprecision explicitly.

A test that catches most of these: read each item aloud. Anything that needs re-reading will be misread by some respondents, and you will never know which.

Response formats

The response format determines what analysis is available to you, so it is an analysis decision as much as a design one.

Number of points. More points give finer discrimination up to a point, beyond which respondents cannot use them meaningfully. Five and seven are the common choices and both are defensible.

Midpoints. Including a neutral option lets genuinely undecided people answer honestly; excluding it forces a direction and can produce false differentiation. Neither is universally right — decide deliberately and be able to justify it.

Label everything, or just the ends? Fully labelled scales are usually interpreted more consistently, since respondents are not left to guess what point three means.

“Not applicable” and “don’t know” are different from a midpoint and from each other. Omitting them forces meaningless answers; including them without thought invites their overuse. Provide them where they are genuinely possible answers.

Keep the direction consistent across the questionnaire, so agreement always means the same thing. And if you reverse-score items, verify the reversal worked before analysing — a quick correlation between a reversed item and its scale will show if it did not.

On whether Likert data can be treated as continuous: a single item is ordinal, and whether summed multi-item scales may be analysed as continuous is genuinely debated with respectable arguments on both sides. State your position, cite it, and where the choice affects your conclusions, show the result holds either way.

Structure, length and order

Order affects answers. Earlier questions frame later ones, and a set of items about problems will colour a subsequent question about overall satisfaction. Group related items, move from general to specific, and consider whether any sequence is doing unintended work.

Length drives dropout. Long questionnaires produce incomplete responses and inattentive answering towards the end, and both damage your data. Every item should earn its place by connecting to a research question — a useful discipline is to map each item to the question it serves and delete anything that maps to nothing. “It would be interesting to know” is not a reason.

Demographics usually go last, unless you need them to filter. They are easy to answer, so they are a poor opening, and sensitive ones benefit from coming after the respondent has invested time.

Tell people what they are in for. A realistic time estimate and a progress indicator reduce abandonment. Underestimating the time is counterproductive: people who feel misled leave.

Test on a phone. Much survey completion happens on mobile devices, and matrix questions in particular often become unusable on a small screen.

Piloting: the cheapest insurance available

Almost every serious questionnaire problem would have been caught by a pilot, and piloting is skipped more often than any other design step.

Two different activities are both worth doing.

Cognitive interviewing — sitting with a handful of people from your target population while they complete the questionnaire and think aloud. This is the single most valuable thing you can do, because it reveals items being interpreted in ways you did not anticipate. Researchers are consistently surprised by what this finds, and it takes a few hours.

A pilot administration — a small trial run under realistic conditions, checking completion time, dropout points, items with high non-response, response distributions that are unusable because everyone answers the same way, and whether your data exports in a form you can actually analyse.

That last point catches people out: run your intended analysis on the pilot data, however few responses. Discovering after full collection that a variable was captured in a form your analysis cannot use is a genuinely painful and entirely avoidable experience.

Report the piloting in your methodology chapter. It is evidence of rigour, and its absence is a reasonable thing for an examiner to ask about.

Response rates and who does not reply

Low response rates are normal, particularly for online surveys, and there is no universal threshold that makes one acceptable.

What matters more than the rate is whether non-response is systematic. If the people who did not reply differ from those who did in ways relevant to your question, your sample is biased regardless of its size — and a large biased sample is not better than a small one, only more confidently wrong.

What helps: a clear explanation of purpose and who you are, a realistic time estimate, a short questionnaire, a plausible reason why their response matters, reminders, and where appropriate an incentive. Institutional endorsement helps in organisational research.

What to report: how many were approached, how many responded, the rate, and anything you know about non-responders — even basic comparison against known population characteristics strengthens the chapter considerably. Acknowledging likely bias and reasoning about its direction is far stronger than presenting a response rate without comment.

Questions researchers ask

Should I use an existing questionnaire or write my own?

Use a validated instrument where one fits your construct and population — you gain published evidence about how it behaves, comparability with other studies, and weeks of time. Writing your own is legitimate when nothing suitable exists, but developing and validating it becomes part of the doctorate rather than a step within it: item generation, expert review, cognitive interviewing, piloting, and then factor analysis and reliability evidence.

What makes a bad survey question?

Asking two things at once, leading wording that presumes an answer, vague quantifiers like “often” that mean different things to different people, assumed knowledge without a filter question, negation, field jargon, and recall demands beyond what people can actually do. Reading each item aloud catches most of them — anything needing a second reading will be misread by some respondents, and you will never know which.

How many points should my rating scale have?

Five or seven are the common choices and both are defensible; beyond a certain point respondents cannot use the distinctions meaningfully. The more consequential decisions are whether to include a midpoint — which lets genuinely undecided people answer honestly but can attract fence-sitting — and whether to label every point, which usually produces more consistent interpretation than labelling only the ends.

How long should a questionnaire be?

Short enough that people finish it attentively, which is shorter than most first drafts. Length drives both dropout and careless answering near the end. The useful discipline is mapping every item to the research question it serves and deleting anything that maps to nothing — “it would be interesting to know” is not a reason to include an item.

Do I need to pilot my questionnaire?

Yes, and it is the cheapest insurance in the whole design. Sit with a few people from your target population while they complete it and think aloud — this reliably reveals items being read in ways you did not intend. Then run a small trial administration and, importantly, run your intended analysis on that pilot data. Discovering after full collection that a variable was captured in a form your analysis cannot use is avoidable and painful.

What is an acceptable response rate?

There is no universal threshold, and rates for online surveys are commonly low. What matters more is whether non-response is systematic: if those who did not reply differ from those who did in ways relevant to your question, the sample is biased regardless of size. Report how many were approached and what you know about non-responders, and reason about the likely direction of any bias rather than presenting a rate without comment.

Related guides

Have the questionnaire reviewed before it goes out

Once responses are in, a badly worded item cannot be repaired — only excluded. Send your draft and a PhD in your field will go through it item by item.

Discuss your instrument