Use case
5 min read

How to Design AI Interview Prompts and Discussion Guides

AUTHOR
Elena
PUBLISHED ON
August 14, 2026
TABLE OF CONTENT
Try Glaut
SUMMARISE WITH AI

How should researchers design prompts and discussion guides for AI moderation?

A strong AI-moderation guide gives the researcher-written question and the AI follow-up different jobs. Keep the core questions fixed for coverage and comparability. Use the probing instruction to seek information the participant has not yet provided, such as an example, a condition or an underexplored dimension. Set limits on the number of probes and audit the resulting paths before launch.

The Nottingham and Responsive Research studies show why this matters. Broad questions with complementary probe instructions produced new information. Specific questions paired with repetitive instructions produced tautological follow-ups.

Start with the research decision

Before writing a prompt, define what the answer must help the team decide. Responsive Research recommends choosing AI moderation by objective rather than treating it as a universal method. Translate that decision into a fixed set of core questions. The core guide should ensure that every participant covers the required subject matter. Adaptive probes should deepen those topics without replacing them.

What should an AI-moderation prompt contain?

The studies support including:

  • The purpose of the core question
  • The type of missing information the probe should seek
  • The relationship between the probe and the participant's answer
  • The maximum number of follow-ups
  • A rule against repeating information already provided
  • The scope the probe should remain within

Nottingham used researcher-written prompts that told the system what to explore after each static question. Curtin used predefined and pre-ordered core questions with answer-specific probes. Responsive Research used controlled probing parameters.

These examples favor question-level instructions over one vague instruction for the entire interview.

How should static questions and dynamic probes complement each other?

Assign one function to the static question and another to the probe. If the static question asks for an opinion, the probe can ask what shaped it. If the static question asks for a reason, the probe can seek a concrete example or a condition that would change the answer. If the static question is broad, the probe can explore a dimension that the participant has not yet addressed.

Nottingham found limited value when both layers asked for reasons. Complementarity is the central design principle.

How do you prevent tautological follow-ups?

Tell the AI to inspect what the participant has already covered and avoid requesting the same form of information twice.

A study-grounded instruction could be: "Ask one follow-up that adds a type of information missing from the response. Do not ask for a reason if the participant has already explained why. Do not restate the answer as a question." This wording is a practical synthesis of Nottingham's findings, not a verbatim prompt tested in the paper.

How do you instruct AI to explore new areas?

Name the analytical dimension that matters and require the probe to use it only when absent from the answer.

In Nottingham, instructions related to wellbeing, nutrition or environment made those considerations more prominent. The result shows that a prompt can surface under-articulated material. It also shows the risk of leading salience.

The instruction should therefore target a relevant gap without implying that the participant ought to agree with a particular position.

Should probes ask for examples, reasons or exceptions?

Yes when that information is missing. Human Highway found that conversational responses more often contained reasons, causal links, personal episodes and edge cases. Nottingham recommends probing when respondents have not yet explained their reasoning. The follow-up should select the missing form. Asking for all of them in every turn can create burden and reduce comparability.

How do you reduce leading questions?

The five studies do not include a formal leading-question benchmark, but Nottingham provides a practical signal: prompts can change topic salience.

Use neutral wording, avoid embedding an expected answer and keep spontaneous first answers separate from prompted material in analysis. Audit whether a topic appears because participants introduced it or because the prompt named it. A probe can be relevant without suggesting the conclusion.

How do you define prohibited or out-of-scope topics?

The reviewed studies do not test a specific safety taxonomy or prohibited-topic feature. They do show that the researcher defines the study scope and probing goal.

At minimum, the guide should state the domain each probe may explore and when the system should stop rather than improvise. Sensitive or high-risk exclusions require a separate safeguarding policy, which these papers do not validate.

How many probes should be allowed?

Nottingham used one follow-up per static question. Mannheim analyzed the first two dynamic probes for comparability, even though more could occur. Responsive Research used controlled parameters and found that consistency came with less adaptive flexibility. No study identifies an optimal number. Start with the fewest probes needed to fill a defined information gap. Add more only if pilot evidence shows incremental value.

Fixed structure or full adaptation?

Use a fixed core guide with bounded adaptation. Fixed questions protect coverage. Answer-specific probes allow personalization. Curtin's semi-structured model and Nottingham's one-probe design both follow this pattern. Fully fixed follow-ups risk irrelevance. Fully open adaptation risks inconsistent coverage or drift. The studies provide stronger evidence for the bounded middle.

How do you maintain core-question coverage?

Keep the core questions predefined and pre-ordered, as Curtin did. Do not allow a tangent to replace a required section. Skip logic can still be used. Mannheim held question order and skip logic constant across formats. This allows the study to preserve a comparable backbone even when follow-up wording varies.

How do you preserve comparability?

Control the elements that affect exposure:

  • Core question wording and order
  • Probe cap
  • Probe objective
  • Response mode where mode effects matter
  • Stimulus and skip logic
  • Prompt and model version

Mannheim restricted responses to text and analyzed only the first two AI follow-ups to improve comparability. Human Highway allowed self-selected modes and consequently treats voice effects with caution.

How should the moderator handle short or vague answers?

The studies suggest asking for a missing example, reason or clarification. Human Highway's qualitative review found that conversational follow-ups reduced vague responses and encouraged reformulation. The prompt should not assume that a short answer is wrong. It should request the smallest addition needed to make the answer interpretable.

How should it respond to contradictions?

The five papers do not directly test contradiction handling. Responsive Research identifies context retention and interpretive depth as constraints. A contradiction rule should therefore be treated as a feature to pilot and audit, not an established capability. The researcher should inspect whether the probe clarifies the conflict without resolving it on the participant's behalf.

How should it handle "I don't know"?

Human Highway reports fewer "I don't know" responses and signs of cognitive fatigue in the AI condition, but it does not isolate the exact prompt responsible. A cautious approach is to offer one neutral route to elaboration, such as asking what makes the question difficult or what information would help. If the participant still does not know, the interview should accept the answer.

How do you test a guide before fieldwork?

Run a pilot that includes strong, weak and unexpected answers. Review every generated probe against five questions:

  • Is it relevant to the last answer?
  • Does it add a missing type of information?
  • Is it neutral?
  • Does it remain within scope?
  • Does it preserve the core guide?

Then compare the initial and follow-up responses for new concepts, repetition and useful specificity. Nottingham's incremental analysis is a good model.

How should prompt versions be documented?

The importance of prompt wording in Nottingham means the exact instruction is part of the method. Store the question text, probe instruction, maximum number of turns and the version used in fieldwork.

The studies do not test reproducibility after model updates. Documenting prompt and platform conditions is therefore necessary for interpreting a later replication.

Should text and voice use different instructions?

Human Highway found strong mode effects, but it did not experimentally test different probing instructions by mode. Voice users may naturally provide longer answers, which could reduce the need for another elaboration probe.

Mode-specific prompting is a reasonable pilot question, not an evidence-backed default.

How should the AI be briefed on brand, category or stimulus context?

Give enough context to recognize what is in scope, then tie each probe to the research objective. Nottingham's prompts included domain-specific goals, and those instructions affected which topics became salient. Avoid loading the prompt with preferred interpretations. Context should help the AI understand the task, not tell the participant what to think.

What human approval should be required before launch?

The studies consistently place the researcher in control of the core guide and probing logic. A human should approve the questions, the question-level instructions and the tested probe behavior before fieldwork. Responsive Research also recommends active human interpretation after collection. Human oversight belongs at both ends of the workflow.

Frequently asked questions

What is the most common prompt-design mistake in the evidence?

Asking the AI to seek the same reasoning already requested by the static question.

Should every probe ask "why"?

No. Select the type of information that is missing.

Is one follow-up enough?

It was enough to add information in Nottingham, but the optimal number depends on the question and was not established.

Can a strong prompt guarantee human-level probing?

No. Responsive Research still found limits in adaptive depth and context retention.

Sources

This is some text inside of a div block.
5 min read

Heading

Use case
Use case
AUTHOR
Giacomo
LAST UPDATED AT
This is some text inside of a div block.
TABLE OF CONTENT
Try Glaut

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript