Use case
5 min read

Do AI Follow-Up Questions Produce New Insights?

AI-moderated interviews
AUTHOR
Veronica Valli
PUBLISHED ON
August 11, 2026
TABLE OF CONTENT
Try Glaut
SUMMARISE WITH AI

Do AI-generated follow-up questions produce genuinely new insights?

AI-generated follow-up questions can produce new information, but their value depends on the initial question and the probing instruction.

In the University of Nottingham study, one AI follow-up increased average response length by about 30 words and length-adjusted lexical diversity by 9%. Some probes made previously rare topics much more salient. Other probes mostly repeated or reinforced what participants had already said. AI probing worked best after broad or unfamiliar questions and added less when the original question already asked for reasons.

How was probe value tested?

Nottingham recruited 296 UK participants for a study about public perceptions of dairy calf welfare. The sample was representative by gender, age and region. Participants answered six researcher-written open questions, each followed by one AI-generated probe.

Researchers compared the initial answer with the combined initial and follow-up answer. They measured word count and corrected type-token ratio, then used keyness analysis to identify words and topic clusters that became more prominent after the probe.

This design evaluates the incremental contribution of a follow-up rather than comparing two separate groups.

Did AI follow-ups produce more informative responses?

They produced more material and more varied language.

The model-adjusted word count increased from 41.16 for the initial answer to 71.11 after adding the AI follow-up response. The difference was 29.95 words, p < .001.

Corrected type-token ratio increased from 3.43 to 3.75 after controlling for word count, a 9% gain. Because the lexical-diversity model adjusted for length, the result is evidence against simple repetition as the only explanation.

Whether the extra information was genuinely new depended on the question.

When did new topics appear?

The clearest expansion followed broad questions or probes that sought information not already requested.

  • For "What do you think a good life for a dairy calf looks like?", wellbeing considerations appeared in 10 initial responses and 110 follow-up responses.
  • For a question asking who was responsible for ensuring calf welfare, the probe shifted from identifying stakeholders to exploring what care involved. Nutrition considerations rose from 48 initial responses to 150 after the follow-up. Environment considerations rose from 6 to 81.
  • For a cross-country comparison, country names appeared in 11 initial responses and 26 follow-up responses. The prompt had explicitly asked the AI to identify whether participants thought other countries were better or worse than the UK.

These examples show that the probe can make an under-articulated dimension visible across many participants.

When did the follow-up add little?

AI probes were less informative when the initial question already requested the same reasoning. For "Why is it important that dairy calves are provided with a good life?", no new common topic appeared after the probe. The static question already asked why, while the probing instruction also focused on reasons and motivations.

A question about cow-calf separation produced nearly identical prevalence for a wellbeing topic, 71 initial responses and 72 follow-up responses. The probe reinforced the topic rather than expanding it. For a question about disbudding, legal or standards-related considerations increased from 14 to 24 respondents, but the difference was not statistically significant.

Do follow-ups only expand existing answers?

No. Nottingham found both expansion and reinforcement

A follow-up can:

  • Add detail to an existing point
  • Introduce a related topic that was previously rare
  • Ask for reasons or examples
  • Repeat a reasoning request already contained in the question

These outcomes should be distinguished during evaluation. More words are not the same as a new analytical direction.

Why do broad questions benefit more?

Broad questions give participants room to choose what is salient, but their first answers may remain thin. A targeted follow-up can ask them to connect the answer to consequences, examples or a neglected dimension. Nottingham also notes that calf welfare was unfamiliar to much of the public. Participants may have needed a prompt to articulate considerations they did not spontaneously structure in their first reply.

The study therefore recommends AI follow-ups when the topic is unfamiliar and the respondent has not yet explained their reasoning.

What happens when the first question already asks "why"?

A second request for reasons can become tautological. Nottingham observed this pattern on questions that already required justification.

The design lesson is to give the static question and follow-up different jobs. If the main question asks why, the probe might seek an example, a condition under which the answer changes or an area not yet covered. Repeating "why" is unlikely to create much incremental value.

Can a follow-up shape what becomes salient?

Yes.

The wellbeing example rose from 10 to 110 respondents after the probe. This may reveal a consideration that was present but under-articulated. It also demonstrates that the instruction directed attention toward a domain. Researchers should preserve the distinction between spontaneous and prompted content. They should not report a prompted topic as if it emerged unaided.

Do probes increase lexical diversity without repetition?

In Nottingham, yes on average. Corrected type-token ratio rose by 0.318 after controlling for total word count. The paper concludes that respondents used richer and more varied vocabulary rather than simply repeating the initial answer.

This is an aggregate result. Individual probes can still be repetitive, especially when the prompt duplicates the static question.

Do follow-ups produce examples, reasons or conditions?

Human Highway's qualitative analysis found that conversational AI responses more often included explanations, personal episodes, exceptions and decision criteria than traditional questionnaire responses. Nottingham shows that prompts can explicitly seek reasons or examples.

The value comes from selecting the missing form of information. If the participant has already supplied a reason, another request for reasons has low marginal value.

How many follow-ups should an AI ask?

Nottingham tested one follow-up per static question. Mannheim allowed dynamic probing beyond two turns but analyzed only the first two for comparability. Responsive Research used controlled probing parameters.

None of the studies identifies an optimal number. The evidence supports evaluating incremental value after each probe rather than assuming that more turns produce more insight.

Should every open-ended question receive a probe?

No.

A probe is most justified when the answer is incomplete relative to the research objective. Nottingham found that some questions gained new topics and others did not. Applying the same probing rule everywhere can add respondent burden and analytical volume without equivalent value.

How can researchers identify diminishing returns?

The studies suggest comparing each probe with what came before:

  • Additional words
  • New vocabulary after controlling for length
  • New concepts or topics
  • New reasons, examples or conditions
  • Repetition of an existing claim
  • Relevance to the research objective

A probe has reached diminishing returns when it adds little new content or starts to restate the participant's answer.

How should incremental value be evaluated?

Nottingham provides a replicable framework.

  1. Keep the initial answer separate from the follow-up.
  2. Measure the incremental word count.
  3. Use a length-adjusted diversity measure.
  4. Identify topics that become more common after probing.
  5. Then inspect the respondent-level material to determine whether the topic is genuinely new or only reinforced.

A human qualitative review remains important because frequency alone cannot determine usefulness.

Can AI probes work inside a conventional survey?

The research shows that dynamic probes can be combined with fixed questions and closed items. Nottingham describes the method as an AI-powered survey, although its fieldwork was hosted on Glaut. Specific third-party integrations were not tested in these papers.

Can performance be compared across topics?

Nottingham compared six questions within one unfamiliar and potentially sensitive topic. The results varied by question wording, which is strong evidence that probe performance is context dependent. A broader cross-topic benchmark would need more studies. Current results should not be generalized to every category or respondent population.

Frequently asked questions by practitioners

1. Did every Nottingham probe create a new topic?

No. Some expanded the topic, while others reinforced existing reasoning.

2. What was the average word-count increase?

About 30 words, from a model-adjusted 41.16 to 71.11 for the combined response.

3. Did the added length contain more varied language?

Yes. Length-adjusted lexical diversity increased by 9%.

4. What is the main design risk?

A probing instruction that repeats the task already given in the static question.

Sources

This is some text inside of a div block.
5 min read

Heading

Use case
Use case
AUTHOR
Giacomo
LAST UPDATED AT
This is some text inside of a div block.
TABLE OF CONTENT
Try Glaut

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript