Use case
5 min read

AI Interviews vs Online Surveys: Which Produces Better Data?

AI-moderated interviews
AUTHOR
Elena
PUBLISHED ON
August 1, 2026
TABLE OF CONTENT
Try Glaut
SUMMARISE WITH AI

AI-moderated interviews produced longer and more varied open-ended responses than static surveys in studies from Mannheim University, Human Highway and Nottingham University. The gains extended beyond word count: researchers observed more unique words or lemmas, more distinct concepts, higher lexical diversity and, in some designs, stronger semantic cohesion or argumentative depth.

That does not make AIMIs better for every research question. Static surveys remain appropriate when standardized measurement and concise responses are the priority. AI follow-ups add the most value when an open question leaves important reasoning unexplored.

The strongest comparative evidence

The University of Mannheim randomly assigned 200 US participants to an AIMI or a static SoSci survey. Both groups answered the same healthy-lifestyle questionnaire. The static condition included two predefined follow-ups after each open question, while the AI condition generated contextual probes. Only the first two AI follow-ups were analyzed.

Compared with the survey condition, AIMI responses contained:

  • 39% more words, 131.52 versus 94.25
  • 51% more unique words, 83.69 versus 55.31
  • 12% higher lexical diversity, 0.704 versus 0.626
  • 36% more unique themes, 8.76 versus 6.42
  • No significant loss in readability or content-word share
  • A 0% gibberish rate versus 10% in the static survey sample

The overall participant-experience score was 4.22 for AIMI and 3.98 for the survey.

Does AI moderation produce longer answers?

Yes, across all three survey comparisons.

  • Mannheim University found 39% more words.
  • Human Highway found that conversational AI responses averaged 32.78 words versus 25.25 in the traditional questionnaire, a 30% difference.
  • Nottingham University found that adding one AI follow-up increased the model-adjusted response from 41.16 to 71.11 words, a 75% increase over the initial response alone.

The studies use different designs, so the percentages should not be compared as if they measure the same treatment. Nottingham compares an initial answer with the combined initial and follow-up answer. Human Highway compares separate panels and includes self-selected voice responses in the overall AI average. Mannheim offers the cleanest randomized text-only comparison.

Do longer answers contain more information?

Often, but length alone cannot establish that.

  • Mannheim found more unique words and more unique themes. The total number of theme mentions did not differ significantly, which suggests that AIMI broadened the range of ideas rather than merely increasing repeated mentions.
  • Human Highway found 24.11% more distinct concepts in the overall AI condition. Semantic cohesion rose from 0.213 to 0.338, and argumentative depth rose from 1.86 to 2.40. Its qualitative review found more explanations, concrete examples and decision criteria.
  • Nottingham controlled lexical-diversity analysis for response length. Corrected type-token ratio increased by 9% after the AI follow-up, supporting the conclusion that respondents added varied language instead of simply restating the first answer.

Do AIMIs generate a wider range of themes?

  • Mannheim found 36% more unique themes. Nottingham found that probes sometimes made previously rare topics much more salient. For a broad question about a good life for a dairy calf, wellbeing considerations rose from 10 respondents in the initial answers to 110 after the follow-up.
  • Human Highway reached a different but compatible result. The set and hierarchy of major themes remained broadly stable between the survey and AI conditions. The conversational format enriched how people expressed those themes rather than creating a different substantive story.
  • The combined interpretation is that AI can broaden individual responses without necessarily changing the overall thematic conclusion.

Are AI responses more coherent and better argued?

  • Human Highway provides the strongest evidence. Semantic cohesion was about 58.7% higher in the overall AI condition, and argumentative depth was 29% higher. Voice responses drove especially large gains.
  • Mannheim did not measure semantic cohesion or argumentative depth. It found that readability was similar across conditions, so richer responses were not significantly harder to read.

These are complementary findings, not interchangeable metrics.

Does AI reduce gibberish or low-effort input?

Mannheim found 10 gibberish cases in the static survey group and none in the final AIMI group. Its AIMI condition used an uncooperative-response detector and continued collection until 100 valid datasets were obtained. This is encouraging evidence for response validity, but it does not isolate whether the improvement came from conversational probing, interface design or active quality filtering. The study also retained static-survey gibberish to keep that sample at 100 while excluding AIMI gibberish before analysis.

Human Highway qualitatively observed fewer vague answers and signs of cognitive fatigue in the AI condition, but did not report the same experimental validity metric.

Do respondents prefer the conversational format?

Both comparative studies reported a better overall experience.

  • Mannheim found a 6% higher aggregated experience score. Participants rated AIMI as more conversational, less repetitive and more likely to make them feel understood. Comfort and ease of expression were high in both groups and did not differ significantly.
  • Human Highway found a mean rating of 8.84 for AI versus 8.09 for the traditional questionnaire. The share rating the experience at least 9 was 66.2% for AI and 43.4% for the survey.

These findings cover the tested questionnaires and samples. They do not prove that every conversational design will be preferred.

Does conversational probing introduce bias?

It can influence salience, which is one form of measurement effect. Human Highway found stable themes across modes, evidence against a large thematic distortion in that study. Nottingham shows why researchers still need caution. A prompt aimed at wellbeing made that topic much more common after the follow-up. The probe may have uncovered a latent consideration, but it also directed attention toward it.

The correct interpretation depends on the research goal. A probe can reveal under-articulated reasoning and still shape what participants discuss. Researchers should report the probing instructions and distinguish spontaneous first answers from prompted elaboration.

Can AIMIs retain quantitative standardization?

Yes at the level of core structure. The studies kept researcher-written questions, order and closed items consistent. Large samples and demographic controls were possible. Personalized probes create different paths after those core questions. Standardization therefore applies to coverage and rules, not to identical wording in every follow-up.

For statistical estimates, the sample design and the closed measures remain decisive. Open-ended themes can be counted, but frequency should not be treated as population prevalence without an appropriate sample and coding method.

When is a traditional survey the better option?

A static survey is a better fit when:

  • The construct is already well specified
  • The main need is comparable closed measurement
  • Open answers are intentionally brief
  • The initial question already captures the required reason
  • Personalized probing could contaminate the measure

Nottingham found that AI probes added little when a specific question already asked respondents to explain why. More interaction is not automatically more informative.

Should AI replace every static open end?

No. Use dynamic probing where the incremental information is worth the added respondent time and analytical volume. The evidence favors broad or unfamiliar questions where respondents have not yet explained their reasoning. It is weaker for specific questions that already request reasons, examples or conditions. A practical design is to reserve AI probes for the questions that carry the greatest decision value.

Can AI interviews be embedded in an existing survey?

The papers show that open-ended AI probing can coexist with closed questions and conventional questionnaire logic. Nottingham describes AI-powered surveys as NLP integrated within a standard survey format, although its study was hosted on Glaut. The five papers do not test specific integrations or iframe implementations. Platform-level integration claims require product documentation rather than this research evidence.

How should researchers compare quality-adjusted cost?

None of the five studies reports a full cost comparison. They provide quality and fieldwork observations, not a validated cost-per-insight model. A defensible comparison should keep recruitment, incentive, interview length, data cleaning and human review visible. Cost should then be assessed against the decision-relevant quality produced, not word count alone.

Frequently asked questions by practictioners

1. Which study offers the cleanest survey comparison?

Mannheim used random assignment, identical question order and text-only responses across two groups of 100.

2. Did all studies find new themes?

No. Mannheim found more unique themes per response. Human Highway found a stable overall thematic structure. Nottingham found new or newly salient topics for some questions, but not others.

3. Does AI make answers harder to read?

Mannheim found no significant difference in reading ease.

4. Can the results be generalized to every topic?

No. The studies covered healthy lifestyles, online reviews and dairy calf welfare, with different samples and designs.

Sources

This is some text inside of a div block.
5 min read

Heading

Use case
Use case
AUTHOR
Giacomo
LAST UPDATED AT
This is some text inside of a div block.
TABLE OF CONTENT
Try Glaut

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript