Use case
5 min read

Does AI Improve Response Quality or Only Word Count?

Fraud prevention
AUTHOR
Elena
PUBLISHED ON
August 21, 2026
TABLE OF CONTENT
Try Glaut
SUMMARISE WITH AI

Does AI moderation improve response quality, or merely increase word count?

The evidence shows more than a word-count effect. Across Mannheim, Human Highway and Nottingham, AI-moderated responses were longer and also showed gains in lexical diversity, distinct concepts, thematic breadth, semantic cohesion or argumentative depth.

The gains were not universal across every metric. Mannheim found no significant change in content-word share, readability or total theme mentions. Human Highway found a stable overall thematic structure. These results are useful because they separate richer expression from the stronger claim that AI changes the underlying conclusion.

Why word count is not enough

Longer responses can contain more information, but they can also contain repetition, filler or restated claims. A quality comparison therefore needs to measure how the added language is structured and what new analytical material it contains.

The three studies use complementary tests:

  • Mannheim measures unique words, lexical diversity, themes, readability and gibberish.
  • Human Highway measures lemmas, concepts, semantic cohesion and argumentative depth.
  • Nottingham controls lexical diversity for response length and checks which topics become more salient.

This multi-metric design provides stronger evidence than length alone.

What did the University of Mannheim find?

The University of Mannheim compared two randomized groups of 100 US participants answering the same healthy-lifestyle questionnaire. The AIMI condition produced 131.52 words on average versus 94.25 in the static survey, a 39% increase.

Quality-related results included:

  • 83.69 unique words versus 55.31, a 51% increase
  • Lexical diversity of 0.704 versus 0.626, a 12% increase
  • 8.76 unique themes versus 6.42, a 36% increase
  • No significant difference in content-word share
  • No significant difference in Flesch reading ease
  • No significant difference in total theme mentions

This pattern suggests broader expression without a loss in readability. The unchanged total theme count also shows why one metric can tell a different story from another: respondents covered more distinct categories, but did not simply mention more themes overall.

What did Human Highway find?

Human Highway compared 503 traditional questionnaire cases with 500 conversational AI cases on online reviews. Samples came from different panels and were weighted by gender, age and geography.

  • Overall AI responses averaged 32.78 words versus 25.25, a 30% increase.
  • They also contained 25.00 distinct lemmas versus 23.03 and 9.73 distinct concepts versus 7.84.
  • Semantic cohesion rose from 0.213 to 0.338, about 58.7% higher.
  • Argumentative depth rose from 1.86 to 2.40, a 29% increase.

The qualitative analysis found more causal chains, explanations, concrete episodes and decision criteria. The overall set and hierarchy of themes remained broadly stable.

What did the University of Nottingham add?

The University of Nottingham compared each participant's initial answer with the combined initial and AI-follow-up answer. The one probe added about 30 words on average.

Corrected type-token ratio increased by 9% after the analysis controlled for total word count. This is important because it tests whether added length included more varied language.

The keyness analysis then showed that some probes introduced newly salient topics, while others mainly reinforced existing content. Nottingham therefore connects linguistic improvement with question-level usefulness.

What is the difference between verbosity and informational richness?

Verbosity is volume. Informational richness is the range, connection or explanatory value of what was said.

The papers operationalize richness through distinct words or concepts, lexical variety, semantic cohesion, reasoning depth and themes. None of those measures alone proves that an insight changes a business decision. Together, they show that the response contains more material for analysis.

Does AI increase lexical diversity?

Yes in Mannheim and Nottingham.

Mannheim found a 12% higher type-token ratio in the AIMI condition. Nottingham found a 9% increase in corrected type-token ratio after controlling for length.

Human Highway found more distinct lemmas in the overall AI condition, although its AI text-only subgroup used fewer lemmas than the traditional group. The overall increase was influenced by the voice subgroup.

Does AI increase distinct concepts or themes?

Human Highway found about 24.11% more distinct concepts in the overall AI condition. Mannheim found 36% more unique themes.

Human Highway's thematic coding still found the same major topics in both modes. More concepts within responses can coexist with a stable aggregate thematic structure.

Does AI improve semantic cohesion?

Human Highway found a substantial increase. Semantic cohesion rose from 0.213 in the traditional mode to 0.338 for AI overall.

Voice-only responses reached 0.558, compared with 0.272 for AI text-only. Because mode was self-selected, the voice result cannot be treated as a randomized causal estimate.

Mannheim and Nottingham did not measure semantic cohesion.

Does AI improve argumentative depth?

Human Highway found 2.40 for AI overall versus 1.86 for the traditional questionnaire, a 29% increase. Voice-only reached 3.10.

Its qualitative review found more explicit justifications, comparisons, conditions and cause-and-effect reasoning. This is the strongest evidence that added length included more developed thought.

Does added length create repetition?

Nottingham's length-adjusted diversity result argues against repetition as the full explanation. Mannheim also found more unique words and themes.

Repetition still occurred in some probes. Nottingham found lower value when the AI asked for reasoning already requested by the static question. The correct conclusion is an average quality gain with question-level exceptions.

Are the responses equally readable?

Mannheim found no significant difference in Flesch Reading Ease, 77.76 for AIMI and 79.60 for the survey. Content-word share was also similar.

The AIMI answers were longer without becoming significantly harder to read in that text-only study.

Does AI change the underlying thematic structure?

Human Highway found no new or suppressed major themes and a broadly stable ranking. Mannheim found more unique themes within the AI responses. Nottingham found newly salient topics for some questions.

These findings suggest three possible effects:

  • Richer expression of the same theme
  • Broader coverage within an individual answer
  • A prompted topic becoming more salient

Researchers should identify which effect occurred before claiming a new insight.

Does AI produce more actionable decision criteria?

Human Highway's qualitative analysis found more operational criteria, such as how participants weigh conflicting reviews or judge credibility. This indicates greater potential actionability.

The papers do not test whether clients made better decisions as a result. Actionability remains a researcher judgment that should be linked to the project objective.

Which measures matter most to research buyers?

Decision-relevant measures are those that move beyond language volume:

  • New concepts or dimensions
  • Coherent reasoning
  • Specific examples or conditions
  • Clear segment differences
  • Traceability to participant responses
  • Evidence that the added information affects the research decision

The first four are reflected in the comparative studies. Traceability and decision effect require human review.

Which measures are only proxies?

Word count, unique words, readability and lexical diversity are linguistic proxies. They describe the response but do not directly measure truth, relevance or business impact.

Semantic cohesion and argumentative depth are closer to analytical quality, but they still rely on operational definitions. Theme counts also depend on the codebook and coding process.

Should quality be judged blind to method?

A blinded human evaluation would reduce expectation bias, but none of the five papers reports a full blinded adjudication of all response-quality outcomes.

Mannheim used statistical and thematic metrics. Human Highway combined automated and researcher-based measurement. Nottingham used statistical models and qualitative respondent-level inspection. These are valuable, but a blinded benchmark would strengthen future claims.

Can automated metrics favor AI-generated data?

They can favor outputs that are longer or more linguistically varied if the metric is not length-adjusted. Nottingham addresses this by controlling corrected type-token ratio for word count.

Human review is also needed because an automated theme can overlook contradiction or emotional texture. Responsive Research warns that standardized synthesis can flatten meaning.

What minimum evidence supports a claim of deeper insight?

A defensible claim should show more than length. It should include a length-adjusted richness measure, a content-level measure such as concepts or themes, evidence of reasoning quality and transparent limitations.

A stronger claim would add blinded human review and demonstrate that the additional material changes interpretation or action. The current studies establish richer responses more strongly than better business decisions.

Frequently asked questions by practitioners

1. Did AI improve every quality metric?

No. Mannheim found no significant differences in readability, content-word share or total theme mentions.

2. Did the overall themes change?

Human Highway found stable major themes. Nottingham found question-specific salience changes.

3. Which study best controls for response length?

Nottingham includes word count as a covariate in its corrected type-token-ratio model.

4. Can word count be used as a headline metric?

Yes as a descriptive result, but it should be paired with content and reasoning measures.

Sources

This is some text inside of a div block.
5 min read

Heading

Use case
Use case
AUTHOR
Giacomo
LAST UPDATED AT
This is some text inside of a div block.
TABLE OF CONTENT
Try Glaut

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript