<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>MeasuringU</title>
	<atom:link href="https://measuringu.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://measuringu.com</link>
	<description>UX Research and Software</description>
	<lastBuildDate>Thu, 01 Oct 2026 20:19:51 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://measuringu.com/wp-content/uploads/2020/11/site-icon.png</url>
	<title>MeasuringU</title>
	<link>https://measuringu.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>ChatGPT vs Humans for Coding Open-Ended Responses, Three Years Later</title>
		<link>https://measuringu.com/chatgpt-vs-humans-for-coding-open-ended-responses-three-years-later/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=chatgpt-vs-humans-for-coding-open-ended-responses-three-years-later</link>
		
		<dc:creator><![CDATA[Will Schiavone, PhD • Eva Sundberg, MPH • Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 05:01:43 +0000</pubDate>
				<category><![CDATA[Usability]]></category>
		<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[comment coding]]></category>
		<category><![CDATA[kappa]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48673</guid>

					<description><![CDATA[Can ChatGPT Replace UX Researchers? We wrote that title into our article three years ago, and some people thought it was click-bait. However, the subtitle spoke of what we thought was a potentially good job for AI—coding open-ended comments from a survey. These comments often get ignored, with only a quick summary or relegation to [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1.jpg"><img fetchpriority="high" decoding="async" class="alignleft wp-image-48718 size-medium" src="https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-300x169.jpg" alt="feature image with robot and human and clipboard" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/09/092926-FeatureImage1.jpg 2000w" sizes="(max-width: 300px) 100vw, 300px" /></a><em>Can ChatGPT Replace UX Researchers?</em></p>
<p>We wrote that title into our article <a href="https://measuringu.com/classification-agreement-between-ux-researchers-and-chatgpt/">three years ago</a>, and some people thought it was click-bait. However, the subtitle spoke of what we thought was a potentially good job for AI—coding open-ended comments from a survey. These comments often get ignored, with only a quick summary or relegation to a word cloud, though they can reveal valuable insights.</p>
<p>Our initial results showed promise. And since then, the AI landscape has exploded. What started as a one-horse race with OpenAI is now a battle among industry heavyweights, with Google’s Gemini and Anthropic’s Claude emerging as serious contenders in their own right. Models have gotten faster, smarter, and significantly better at reasoning. With the tech maturing so rapidly, we decided it was time to revisit our original experiment and see where things stand today.</p>
<p>Back in 2023, our study found substantial agreement among human coders, among repeated runs of ChatGPT-4, and between the two (mean kappas from .632 to .704), suggesting that ChatGPT could be useful for helping to automate this analysis. But how much better has ChatGPT gotten, and how do the new industry leaders stack up against one another?</p>
<p>To find out, we’re kicking off a series of benchmarking tests, starting right where we began with re-running our original dataset and prompt through version 5.6 of ChatGPT. In upcoming posts, we’ll put Gemini and Claude through the exact same paces.</p>
<p>In short, we found that ChatGPT-5.6 not only matched human levels of interrater agreement and achieved near-perfect self-consistency, but has also evolved into a far more granular analyst, subdividing core themes into small subcategories.</p>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2>Study Details</h2>
<p>To maintain direct comparability, we used the SUPR-Q<sup>®</sup> benchmarking data from our original 2023 study. The dataset comprises participant-reported problems and frustrations from their most recent visit to one of three websites:</p>
<ul>
<li>Office Depot (52 statements from <a href="https://measuringu.com/office-supply-benchmark-2023/">Office Supplies Websites Survey</a>)</li>
<li>AT&amp;T (52 statements from <a href="https://measuringu.com/wireless-benchmark-2023/">Wireless Service Provider Websites Survey</a>)</li>
<li>Food Lion (50 statements from <a href="https://measuringu.com/grocery-benchmark-2022/">Grocery Websites Survey</a>)</li>
</ul>
<p>All datasets and prompts remained identical to the 2023 evaluation (see the <a href="_The_Verbatim">appendix</a> for the full prompt). On August 11, 2026, we evaluated each dataset across three separate runs using ChatGPT-5.6 Sol set to “high” thinking. Each run was executed in a fresh chat session with the “Memory” feature disabled to ensure independence between trials. For our comparison, we retained the baseline coding from the three human evaluators (1, 2, 3) and the three original ChatGPT-4 runs (A, B, C), and designated the three new ChatGPT-5.6 Sol runs as D, E, and F.</p>
<h3>How we aligned themes this time (and how the approach differs from 2023)</h3>
<p>In our initial 2023 study, ChatGPT-4 consistently grouped survey responses into six to nine broad categories. In contrast, ChatGPT-5.6 configured with “high” thinking frequently split responses into finer subcategories. While these granular divisions preserve the same core conceptual insights, they introduce practical analytical trade-offs. For instance, across all new runs on the Office Supplies data, the “Out of Stock” category was subdivided into two to four narrower groups. This added granularity may offer researchers deeper contextual nuance, but it carries the risk of fragmenting macro-level themes if an LLM over-subdivides this data.</p>
<p>This structural shift required us to adapt our alignment methodology. In 2023, we aligned themes using a bottom-up approach based strictly on the highest number of overlapping statements across runs. However, applying the overlap-based approach to the new outputs would have artificially depressed agreement metrics. Instead, we evaluated the new runs against the original framework, manually combining and mapping subcategories that logically aligned with the baseline structure. See Table 1 for an example of how the new ChatGPT-5.6 categories were combined to align with our existing themes.</p>
<p>By shifting from a purely data-driven overlap model to a top-down frame, we preserved structural continuity across iterations that allowed for a more direct, apples-to-apples comparison between ChatGPT-5.6 and our original baseline codes.</p>
<p><strong>Office Supplies Data: Theme Alignment</strong></p>

<table id="tablepress-1080" class="tablepress tablepress-id-1080">
<thead>
<tr class="row-1">
	<th class="column-1">Category</th><th class="column-2">GPT D</th><th class="column-3">GPT E</th><th class="column-4">GPT F</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">No Issues</td><td class="column-2">No problems or frustrations</td><td class="column-3">No problems or frustrations reported</td><td class="column-4">No problems or frustrations reported</td>
</tr>
<tr class="row-3">
	<td class="column-1">Appearance</td><td class="column-2">Unappealing, bland, or outdated visual design</td><td class="column-3">Unappealing, bland, or outdated visual design;<br />
Product image quality<br />
</td><td class="column-4">Visual design is bland, dated, or unappealing</td>
</tr>
<tr class="row-4">
	<td class="column-1">Navigation</td><td class="column-2">Navigation and browsing difficulties; <br />
Too many or overly complex product categories; <br />
Search, filtering, and product-findability problems</td><td class="column-3">Navigation and information architecture; <br />
Search, filtering, and product findability<br />
</td><td class="column-4">Navigation, browsing, and category organization; <br />
Search, filtering, and product-finding difficulties<br />
</td>
</tr>
<tr class="row-5">
	<td class="column-1">Cluttered</td><td class="column-2">Cluttered, busy, or overwhelming interface; <br />
Distracting ads, promotions, pop-ups, or graphics</td><td class="column-3">Cluttered, overwhelming, or distracting interface</td><td class="column-4">Cluttered, busy, or overwhelming interface; <br />
Promotions, ads, recommendations, and pop-ups are distracting<br />
</td>
</tr>
<tr class="row-6">
	<td class="column-1">Out of Stock</td><td class="column-2">Inventory information is inaccurate; <br />
Website does not fully reflect local-store inventory; <br />
Out-of-stock, unavailable, or limited product selection; <br />
In-store pickup availability limitations</td><td class="column-3">Store inventory accuracy and online/in-store synchronization; <br />
Product availability, out-of-stock items, and pickup limitations; <br />
Limited or incomplete product assortment<br />
</td><td class="column-4">Store inventory and website availability mismatch; <br />
Product availability, stock, and assortment limitations<br />
</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left"><strong>Table 1:</strong> Alignment of themes for ChatGPT-5.6 runs in the Office Supplies dataset.</p>
<h2>Study Results</h2>
<p>We used the baseline metrics from our 2023 study as a starting point to evaluate interrater agreement and assess how qualitative analysis has evolved with the newest generation of ChatGPT.</p>
<h3>Comparing the number of themes</h3>
<p>As we’ve previously discussed, qualitative coding often reflects <a href="https://en.wikipedia.org/wiki/Lumpers_and_splitters">a fundamental tension</a> between lumpers (where coders group smaller concepts into broad themes) and splitters (where coders use more granular categories). In our initial study, we identified two distinct baseline behaviors: substantive theme granularity (the number of core categories containing more than one statement) and fringe-theme sensitivity (the willingness to create standalone categories for single, isolated comments).</p>
<p>In 2023, <strong>ChatGPT-4 generated almost half as many total themes as human coders</strong>, averaging seven total themes per dataset compared to thirteen for humans. However, this gap was driven almost entirely by fringe sensitivity. Human coders frequently created dedicated categories for one-off comments, accounting for nearly half of their total generated themes. Once those single-statement themes were removed, human coders and ChatGPT-4 converged at a similar macro baseline of six or seven substantive themes. Both operated as macro-lumpers for primary patterns.</p>
<p>Human coders were lumpers at the macro level, splitters at the margins. Human researchers built high overall category counts by isolating single-statement fringe themes and maintaining lower resolution on core patterns. In our 2026 re-run, ChatGPT-5.6’s approach to qualitative coding evolved, demonstrating a distinct shift towards aggressive splitting of substantive themes.</p>
<p>ChatGPT-5.6 closed the theme gap with human coders, generating a similar volume (15.6 for ChatGPT-5.6 versus 13.3 for humans) by mainly subdividing core themes into smaller buckets of two to four statements. On average, the model generated almost 70% more substantive themes than humans and nearly double that of ChatGPT-4. For example, where 2023 runs grouped feedback under a single “Cluttered” category, ChatGPT-5.6 split the node into two distinct subthemes: “Cluttered, busy, or overwhelming interface” and “Distracting ads, promotions, pop-ups, or graphics.” Table 2 shows the full counts of themes generated by each coder and AI; humans and ChatGPT-5.6 often had many more themes than the ChatGPT-4 runs. Table 3 shows that removing the one-off categories made the human coders’ major themes more comparable to ChatGPT-4, while ChatGPT-5.6 retained about twice as many themes (mainly due to categories with two to four statements).</p>

<table id="tablepress-1081" class="tablepress tablepress-id-1081 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<td class="column-1"></td><td class="column-2"></td><th class="column-3">Grocery</th><td class="column-4"></td><th class="column-5">Wireless</th><td class="column-6"></td><th class="column-7">Office Supplies</th><td class="column-8"></td>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1"></td><td class="column-2"></td><td class="column-3">Number of Themes</td><td class="column-4">Number of Themes with 1 statement</td><td class="column-5">Number of Themes</td><td class="column-6">Number of Themes with 1 statement</td><td class="column-7">Number of Themes</td><td class="column-8">Number of Themes with 1 statement</td>
</tr>
<tr class="row-3">
	<td class="column-1"></td><td class="column-2">Coder 1</td><td class="column-3">12</td><td class="column-4"> 6</td><td class="column-5">14</td><td class="column-6">6</td><td class="column-7">13</td><td class="column-8"> 5</td>
</tr>
<tr class="row-4">
	<td class="column-1"></td><td class="column-2">Coder 2</td><td class="column-3">25</td><td class="column-4">16</td><td class="column-5">16</td><td class="column-6">8</td><td class="column-7">17</td><td class="column-8">10</td>
</tr>
<tr class="row-5">
	<td class="column-1"></td><td class="column-2">Coder 3</td><td class="column-3"> 7</td><td class="column-4"> 0</td><td class="column-5">10</td><td class="column-6">4</td><td class="column-7"> 6</td><td class="column-8"> 1</td>
</tr>
<tr class="row-6">
	<td class="column-1"></td><td class="column-2"><strong>Coder Average</td><td class="column-3"><strong>14.7</td><td class="column-4"> <strong>7.3</td><td class="column-5"><strong>13.3</td><td class="column-6"><strong>6</td><td class="column-7"><strong>12</td><td class="column-8"> <strong>5.3</td>
</tr>
<tr class="row-7">
	<td rowspan="4" class="column-1">ChatGPT 4 2023<center></td><td class="column-2">ChatGPT A</td><td class="column-3">8</td><td class="column-4">1</td><td class="column-5">8</td><td class="column-6">2</td><td class="column-7">6</td><td class="column-8">0</td>
</tr>
<tr class="row-8">
	<td class="column-2">ChatGPT B</td><td class="column-3">6</td><td class="column-4">0</td><td class="column-5">7</td><td class="column-6">0</td><td class="column-7">6</td><td class="column-8">0</td>
</tr>
<tr class="row-9">
	<td class="column-2">ChatGPT C</td><td class="column-3">7</td><td class="column-4">0</td><td class="column-5">8</td><td class="column-6">1</td><td class="column-7">8</td><td class="column-8">2</td>
</tr>
<tr class="row-10">
	<td class="column-2"><strong>ChatGPT Average</td><td class="column-3"><strong>7.0</td><td class="column-4"><strong>0.3</td><td class="column-5"><strong>7.7</td><td class="column-6"><strong>1</td><td class="column-7"><strong>6.3</td><td class="column-8"><strong>0.7</td>
</tr>
<tr class="row-11">
	<td rowspan="4" class="column-1">ChatGPT 5.6 2026</td><td class="column-2">ChatGPT D</td><td class="column-3">22</td><td class="column-4"> 9</td><td class="column-5">14</td><td class="column-6">2</td><td class="column-7">14</td><td class="column-8">1</td>
</tr>
<tr class="row-12">
	<td class="column-2">ChatGPT E</td><td class="column-3">24</td><td class="column-4">12</td><td class="column-5">15</td><td class="column-6">2</td><td class="column-7">13</td><td class="column-8">3</td>
</tr>
<tr class="row-13">
	<td class="column-2">ChatGPT F</td><td class="column-3">14</td><td class="column-4"> 1</td><td class="column-5">13</td><td class="column-6">2</td><td class="column-7">11</td><td class="column-8">0</td>
</tr>
<tr class="row-14">
	<td class="column-2"><strong>ChatGPT Average</td><td class="column-3"><strong>20</td><td class="column-4"> <strong>7.3</td><td class="column-5"><strong>14</td><td class="column-6"><strong>2</td><td class="column-7"><strong>12.7</td><td class="column-8"><strong>1.3</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left"><strong>Table 2:</strong> Number of category themes generated by coder and ChatGPT runs.</p>

<table id="tablepress-1082" class="tablepress tablepress-id-1082 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<td class="column-1"></td><td class="column-2"></td><th class="column-3">Grocery Themes</th><th class="column-4">Wireless Themes</th><th class="column-5">Office Supplies Themes</th><th class="column-6">Average</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1"></td><td class="column-2">Coder 1</td><td class="column-3"> 6</td><td class="column-4"> 8</td><td class="column-5"> 8</td><td class="column-6"> 7.3</td>
</tr>
<tr class="row-3">
	<td class="column-1"></td><td class="column-2">Coder 2</td><td class="column-3"> 9</td><td class="column-4"> 8</td><td class="column-5"> 7</td><td class="column-6"> 8</td>
</tr>
<tr class="row-4">
	<td class="column-1"></td><td class="column-2">Coder 3</td><td class="column-3"> 7</td><td class="column-4"> 6</td><td class="column-5"> 5</td><td class="column-6"> 6</td>
</tr>
<tr class="row-5">
	<td class="column-1"></td><td class="column-2"><strong>Coder Average</td><td class="column-3"> <strong>7.3</td><td class="column-4"> <strong>7.3</td><td class="column-5"> <strong>6.7</td><td class="column-6"> <strong>7.1</td>
</tr>
<tr class="row-6">
	<td rowspan="4" class="column-1">ChatGPT 4 2023</td><td class="column-2">ChatGPT A</td><td class="column-3"> 7</td><td class="column-4"> 6</td><td class="column-5"> 6</td><td class="column-6"> 6.3</td>
</tr>
<tr class="row-7">
	<td class="column-2">ChatGPT B</td><td class="column-3"> 6</td><td class="column-4"> 7</td><td class="column-5"> 6</td><td class="column-6"> 6.3</td>
</tr>
<tr class="row-8">
	<td class="column-2">ChatGPT C</td><td class="column-3"> 7</td><td class="column-4"> 7</td><td class="column-5"> 6</td><td class="column-6"> 6.7</td>
</tr>
<tr class="row-9">
	<td class="column-2"><strong>ChatGPT Average</td><td class="column-3"> <strong>6.7</td><td class="column-4"> <strong>6.7</td><td class="column-5"> <strong>6</td><td class="column-6"> <strong>6.4</td>
</tr>
<tr class="row-10">
	<td rowspan="4" class="column-1">ChatGPT 5.6 2026</td><td class="column-2">ChatGPT D</td><td class="column-3">13</td><td class="column-4">12</td><td class="column-5">13</td><td class="column-6">12.7</td>
</tr>
<tr class="row-11">
	<td class="column-2">ChatGPT E</td><td class="column-3">12</td><td class="column-4">13</td><td class="column-5">10</td><td class="column-6">11.7</td>
</tr>
<tr class="row-12">
	<td class="column-2">ChatGPT F</td><td class="column-3">13</td><td class="column-4">11</td><td class="column-5">11</td><td class="column-6">11.7</td>
</tr>
<tr class="row-13">
	<td class="column-2"><strong>ChatGPT Average</td><td class="column-3"><strong>12.7</td><td class="column-4"><strong>12</td><td class="column-5"><strong>11.3</td><td class="column-6"><strong>12</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left"><strong>Table 3:</strong> Number of category themes by coder and ChatGPT containing more than one statement.</p>
<h3>Interrater agreement between humans and ChatGPT runs</h3>
<p>To evaluate how accurately ChatGPT-5.6 aligned with human judgment, and how reliable it was across repeated runs, we computed kappa for each pairing of raters, as we did in 2023. We <a href="https://measuringu.com/assessing-interrater-reliability-in-ux-research/">averaged kappas</a> across themes within products and then across products to get the overall results shown in Table 4.</p>
<p><strong><em>A brief reminder about kappa.</em></strong> There are different methods for assessing the magnitude of interrater agreement. One of the best-known is the kappa statistic (Fleiss, 1971). Kappa measures the extent of agreement among raters that exceeds estimates of chance agreement. Kappa values can be between −1 (perfect disagreement) and 1 (perfect agreement) and are often interpreted with the Landis and Koch guidelines (poor agreement: ≤ 0, slight: 0.01–0.20, fair: 0.21–0.40, moderate: 0.41–0.60, substantial: 0.61–0.80, almost perfect agreement: 0.81–1.00).</p>

<table id="tablepress-1083" class="tablepress tablepress-id-1083 tbody-has-connected-cells">
<tbody>
<tr class="row-1">
	<td class="column-1"></td><td class="column-2"></td><td class="column-3"></td><td class="column-4"></td><td colspan="3" class="column-5"><div style="text-align: center;"><strong>2023</strong></div></td><td colspan="2" class="column-8"><div style="text-align: center;"><strong>2026</strong></div></td>
</tr>
<tr class="row-2">
	<td class="column-1"><strong>Rater</td><td class="column-2"><strong>Coder 1</td><td class="column-3"><strong>Coder 2</td><td class="column-4"><strong>Coder 3</td><td class="column-5"><strong>ChatGPT A</td><td class="column-6"><strong>ChatGPT B</td><td class="column-7"><strong>ChatGPT C</td><td class="column-8"><strong>ChatGPT D</td><td class="column-9"><strong>ChatGPT E</td>
</tr>
<tr class="row-3">
	<td class="column-1">Coder 2</td><td class="column-2">0.703</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td><td class="column-7"></td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-4">
	<td class="column-1">Coder 3</td><td class="column-2">0.726</td><td class="column-3">0.683</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td><td class="column-7"></td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-5">
	<td class="column-1">ChatGPT A</td><td class="column-2">0.699</td><td class="column-3">0.608</td><td class="column-4">0.662</td><td class="column-5"></td><td class="column-6"></td><td class="column-7"></td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-6">
	<td class="column-1">ChatGPT B</td><td class="column-2">0.643</td><td class="column-3">0.548</td><td class="column-4">0.615</td><td class="column-5">0.684</td><td class="column-6"></td><td class="column-7"></td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-7">
	<td class="column-1">ChatGPT C</td><td class="column-2">0.681</td><td class="column-3">0.607</td><td class="column-4">0.624</td><td class="column-5">0.784</td><td class="column-6">0.584</td><td class="column-7"></td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-8">
	<td class="column-1">ChatGPT D</td><td class="column-2">0.756</td><td class="column-3">0.754</td><td class="column-4">0.734</td><td class="column-5">0.656</td><td class="column-6">0.610</td><td class="column-7">0.612</td><td class="column-8"></td><td class="column-9"></td>
</tr>
<tr class="row-9">
	<td class="column-1">ChatGPT E</td><td class="column-2">0.767</td><td class="column-3">0.738</td><td class="column-4">0.732</td><td class="column-5">0.665</td><td class="column-6">0.627</td><td class="column-7">0.642</td><td class="column-8">0.952</td><td class="column-9"></td>
</tr>
<tr class="row-10">
	<td class="column-1">ChatGPT F</td><td class="column-2">0.779</td><td class="column-3">0.744</td><td class="column-4">0.740</td><td class="column-5">0.654</td><td class="column-6">0.635</td><td class="column-7">0.643</td><td class="column-8">0.949</td><td class="column-9">0.949</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left"><strong>Table 4:</strong> Kappa matrix showing chance-corrected agreement for each pairing of human coders and ChatGPT runs, averaged over fourteen themes pulled from three datasets, expanded to include ChatGPT-5.6 runs (D, E, F).</p>
<p>Looking at the full pairwise matrix in Table 4, three clear trends emerge. First, ChatGPT-5.6 consistently achieved higher agreement with human coders than ChatGPT-4 did. Second, cross-generational agreement between GPT-4 and GPT-5.6 remained firmly substantial, confirming underlying thematic continuity despite the newer model’s increased granularity. Most strikingly, ChatGPT-5.6 demonstrated near-perfect self-consistency across independent runs, elevating intra-model agreement to a remarkable 0.949–0.952.</p>
<h3>Human Coders vs. ChatGPT Runs vs. Combined</h3>
<p>To test whether these shifts were statistically meaningful, we aggregated the pairwise kappas into group-level averages. Below are the average kappas for each pairing type, and these same values are shown visually in Figure 1.</p>
<ul>
<li><strong>Humans Coders:</strong> .704 (substantial) — unchanged from 2023</li>
<li><strong>ChatGPT-4:</strong> .684 (substantial) — unchanged from 2023</li>
<li><strong>ChatGPT-5.6:</strong> .950 (almost perfect)</li>
<li><strong>Humans <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2194.png" alt="↔" class="wp-smiley" style="height: 1em; max-height: 1em;" /> ChatGPT-4:</strong> .632 (substantial) — unchanged from 2023</li>
<li><strong>Humans <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2194.png" alt="↔" class="wp-smiley" style="height: 1em; max-height: 1em;" /> ChatGPT-5.6:</strong> .749 (substantial)</li>
<li><strong>GPT-4 <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2194.png" alt="↔" class="wp-smiley" style="height: 1em; max-height: 1em;" /> ChatGPT-5.6:</strong> .638 (substantial)</li>
</ul>
<p><a href="https://measuringu.com/wp-content/uploads/2026/09/Figure-1.png" rel="attachment wp-att-48688"><img decoding="async" class="alignnone wp-image-48688 size-full" src="https://measuringu.com/wp-content/uploads/2026/09/Figure-1.png" alt="Average agreement between ChatGPT and human coders." width="2157" height="1210" srcset="https://measuringu.com/wp-content/uploads/2026/09/Figure-1.png 2157w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-300x168.png 300w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-1024x574.png 1024w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-768x431.png 768w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-1536x862.png 1536w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-2048x1149.png 2048w, https://measuringu.com/wp-content/uploads/2026/09/Figure-1-600x337.png 600w" sizes="(max-width: 2157px) 100vw, 2157px" /></a></p>
<p class="wp-caption-text" style="text-align: left"><strong>Figure 1:</strong> Average agreement between ChatGPT and human coders.</p>
<p>There was a statistically significant difference among these group means (<em>F</em>(5,30) = 40.5, <em>p</em> &lt; .001). We used Tukey’s Honest Significant Difference test to look for specific differences while controlling for error. We found that ChatGPT-5.6 agreed with itself much more than any other pairing: its self-agreement was 0.20 to 0.31 greater than the other kappas (<em>p</em> &lt; .001). Human coders also agreed significantly more with ChatGPT-5.6 than they did with ChatGPT-4 (<em>p</em> &lt; .001). Crucially, human-to-ChatGPT-5.6 agreement (0.749) was statistically indistinguishable from human-to-human agreement (0.704, <em>p</em> = .475).</p>
<p>These findings suggest that ChatGPT has improved to a point where it’s highly consistent and agrees considerably well with humans. Consistency is, of course, good, but you can also be consistently wrong.</p>
<h3>No Issues vs. Other Themes</h3>
<p>Consistent with our 2023 findings, interrater agreement remained exceptionally high for the &#8220;No Issues&#8221; category, maintaining an almost perfect mean kappa of .937 across all three datasets. However, there were some cases where ChatGPT-5.6 called out unclear statements and pulled them out separately from the main analysis. This behavior is illustrated by how different runs of ChatGPT-5.6 evaluated two ambiguous participant responses in the Wireless dataset, Statement 16 (&#8220;Na&#8221;) and Statement 44 (&#8220;Not sure.&#8221;):</p>
<p><strong><u>Wireless – ChatGPT D:</u></strong></p>
<p>Here, the model created explicit reasoning notes, observing that &#8220;&#8216;Na&#8217; provides no usable information. It may mean &#8216;N/A,&#8217; but that is not explicit enough to confidently code as &#8216;no problems,’&#8221; and &#8220;&#8216;Not sure&#8217; does not identify a specific problem or indicate clearly that no problem occurred.&#8221;</p>
<p><strong><u>Wireless – ChatGPT E:</u></strong></p>
<p>During this run, the model created a separate “Uncategorized Statements” bucket, designating statement 44 as an off-topic/service issue, while identifying statement 16 as a non-substantive response.</p>
<p><strong><u>Wireless – ChatGPT F:</u></strong></p>
<p>In this case, the model created a dedicated category titled “Statements without enough information to classify,” grouping 16 and 44 together because they lacked substantive feedback.</p>
<p>Humans often included these statements in the &#8220;No Issues&#8221; theme. Depending on your lens of analysis, either classification could be justified. Subtle shifts in classification like this could lead to lower agreement. Fortunately, this type of model variance can likely be mitigated with a more explicit prompt instruction regarding edge-case handling.</p>
<h2>Summary and Discussion</h2>
<p>Our key conclusions from these analyses were:</p>
<p><strong>Agreement with human coders improved with the newer ChatGPT-5.6 model; human-to-ChatGPT-5.6 agreement was as high as human-to-human. </strong>Average interrater agreement increased, with kappa going from .632 to .749 when comparing ChatGPT to human coders. While ChatGPT-4 already fell within the established guidelines for “substantial” agreement, this improvement suggests that using ChatGPT-5.6 for coding open-ended survey responses is now virtually as reliable as comparing multiple human researchers.</p>
<p><strong>ChatGPT-5.6 responses seem to be dramatically more consistent in this setting. </strong>In addition to achieving better agreement with human coders, ChatGPT-5.6 exhibited remarkable agreement across repeated runs. In contrast to our 2023 testing, running open-ended comments through ChatGPT-5.6 is now sufficiently consistent that a single analysis run will likely yield the same taxonomic output.</p>
<p><strong>ChatGPT-5.6 responses may split categories into further detail. </strong>In this setting, ChatGPT-5.6 demonstrated the ability to construct more precise subcategories while maintaining structural alignment with the overarching core themes identified by humans.</p>
<h3>Caveats</h3>
<p><strong>Prompt frozen at 2023 for comparability. </strong>To isolate model performance over time, we controlled prompt structure and data inputs. However, optimized prompting techniques customized for newer reasoning architectures could likely yield even stronger performance.</p>
<p><strong>Alignment method changed from 2023. </strong>To align outputs with pre-existing categories, we shifted to a top-down alignment strategy. Because ChatGPT-5.6 frequently split baseline categories into narrower subthemes, those subcategories were conceptually combined for cross-model evaluation.</p>
<p><strong>Small datasets, one prompt, three runs, one model. </strong>In 2023, LLMs struggled with high-input tasks due to limited context windows, forcing us to rely on compact datasets (50–52 statements). To preserve comparability, we retained these small datasets, tested a single prompt across three runs, and limited this initial phase to OpenAI models. Keep this in mind when generalizing these findings to open-ends in different contexts.</p>
<h3>What’s Next?</h3>
<p>In this study, we tested a fairly straightforward replication of our original analysis. Given the large improvement in how ChatGPT performed, further expansions are justified. We plan to continue to explore and test ideas like:</p>
<ul>
<li>How can we improve our prompt to maximize accuracy and ease of use?</li>
<li>How do Claude, Gemini, Meta AI, or other models perform on qualitative coding?</li>
<li>How do LLMs perform on larger datasets now that context windows have improved?</li>
</ul>
<h2>Appendix: The Verbatim Prompt</h2>
<p>As a UX researcher, you are tasked with analyzing the dataset provided below, which contains answers to the question, “What are some problems or frustrations you’ve had with the XXXX website?”. Your goal is to classify each numbered statement according to common themes. Create as many categories as necessary to group similar statements together. If a statement fits multiple categories, include it in all relevant categories.</p>
<p>START DATASET:</p>
<p>[dataset numbered by participant]</p>
<p>END DATASET</p>
<p>After analyzing the dataset, follow these steps:</p>
<ol>
<li>List the categories you have created, along with a brief description for each category.</li>
<li>For each category, list the numbers of the statements that belong to it.</li>
<li>If there are any statements that do not fit into any of the categories you have created, list their numbers separately.<strong><br />
</strong></li>
</ol>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>How Well Do Digital Twins Predict SUS Scores?</title>
		<link>https://measuringu.com/how-well-do-synthetic-users-predict-sus-scores/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=how-well-do-synthetic-users-predict-sus-scores</link>
		
		<dc:creator><![CDATA[Lucas Plabst, PhD • Jeff Sauro, PhD • Jim Lewis, PhD]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 01:08:02 +0000</pubDate>
				<category><![CDATA[Survey]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[SUS]]></category>
		<category><![CDATA[Synthetic user]]></category>
		<category><![CDATA[Synthetic users]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48634</guid>

					<description><![CDATA[Participant recruitment is one of the most costly (and time-consuming) withdrawals from a research budget. If AI can generate findings by learning about your users, it could save a lot of money and generate study results in milliseconds instead of months. Synthetic users are the broad name to describe AI’s role in mimicking participants’ attitudes [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1.jpg"><img decoding="async" class="alignleft wp-image-48658 size-medium" src="https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-300x169.jpg" alt="Feature image" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/09/092226-FeatureImage1.jpg 2000w" sizes="(max-width: 300px) 100vw, 300px" /></a>Participant recruitment is one of the most costly (and time-consuming) withdrawals from a research budget. If AI can generate findings by learning about your users, it could save a lot of money and generate study results in milliseconds instead of months.</p>
<p><em>Synthetic users</em> are the broad name to describe AI’s role in mimicking participants’ attitudes and intentions. But the term is used rather loosely. We’ve <a href="https://measuringu.com/what-are-the-different-types-of-synthetic-users/">earlier defined synthetic users</a> as more of an umbrella term (like a genus) than a type (like a species). The genus encompasses digital twins, research-grounded, persona-based, and demographic-based synthetic users, and AI proto personas.</p>
<p>We’ve written about pro-synthetic and anti-synthetic user attitudes in the UX community, and we <a href="https://measuringu.com/review-of-experiments-with-synthetic-users/">published a literature review of experiments with synthetic users</a>. Much of the literature has reported numerous discrepancies between synthetic and human results, including reduced variance, misalignment of means/percentages, distorted correlations, inaccurate regression coefficients, and shallow experiential narratives. Of the different types of synthetic users, digital twins are widely believed to be more accurate than other types of synthetic user despite their accuracy being <a href="https://www.science.org/doi/10.1126/sciadv.aeh8260">an open research question</a>.</p>
<p>But what can be helpful is an understanding of how effective synthetic users could be in the context of UX research. UX research has <a href="https://measuringu.com/taxonomy-ux-research-methods/">numerous methods</a> and <a href="https://measuringu.com/what-are-ux-deliverables/">deliverables</a>. One of those methods, the <a href="https://measuringu.com/benchmark-intro/">retrospective UX benchmark study</a>, uses the same <strong>attitudinal</strong> UX metrics as task-based UX testing (such as the SUS, SUPR-Q<sup>®</sup>, or UX-Lite<sup>®</sup>), but participants are asked to reflect on their prior experience (as opposed to being presented with tasks), and they answer open-ended questions about their experiences. This sort of study looks a lot like a typical market-research survey, where you collect participant demographic details, attitudes, and intentions. We collect a lot of data with this method, making it a good candidate on which to test the promise of synthetic users.</p>
<p>In this article, we describe our experiment comparing quantitative results of one of our industry benchmark surveys with data generated by two types of digital twins.</p>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2><span lang="EN">The Original Study</span></h2>
<p>In May 2026, we conducted a <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">retrospective benchmark study of four AI-based chat software products</a> with 420 U.S.-based panel participants. This study included the metrics we typically collect in our <a href="https://measuringu.com/consumer-software-ux-2025/">standard UX and NPS study of consumer software</a>.</p>
<p>There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several questionnaires, including the <a href="https://measuringu.com/nps-ux/">NPS</a>, <a href="https://measuringu.com/10-things-sus/">SUS</a>, <a href="https://measuringu.com/from-umux-lite-to-ux-lite/">UX-Lite</a>, and <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" />. The AI-based chat products and sample sizes were:</p>
<p>● ChatGPT: 113<br />
● Claude: 103<br />
● Gemini: 101<br />
● Grok: 103</p>
<h2><span lang="EN">Replication with Digital Twins</span></h2>
<p>For these experiments, we assessed how well synthetic users could reproduce one of the key metrics collected in the retrospective study, the System Usability Scale (<a href="https://measuringu.com/10-things-sus/">SUS</a>). The SUS is one of the most widely used standardized questionnaires for assessing perceived usability, made up of ten five-point items varying equally in positive and negative tone. The composite SUS score is interpolated to a 0–100-point scale for which there are <a href="https://www.researchgate.net/publication/324116412_The_System_Usability_Scale_Past_Present_and_Future">well-known methods</a> for classifying scores as good or poor, including the Sauro–Lewis Curved Grading Scale (see the <a href="#_The_Sauro-Lewis">Appendix</a>).</p>
<p>We used Claude Sonnet 5 to construct two types of <strong>digital twins</strong> based on each person&#8217;s responses to the original survey (<em>n</em> = 420). We fed the model each participant’s survey responses that included questions on demographics, chatbot use, brand attitude, likelihood to recommend, and open-ended comments about what they like and dislike about chatbots. Most importantly, <strong>we left out the participants&#8217; SUS scores</strong>.</p>
<p>Claude constructed a 700-word max persona written in the first person in the participants’ voice explaining who they are, based on the “summary agent” approach of Park et al. (<a href="https://arxiv.org/abs/2411.10109">2024</a>).</p>
<p>The persona was made up of six sections:</p>
<ol>
<li>Who I am</li>
<li>What matters to me</li>
<li>How I talk</li>
<li>How I decide and answer</li>
<li>What I would have to guess</li>
<li>What is not known or inferred about me</li>
</ol>
<p>We built the synthetic users with two levels of fidelity: summary twins or full twins. In <a href="https://measuringu.com/what-are-the-different-types-of-synthetic-users/">our taxonomy</a>, both of these types of synthetic users are digital twins because they were derived from individual respondents’ data. Summary twins were given the 700-word persona summary (no direct access to any of the original survey responses); the full twins were supplied with all the participants&#8217; survey responses (except for their SUS ratings) in addition to the summary. We did this to see whether giving a digital twin access to more specific information about a person would lead to better predictions of the SUS.</p>
<p>Each twin was prompted to role-play as its participant and to respond to the SUS as its human would have. We then matched each twin with its human counterpart and analyzed the differences in SUS scores.</p>
<h3><span lang="EN">Results</span></h3>
<h4><span lang="EN">Digital twin means were systematically higher than human means</span></h4>
<p>Both the full and summary twins scored statistically significantly higher SUS scores than the human responses (all <em>p</em> &lt; .05), ranging from about 2.9 to 7.8 points higher when computed by product (Table 1 and Figure 1). The differences by product between full and summary twins were not statistically significant (all <em>p</em> &gt; .36). In other words, the twins were more like each other than they were like the humans on which they were based.</p>

<table id="tablepress-1077" class="tablepress tablepress-id-1077">
<thead>
<tr class="row-1">
	<th class="column-1">Chatbot</th><th class="column-2">Human Mean</th><th class="column-3">Summary Twin Mean</th><th class="column-4">Full Twin Mean</th><th class="column-5">Summary – Human</th><th class="column-6">Full – Human</th><th class="column-7">Full – Summary</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">ChatGPT</td><td class="column-2">81.5</td><td class="column-3">85.7</td><td class="column-4">88.3</td><td class="column-5">4.2</td><td class="column-6">6.8</td><td class="column-7">2.6</td>
</tr>
<tr class="row-3">
	<td class="column-1">Claude</td><td class="column-2">78.9</td><td class="column-3">84.6</td><td class="column-4">86.2</td><td class="column-5">5.7</td><td class="column-6">7.3</td><td class="column-7">1.6</td>
</tr>
<tr class="row-4">
	<td class="column-1">Gemini</td><td class="column-2">79.0</td><td class="column-3">81.9</td><td class="column-4">84.2</td><td class="column-5">2.9</td><td class="column-6">5.2</td><td class="column-7">2.3</td>
</tr>
<tr class="row-5">
	<td class="column-1">Grok</td><td class="column-2">78.4</td><td class="column-3">83.2</td><td class="column-4">86.2</td><td class="column-5">4.8</td><td class="column-6">7.8</td><td class="column-7">3.0</td>
</tr>
<tr class="row-6">
	<td class="column-1">Average</td><td class="column-2">79.5</td><td class="column-3">83.9</td><td class="column-4">86.2</td><td class="column-5">4.4</td><td class="column-6">6.7</td><td class="column-7">2.3</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Comparison of means and mean differences by product.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092226f1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48664" src="https://measuringu.com/wp-content/uploads/2026/09/092226f1-300x153.jpg" alt="Graph of mean SUS by product and source with 95% confidence intervals." width="882" height="450" srcset="https://measuringu.com/wp-content/uploads/2026/09/092226f1-300x153.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092226f1-1024x522.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/092226f1-768x392.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092226f1-1536x783.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/092226f1-2048x1045.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/09/092226f1-600x306.jpg 600w" sizes="auto, (max-width: 882px) 100vw, 882px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Graph of mean SUS by product and source with 95% confidence intervals.</p>
<h4><span lang="EN">More information led to more discrepancy in means</span></h4>
<p>Contrary to our expectation, giving the full twins <strong>more</strong> evidence (the actual responses) on which to base their responses did not lead to better predictions. The summary twin SUS scores differed from human scores by 4.4 on average, while the full twin scores differed from human scores by 6.7 points, both statistically significant differences. But is that difference a lot?</p>
<p>There are a few ways to interpret it. First, Table 1 shows that that number of points represents a 4.4% to 6.7% difference on the SUS’s 100-point scale, which can be thought of as an average error. Or those could be rephrased as 95.6% and 93.3% accurate predictions, which doesn’t sound that bad as long as the decisions that need to be made with the data can tolerate that level of precision. It’s also possible that if products have lower SUS scores, then the accuracy of synthetic responses would be lower (something we plan to follow up on).</p>
<p>A second way is to convert the scores into grades. Table 2 shows the SUS grades (see the Appendix for the grading scale) for the means by product and source. The predicted SUS grade was off by at least half a letter grade for all products and in one case off by a full letter grade (Grok).</p>
<p>This illustrates how the upward shift in scores for digital twins distorted the typical interpretation of the SUS, especially for the full twins. A stakeholder looking at these human grades would interpret the human scores as pretty good but not fantastic. The same stakeholder looking at the full twin grades would think the products are all superb.</p>

<table id="tablepress-1078" class="tablepress tablepress-id-1078">
<thead>
<tr class="row-1">
	<th class="column-1">Chatbot</th><th class="column-2">Human</th><th class="column-3">Summary Twin</th><th class="column-4">Full Twin</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">ChatGPT</td><td class="column-2">A</td><td class="column-3">A+</td><td class="column-4">A+</td>
</tr>
<tr class="row-3">
	<td class="column-1">Claude</td><td class="column-2">A−</td><td class="column-3">A+</td><td class="column-4">A+</td>
</tr>
<tr class="row-4">
	<td class="column-1">Gemini</td><td class="column-2">A−</td><td class="column-3">A</td><td class="column-4">A+</td>
</tr>
<tr class="row-5">
	<td class="column-1">Grok</td><td class="column-2">B+</td><td class="column-3">A</td><td class="column-4">A+</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 2:</strong> SUS grades according to Sauro–Lewis Curved Grading Scale (see the Appendix).</p>
<h4><span lang="EN">Small but consistent distortion at the item level adds up</span></h4>
<p>To understand what’s driving the differences in scores, we looked at responses to the ten individual SUS items. Both twin models agreed more with the positive-tone items and disagreed more with the negative-tone items (Figure 2). This could be an artifact of the way that LLMs are trained, with a tendency to be overly agreeable. The mean response to each SUS item (five-point Likert scale) was off by an average of about half a point. Because the distortion was systematic rather than random, the deviations added up across the ten items rather than canceling out.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092226f2.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48661 size-full" src="https://measuringu.com/wp-content/uploads/2026/09/092226f2.jpg" alt="Mean scores for SUS items by source with 95% confidence intervals (n = 420)." width="980" height="577" srcset="https://measuringu.com/wp-content/uploads/2026/09/092226f2.jpg 980w, https://measuringu.com/wp-content/uploads/2026/09/092226f2-300x177.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092226f2-768x452.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092226f2-600x353.jpg 600w" sizes="auto, (max-width: 980px) 100vw, 980px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Mean scores for SUS items by source with 95% confidence intervals (<em>n</em> = 420).</p>
<h4><span lang="EN">Human-twin correlations were significant but far from perfect</span></h4>
<p>Figure 3 shows the correlations (with 95% confidence intervals) between the human and digital twin SUS scores. The human-twin correlations were lower than you would expect if the digital twins were faithfully reproducing the human scores (although significantly higher than the correlation of .20 reported by Peng et al., <a href="https://www.science.org/doi/10.1126/sciadv.aeh8260">2026</a>). The correlation between the summary and full twins was very high. The difference of .03 in the correlations of summary and full twins with human data was not statistically significant.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092226f3-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48662" src="https://measuringu.com/wp-content/uploads/2026/09/092226f3-300x186.jpg" alt="Correlations between human and digital twin SUS scores (95% confidence intervals)." width="725" height="450" srcset="https://measuringu.com/wp-content/uploads/2026/09/092226f3-300x186.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092226f3-1024x636.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/092226f3-768x477.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092226f3-1536x954.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/092226f3-2048x1272.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/09/092226f3-600x373.jpg 600w" sizes="auto, (max-width: 725px) 100vw, 725px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> Correlations between human and digital twin SUS scores (95% confidence intervals).</p>
<h4><span lang="EN">Digital twin data was less variable than human data</span></h4>
<p>Figure 4 shows compression of the variability of the digital twin responses relative to the human responses, with the twins’ responses tending to converge more tightly around the mean. Summary twin means ran 2.9 to 5.7 points above human means; full twins ran 5.2 to 7.7 points above. The standard deviations of the individual ratings were 15.1 for humans, 12.6 for summary twins, and 11.6 for full twins.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/09/092226f4.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48663 size-full" src="https://measuringu.com/wp-content/uploads/2026/09/092226f4.jpg" alt="Distributions of SUS scores by product for Human, Summary Twin, and Full Twin." width="1016" height="704" srcset="https://measuringu.com/wp-content/uploads/2026/09/092226f4.jpg 1016w, https://measuringu.com/wp-content/uploads/2026/09/092226f4-300x208.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/092226f4-768x532.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/092226f4-600x416.jpg 600w" sizes="auto, (max-width: 1016px) 100vw, 1016px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> Distributions of SUS scores by product for Human, Summary Twin, and Full Twin (dots were spread horizontally to separate ties).</p>
<h2><span lang="EN">Summary and Discussion</span></h2>
<p>We generated digital twins by creating summaries based on data collected from 420 human respondents to a retrospective study on attitudes toward four AI products: ChatGPT, Claude, Gemini, and Grok. Two types of digital twins were created for each respondent: one based only on the summary (summary twin) and one based on the summary plus all the specific data used to create the summary (full twin). Because the goal of this research was to compare human with digital twin SUS scores, the human responses to the SUS were withheld from the data used to create the summaries and from the specific data provided to the full twins.</p>
<p>Our key findings were:</p>
<p><strong>Digital twin means were reasonably accurate, but systematically higher than human means. </strong>Averaging across the four products, SUS scores for summary twins were 4.4 points higher than human means (95.6% accuracy); full twin scores were 6.7 points higher (93.3% accuracy). We found this pattern to be statistically significant for all four products, and it was driven by systematic acquiescence by the digital twins at the item level.</p>
<p>Whether these levels of accuracy are good enough largely depends on the research context. We are cautious about generalizing these results beyond the fairly high SUS scores in this one study. It’s possible that because the SUS scores were already very high, a ceiling effect pushed the digital twin means closer to the human means than would be the case if the human SUS scores were lower.</p>
<p><strong>Human responses were more variable than the digital twins.</strong> As reported in the literature, we also saw lower variability of the digital twin–generated SUS scores. The standard deviations of the individual ratings were 15.1 for humans, 12.6 for summary twins, and 11.6 for full twins.</p>
<p><strong>Having more information did not improve prediction accuracy.</strong> We expected the means for the full twins to correspond closer than the summary twins to the human data. That didn’t happen. At the product level, the means of the full twins were on average about 2.2 points higher than the summary twins (and even farther from the human means).</p>
<p><strong>The inflation of SUS scores affected their interpretation.</strong> At higher SUS levels, the deviations of three to eight points across products (the range from Table 1 across both types of digital twins) are enough to change the products’ grades. The grades for the human scores ranged from B+ to A: good but not fantastic. For the full twins, the grades were all A+, leading to a very different interpretation of the perceived usability of the products.</p>
<p><strong>Correlations between human and digital twin scores were significant but far from perfect.</strong> The correlations were 0.56 for human/summary twins and 0.59 for human/full twins (no significant difference between these two). The correlation between summary and full twins was 0.92 (significantly higher than their correlations with the human data).</p>
<p><strong>Bottom line:</strong> We designed these digital twins to give them the best possible chance to accurately estimate held-out SUS scores relative to other types of synthetic users that are less grounded in specific human data. We were frankly surprised that deviations from human scores (ground truth) were consistently greater for the full twins (more information) than for the summary twins (less information). We were also surprised that this approach generated SUS scores that were within a few points of the human data.</p>
<p>This is the first in a series of similar experiments that we plan to conduct and report using different datasets (with more varied SUS scores) and different ways of creating synthetic users. Stay tuned …</p>
<h2 id="_The_Sauro-Lewis"><span lang="EN">Appendix: The Sauro-Lewis Curved Grading Scale</span></h2>

<table id="tablepress-964" class="tablepress tablepress-id-964">
<thead>
<tr class="row-1">
	<th class="column-1">SUS Score Range</th><th class="column-2">Grade</th><th class="column-3">Percentile Range</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">84.1–100</td><td class="column-2">A+</td><td class="column-3">96–100</td>
</tr>
<tr class="row-3">
	<td class="column-1">80.8–84.0</td><td class="column-2">A</td><td class="column-3">90–95</td>
</tr>
<tr class="row-4">
	<td class="column-1">78.9–80.7</td><td class="column-2">A−</td><td class="column-3">85–89</td>
</tr>
<tr class="row-5">
	<td class="column-1">77.2–78.8</td><td class="column-2">B+</td><td class="column-3">80–84</td>
</tr>
<tr class="row-6">
	<td class="column-1">74.1–77.1</td><td class="column-2">B</td><td class="column-3">70–79</td>
</tr>
<tr class="row-7">
	<td class="column-1">72.6–74.0</td><td class="column-2">B−</td><td class="column-3">65–69</td>
</tr>
<tr class="row-8">
	<td class="column-1">71.1–72.5</td><td class="column-2">C+</td><td class="column-3">60–64</td>
</tr>
<tr class="row-9">
	<td class="column-1">65.0–71.0</td><td class="column-2">C</td><td class="column-3">41–59</td>
</tr>
<tr class="row-10">
	<td class="column-1">62.7–64.9</td><td class="column-2">C−</td><td class="column-3">35–40</td>
</tr>
<tr class="row-11">
	<td class="column-1">51.7–62.6</td><td class="column-2">D</td><td class="column-3">15–34</td>
</tr>
<tr class="row-12">
	<td class="column-1"> 0.0–51.6</td><td class="column-2">F</td><td class="column-3"> 0–14</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 1:</strong> The Sauro-Lewis curved grading scale for the SUS.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Do Participants Say Less to AI Than to Human Moderators?</title>
		<link>https://measuringu.com/do-participants-say-less-to-ai-than-to-human-moderators/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=do-participants-say-less-to-ai-than-to-human-moderators</link>
		
		<dc:creator><![CDATA[Lucas Plabst, PhD • Jeff Sauro, PhD • Eva Sundberg, MPH • Jim Lewis, PhD]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 21:08:23 +0000</pubDate>
				<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Moderator]]></category>
		<category><![CDATA[moderated research]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48604</guid>

					<description><![CDATA[&#8220;Are AI moderators as good as human moderators?&#8221; That&#8217;s a good question. It&#8217;d be good to have data to help answer that. Instead, we have claims from vendors selling AI moderator technology. That&#8217;s not nothing, but it&#8217;s hardly the objective data you&#8217;d want before you retire your research team from moderating. When we encounter interesting [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48621 size-medium" src="https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-300x169.jpg" alt="Feature image" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/09/091526-FeatureImage1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>&#8220;Are AI moderators as good as human moderators?&#8221;</p>
<p>That&#8217;s a good question. It&#8217;d be good to have data to help answer that. Instead, we have claims from vendors selling AI moderator technology. That&#8217;s not nothing, but it&#8217;s hardly the objective data you&#8217;d want before you retire your research team from moderating.</p>
<p>When we encounter interesting claims like that, we generally take the following approach:</p>
<ol>
<li>Define the claim.</li>
<li>Assess evidence and the quality (focusing on peer-reviewed publications).</li>
<li>Conduct our own experiment.</li>
<li>Refine and revise our opinion of the merits of the claim.</li>
</ol>
<p>We completed step two in <a href="https://measuringu.com/effectiveness-of-ai-moderators-a-literature-review">our previous article</a>, where we reviewed the published literature. There wasn’t much literature to review, and some of the findings may be a bit dated. But even newer work is lacking because most studies compared AI moderators to surveys or static follow-up questions, not to a live human moderator. Only two studies actually put an AI moderator head-to-head against a human one.</p>
<p>One was a small classroom-style study (Wuttke et al., <a href="https://aclanthology.org/2025.latechclfl-1.17.pdf">2025</a>). Student pairs each ran one AI-led and one human-led interview on politics and democracy, netting just five AI-led and six human-led sessions, with students standing in for trained moderators.</p>
<p>The other (Zhu et al., <a href="https://dl.acm.org/doi/pdf/10.1145/3772318.3791653">2026</a>) was a properly powered randomized control trial (RCT) with 30 sessions per condition, pitting an agentic voice moderator against a human moderator in think-aloud usability testing of a note-taking app. It&#8217;s a stronger study, but think-aloud moderation is a lower-demand job than in-depth-interview moderation. In a think-aloud session, the moderator&#8217;s main task is prompting someone to keep narrating what they&#8217;re already doing. It&#8217;s not that far removed from unmoderated think-aloud, where there&#8217;s no moderator at all, just a periodic &#8220;keep talking&#8221; nudge or “tell me more.” A semi-structured or in-depth interview asks much more of a moderator: <a href="https://measuringu.com/five-ways-to-use-rapport-when-moderating/">building rapport</a> from a cold start, deciding in real time whether an offhand comment is worth five extra minutes, and steering a conversation that has no task to anchor it.</p>
<p>In the Zhu et al. study, they found no difference in the number of words participants produced, but the <strong>AI</strong> <strong>moderator talked more, faster, and more often </strong>during the session itself. The AI moderator also scored significantly worse than the human moderator on context-aware follow-up and on building trust, even though the two were indistinguishable on procedural adherence.</p>
<p>That review left us with five questions that the literature couldn&#8217;t settle:</p>
<ul>
<li>Do participants say less to AI than to human moderators?</li>
<li>Do participants provide the same depth of insights to AI and human moderators?</li>
<li>Does AI moderate as competently as a human?</li>
<li>Do participants want to talk to AI moderators?</li>
<li>Do findings generalize across models or prompts?</li>
</ul>
<p>Each of those questions needs its own analysis, and all require more data. To help us (and the industry) answer these questions, we built our own AI moderator and ran a controlled experiment. In this article, we’ll cover the setup of that experiment and address the first question of whether participants say less, more, or about the same to AI moderators as human moderators. Of course, speaking more isn’t necessarily better, but if humans systematically say LESS to an AI moderator, that would create fewer opportunities for uncovering insights.</p>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2>Experimental Study Setup</h2>
<p>To test an AI moderator, you need a few things: an AI moderator, human moderators, a study goal and discussion guide, and human participants. The last three are staples of UX research; AI moderators are the new additions.</p>
<h3>AI Moderator</h3>
<p>There are AI moderators on the market, but we wanted to build our own for two reasons. First, we wanted to understand how prompting and programming decisions affect the output around probing decisions, especially as models change. Second, we wanted the AI moderator to look as realistic as possible (and uncannily similar to the real human moderators in the study). We built our AI moderator using an AI agent and created photorealistic avatars with voices of two of our researchers (see Video 1).</p>
<div class="ast-oembed-container " style="height: 100%;"><iframe loading="lazy" title="AI-vs-Humans" src="https://player.vimeo.com/video/1227406419?h=ed3f31be28&amp;dnt=1&amp;app_id=122963" width="1170" height="658" frameborder="0" allow="autoplay; fullscreen; picture-in-picture; clipboard-write; encrypted-media; web-share" referrerpolicy="strict-origin-when-cross-origin"></iframe></div>
<p class="wp-caption-text" style="text-align: left;"><strong>Video 1:</strong> A quick comparison between our AI and human moderators.</p>
<p>The AI moderator used Gemini 3 Flash. To minimize latency and the risk of hallucinations due to the context window overflowing, the AI moderator was made up of several agents working sequentially. Each agent covered a broad topic, summarized its findings, and then handed those findings off to the next agent. This happened behind the scenes, so to participants, it seemed like one fluid conversation with latency times about 400ms longer than human moderator latency.</p>
<h2>Research Topic and Discussion Guide: Sleep Tech</h2>
<p>In the study, instead of asking a simple set of questions (like asking people about their AI usage), we wanted to simulate a realistic exploratory study that would benefit from an in-depth interview. We generated a study for a fictitious company looking to build a product to help with waking up and sleep management. We selected sleep tech as the research topic because it was both realistic (there are many <a href="https://www.forbes.com/sites/forbes-personal-shopper/article/best-sleep-tech/">sleep tech products</a>) and relatively easy to find recruits because, well, everyone sleeps.</p>
<p>We built discussion and moderator probing guides (Appendix A) that generated questions to help understand people’s sleep patterns, current pain points, waking habits, and any techniques and tech they may use.  The goal of the research was to find current frustrations with people’s wake-up methods and to identify potential new product opportunities for waking up. For details, see the appendix.</p>
<h3>Recruitment and Screener</h3>
<p>Twenty participants were recruited using a convenience sample from the Denver metro area to participate in a 30-minute in-person interview. Before the interview, participants completed a screener covering topics like caffeine use, melatonin intake, sleep accessories, and living arrangements (e.g., having kids or pets) that may affect sleep. We also captured AI sentiment (favorable or unfavorable) so we could assess and evenly balance AI attitudes across conditions. If the AI moderator underperformed, it wouldn&#8217;t be because we stacked the AI-moderated group with AI skeptics. Both human and AI moderators were instructed to probe beyond the discussion guide whenever it served the research goals.</p>
<p>Unlike a typical study, neither human nor AI moderators had access to the screener data beforehand. We withheld that information to see how well the moderators would be able to uncover prior established facts. This would become a key quantitative measure of probing effectiveness that we’ll cover in a future article.</p>
<h3>Participant Assignment: 2 × 2 Factorial Design</h3>
<p>We randomly assigned the 20 participants to a human moderator or AI moderator condition, crossed with moderator gender. This was a full 2 × 2 between-subjects design with five participants per cell. Two MeasuringU researchers, one female and one male, ran the human sessions. To keep the conditions as similar as possible, the AI moderators were built using the voices and likenesses of the human moderators, resulting in ten AI moderated sessions and ten human moderated sessions.</p>
<p>Participants were greeted upon arrival at our Denver labs and escorted by a different MeasuringU researcher who did not participate in the sessions. Participants didn’t know ahead of time if they were to be paired with a human or AI moderator. Participants sat in our lab and the moderators connected from the other side of the one-way mirror, avoiding direct human contact with the participants. Participants were connected to a video call with either the AI or the human moderator.</p>
<p>The 20 participants split evenly across the two conditions (ten AI-moderated, ten human-moderated). Gender skewed slightly male in both groups (7/3 in the AI condition, 6/4 in the human condition) and ages spanned the same 25–65+ range in each. Attitude toward AI (Enthusiast, Neutral, or Skeptic) was reasonably balanced across conditions, so a weaker AI-moderator showing couldn&#8217;t be chalked up to one group being stacked with AI skeptics.</p>

<table id="tablepress-1075" class="tablepress tablepress-id-1075">
<thead>
<tr class="row-1">
	<th class="column-1">Attitude toward AI</th><th class="column-2"><center>AI moderator (<i>n</i> = 10)</th><th class="column-3"><center>Human moderator (<i>n</i> = 10)</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">Enthusiast</td><td class="column-2"><center>4</td><td class="column-3"><center>6</td>
</tr>
<tr class="row-3">
	<td class="column-1">Neutral</td><td class="column-2"><center>3</td><td class="column-3"><center>2</td>
</tr>
<tr class="row-4">
	<td class="column-1">Skeptic</td><td class="column-2"><center>3</td><td class="column-3"><center>2</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Attitude toward AI by moderator condition (<em>n</em> = 20).</p>
<h2>Study Results</h2>
<p>So, do participants say less to AI than to human moderators?</p>
<p>In short, <strong>yes</strong>. And not just a little less, <strong>a lot less</strong>. We used several methods to assess how much participants spoke to AI versus human moderators (Table 2).</p>

<table id="tablepress-1076" class="tablepress tablepress-id-1076">
<thead>
<tr class="row-1">
	<td class="column-1"></td><th class="column-2">AI</th><th class="column-3">Human</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">Session Duration (Min)</td><td class="column-2">14</td><td class="column-3">22</td>
</tr>
<tr class="row-3">
	<td class="column-1">Moderator Words Per Minute (WPM)</td><td class="column-2">99</td><td class="column-3">55</td>
</tr>
<tr class="row-4">
	<td class="column-1">Participant Words Per Minute (WPM)</td><td class="column-2">54</td><td class="column-3">98</td>
</tr>
<tr class="row-5">
	<td class="column-1">Participant % Share of Talking</td><td class="column-2">34%</td><td class="column-3">64%</td>
</tr>
<tr class="row-6">
	<td class="column-1">Participant to Moderator Word Ratio</td><td class="column-2">0.5</td><td class="column-3">1.8</td>
</tr>
<tr class="row-7">
	<td class="column-1">Delay in Responding (seconds)</td><td class="column-2">1.9</td><td class="column-3">1.5</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 2:</strong> Comparison of AI and human results.</p>
<p><strong>AI Moderated Sessions Were 37% Shorter</strong></p>
<p>Participants were scheduled for 30-minute slots, but as is typical with research sessions, that 30 minutes included set-up time and some buffer time. Once the interview started with the AI or human moderators, participants were told the interview would last about 20 minutes.</p>
<p>Human sessions ran on average 22 minutes, whereas AI-moderated sessions lasted only 14 minutes (37% shorter). Or put another way, participants spent 57% more time with human moderators than AI moderators despite having the same script and planned session duration.</p>
<p><strong>AI Moderators Spoke 80% More Than Human Moderators</strong></p>
<p>Because average session duration differed a lot between human and AI moderators, we normalized the words spoken by both moderators and participants by dividing the total number of words spoken per session by the total session minutes to derive words spoken per minute (WPM). Human moderators spoke on average 55 WPM compared to AI moderators’ 99 WPM. That is, on average, AI moderators were speaking 80% more than human moderators.</p>
<p><strong>Participants Spoke 45% Less to AI Moderators</strong></p>
<p>Moderators speaking more isn’t necessarily bad as long as participants are also speaking. The data told another story. Participants spoke 98 WPM to human moderators compared to 54 WPM to AI moderators (45% less—see Video 2 for examples).</p>
<div class="ast-oembed-container " style="height: 100%;"><iframe loading="lazy" title="AI-responses-items-edited" src="https://player.vimeo.com/video/1227406418?h=66620a14aa&amp;dnt=1&amp;app_id=122963" width="1170" height="658" frameborder="0" allow="autoplay; fullscreen; picture-in-picture; clipboard-write; encrypted-media; web-share" referrerpolicy="strict-origin-when-cross-origin"></iframe></div>
<p class="wp-caption-text" style="text-align: left;"><strong>Video 2:</strong> Examples of different participant responses to the same question with an AI and a human moderator.</p>
<p><strong>AI Moderators Dominated the Talking </strong></p>
<p>Another way to look at the amount of time participants speak to AI moderators versus humans is to look at the ratio of words spoken per session. The ratio tells a similar story, with participants speaking 64% of the words in human-moderated sessions compared to 34% in AI-moderated sessions. In nine of the ten human-moderated sessions, participants out-talked the moderator; in ten AI-moderated sessions, participants spoke more than the moderator only three times.</p>
<p>We also looked at the ratio of words spoken by participant to moderator. Like the other metrics, AI moderators just spoke more. For every human moderator word, we got 1.8 words out of participants. With the AI moderators, every word generated only 0.5 words (less than 1/3 than in human sessions).</p>
<p>This result differs from the Zhu et al. study, where they found that the total participant word count didn&#8217;t change much between conditions. That could be a function of their research context being think-aloud during a usability test versus in-depth interviewing (which requires the moderator to draw more out of participants).</p>
<p><strong>AI Moderators Spoke a Bit Faster and Responded Slightly Slower</strong></p>
<p>We measured the time actually spent talking and found the AI moderator was about 11% faster on average than the human moderators (consistent with Zhu et al.). Participants spoke at the same rate no matter who they were talking to, so what changed was who had the floor, not the speed of talking.</p>
<p>When we looked at the delay between the end of participant speech and the beginning of moderator speech, we saw a slight difference in response time. AI moderators took 1.9 seconds to respond versus 1.5 seconds for human moderators. Human moderators were slightly faster in responding on average but it’s not clear whether participants even perceived differences this small (400ms on average).</p>
<h2>Summary and Discussion</h2>
<p>A controlled experiment with 20 participants randomly assigned to human or AI moderator in-depth interview conditions found that participants spoke less to AI moderators. Specifically, AI moderators’ sessions were 37% shorter, participants spoke 45% less to AI moderators, and AI moderators dominated the conversations, speaking 64% of the words. Human moderators spoke only about 34% of the words, a ratio that matches the typical expectation of moderators: let the participant do the talking. This is just one study, but the results match those of Zhu et al., who also found that AI moderators <strong>talked more, faster, and more often</strong> than human moderators.</p>
<p>Of course, the number of words and amount of speaking doesn’t mean AI moderators are bad or useless. In fact, future prompt modifications could improve results. And even if participants still talk less to AI moderators than humans, if the same or more information is extracted from participants and the same conclusions are reached, then it might not matter who’s talking too much. We’ll examine the content of the sessions and what’s extracted in future articles where we dig more into the experimental data.</p>
<h2>Appendix</h2>
<h3>Discussion Guide</h3>
<p><strong>Sleep schedule:</strong> Tell me a bit about your sleep schedule. When do you usually go to sleep and wake up? <em>Probe if needed:</em> Do you wake up at the same time every day, or do you wake up at different times depending on the day?</p>
<p><strong>Wake-up method:</strong> How do you make sure you&#8217;ll wake up at the time you intend to? <em>(Most likely they will mention some kind of alarm clock, app, or device.)</em> What is it, how does it work, and does it ever fail? <em>(If they say someone else wakes them up)</em> How do they wake you up? Are there ever any days when they aren&#8217;t available to wake you? What do you do to wake up on time on those days? <em>(If they previously mentioned they wake up at a different time on different days)</em> Since you wake up at a different time depending on the day, what do you do to make sure you&#8217;ll wake up when you want to each day?</p>
<p><strong>Sleeping arrangement:</strong> Do you sleep alone, or does anyone else sleep in the same room with you? <em>(If they share a room)</em> Do they wake up at the same time you do? <em>(If no)</em> How do you make sure each of you wakes up on time without disturbing the other?</p>
<p><strong>Current experience:</strong> How has your experience been so far using your current wake-up method? What do you like about using it to wake up? What problems or difficulties have you experienced? What workarounds or backup solutions have you used after running into those problems?</p>
<p><strong>Ideal wake-up:</strong> What would be your ideal way to be woken up in the morning?</p>
<h3>Moderator Probing Guide</h3>
<p>1. GOAL: Cover ALL required probe areas AND uncover deeper insights related to RESEARCH GOALS.</p>
<p>2. REQUIRED PROBE AREAS must be addressed (non-negotiable, regardless of time) &#8211; adapt phrasing based on context.</p>
<ul>
<li style="list-style-type: none;">
<ul>
<li>If already mentioned: Confirm or dig deeper rather than re-asking.</li>
<li>If not yet covered: Ask directly but naturally</li>
</ul>
</li>
</ul>
<p>3. WHEN TO MOVE ON FROM A THREAD:</p>
<ul>
<li style="list-style-type: none;">
<ul>
<li>You&#8217;ve exhausted it (participant repeating themselves or has nothing more to add).</li>
<li>The participant explicitly says they don&#8217;t know or has nothing more to add.</li>
<li>It&#8217;s a minor detail that doesn&#8217;t advance research understanding.</li>
</ul>
</li>
</ul>
<p>4. BALANCE: Cover all required probes + follow valuable tangents related to research goals. Don&#8217;t chase irrelevant details or go too in depth on minor details.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Effectiveness of AI Moderators: A Literature Review</title>
		<link>https://measuringu.com/effectiveness-of-ai-moderators-a-literature-review/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=effectiveness-of-ai-moderators-a-literature-review</link>
		
		<dc:creator><![CDATA[Lucas Plabst, PhD • Jeff Sauro, PhD • Jim Lewis, PhD]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 02:52:24 +0000</pubDate>
				<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Moderating]]></category>
		<category><![CDATA[Moderation]]></category>
		<category><![CDATA[moderator]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48460</guid>

					<description><![CDATA[AI can replace research moderators. Maybe yes; maybe no. When we encounter interesting claims like that, we generally take the following approach: Define the claim. Assess evidence and the quality (focusing on peer-reviewed publications). Conduct our own experiments. Refine and revise our opinion of the claim. When it comes to AI moderators, we are now [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1.jpg" rel="attachment wp-att-24969"><img loading="lazy" decoding="async" class="alignleft wp-image-48486 size-medium" src="https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-300x169.jpg" alt="Feature image" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/09/090826-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a><em>AI can replace research moderators.</em></p>
<p>Maybe yes; maybe no.</p>
<p>When we encounter interesting claims like that, we generally take the following approach:</p>
<ol>
<li>Define the claim.</li>
<li>Assess evidence and the quality (focusing on peer-reviewed publications).</li>
<li>Conduct our own experiments.</li>
<li>Refine and revise our opinion of the claim.</li>
</ol>
<p>When it comes to AI moderators, we are now on step two in that process. We covered step one in a <a href="https://measuringu.com/can-we-trust-ai-to-moderate-ux-interviews/">previous article</a> in which we asked whether AI can reasonably be used to moderate UX interviews.</p>
<p>We&#8217;ve also written previously about <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">what separates an adequate moderator from an excellent one</a>, and asking good questions is only one part of that list. It starts before the session, with understanding why a stakeholder wanted the study at all—that&#8217;s what tells a skilled moderator when to probe and when to go off script. In the session, it means knowing when to assist a stuck participant without contaminating the task, when someone is misrepresenting who they are, when to stop probing and move on, and how to handle observers who want more than the discussion guide asks for.</p>
<p>Some data suggest that <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5395709">AI could be used at scale to interview job candidates</a>, but the interviews in that case were straightforward questions and answers for a low-level position. This is qualitatively different from semi-structured UX sessions conducted by skilled UX moderators who build <a href="https://measuringu.com/five-ways-to-use-rapport-when-moderating/">rapport</a> and can judge when to go off script to follow up an interesting lead versus when to stick to the guide.</p>
<p>For this article, we reviewed the published literature to investigate evidence for and against the effectiveness of AI moderators conducting interviews. Specifically, we examined five key questions about the advantages (or lack thereof) of using generative AI as research moderators:</p>
<ul>
<li>Do participants say less to AI than to human moderators?</li>
<li>Do participants provide the same depth of insights to AI and human moderators?</li>
<li>Does AI moderate as competently as a human?</li>
<li>Do participants want to talk to AI moderators?</li>
<li>Do findings generalize across models or prompts?</li>
</ul>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2>How we picked these studies</h2>
<p>We searched the peer-reviewed literature for work on AI-led interviewing published between 2024 and mid-2026. To ensure a reasonable approximation of the capabilities of current LLM agents, all papers in this review used ChatGPT-3.5 or a later model.</p>
<p>We prioritized peer-reviewed papers but included two exceptions: Chopra and Haaland&#8217;s <a href="https://www.ifo.de/en/cesifo/cesifo-homepage">CESifo</a> working paper (a large, empirical AI-interview dataset with <em>n</em> = 766) and <a href="https://www.nngroup.com/articles/ai-interviewers/">Rosala&#8217;s Nielsen Norman Group report</a> (the only source describing real-world practitioner tool use), for a total of ten sources. For each question in the following sections, findings are tagged as <em>pro AI</em> (apparent advantage for AI moderation), <em>neutral</em> (no apparent advantage or disadvantage for AI), or <em>anti AI</em> (apparent disadvantage for AI). All studies are summarized in the appendix.</p>
<h2>Do participants say less to AI than to human moderators (number of words, engagement, elaboration)?</h2>
<p>There are some claims that people talk more to AI moderators about sensitive topics, but for UX research, which doesn’t usually involve sensitive topics, this is less likely. Consequently, we don’t expect an AI moderator to generate more talking from a participant, but we would be concerned if people spoke <em>less</em> to AI moderators.</p>
<p>There were two studies with relevant findings:</p>
<p>The results of Wuttke et al. (<a href="https://aclanthology.org/2025.latechclfl-1.17.pdf">2025</a>) were mixed. They ran a small classroom pilot in which student pairs each completed one AI-led (GPT-4 Turbo) and one human-led interview on politics and democracy, analyzing <strong>six</strong> human-led and <strong>five</strong> AI-led sessions.</p>
<p>In their Table 1, they showed participants’ raw response lengths were longer when the moderator was AI (52 vs. 33 words/answer; <strong>pro AI</strong>), but ratings of engagement and elaboration were better for human moderators (<strong>anti AI</strong>). Note that the human moderators in this study were students, not seasoned UX professionals.</p>
<p>Zhu and colleagues (<a href="https://dl.acm.org/doi/pdf/10.1145/3772318.3791653">2026</a>) ran a randomized controlled trial comparing an agentic voice moderator (using Chinese ByteDance technologies like Doubao-1.5-pro-32k-250115) against a human moderator in think-aloud testing of a note-taking app (<em>n</em> = 60). They found no significant difference in measures of total words in responses to their AI and human moderators (Table 6, <strong>neutral</strong>).</p>
<h3>Other findings of interest</h3>
<p>Cuevas and colleagues (<a href="https://dl.acm.org/doi/pdf/10.1145/3710947">2025</a>) ran a large study (<em>n</em> = 399) pitting two LLM-based interview chatbots using GPT-3.5-turbo against a hard-coded-question baseline. They found no significant difference in the number of words per session.</p>
<p>Chopra and Haaland (<a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">2026</a>) ran two large-sample open-ended qualitative interviews (one on stock market nonparticipation in 2023, <em>n</em> = 381; one on attitudes toward U.S. tariffs in 2025, <em>n</em> = 385) with a multi-agent AI chatbot, comparing conditions with different levels of probing. They reported 29 words per minute (wpm) in response to their AI interviewer in their stock market study and compared that to a published benchmark for human-led chat-based interviews (20 wpm; Namey et al., <a href="https://journals.sagepub.com/doi/pdf/10.1177/1525822X19886839?casa_token=wpbwKGxIhGEAAAAA:IBZr1yU9-81ZGU8ctI7LZkIpTxDhS-HLLvAkzuDl9Rf0mcahR8agejaTnQLRzTKfuFI4jnexxj65NQ">2020</a>). However, there appears to be an error in the wpm calculations for the AI interviewer because the mean number of words was 654 and the mean number of minutes was 33, so the actual wpm was 19.8 (654 / 33). We classified this finding as not relevant due to the apparent calculation error and the weakness of this type of uncontrolled cross-study comparison.</p>
<p><strong>Bottom line:</strong> The evidence is mixed regarding whether participants say more to AI than human moderators, but there is no compelling evidence that they say less.</p>
<h2>Do participants provide as much depth of insights to AI and human moderators (richness, novel insight)?</h2>
<p>Quantity of words does not equate to quality of insights. A <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">good moderator</a> should be able to get participants to go deep on topics to help generate insights for stakeholders. Do AI moderators go as deep as humans?</p>
<p>The relevant studies were Wuttke et al. (<a href="https://aclanthology.org/2025.latechclfl-1.17.pdf">2025</a>) and Zhu et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3772318.3791653">2026</a>).</p>
<p>The classroom study by Wuttke et al. (their Table 1) reported lower human-coded ratings of specificity (level of detail in the response) for their AI moderator (<strong>anti AI</strong>).</p>
<p>Zhu et al. reported “The analysis indicates that the depth of participants’ analytical thinking and problem-solving articulation remained similar across conditions” (Section 4.3, <strong>neutral</strong>).</p>
<h3>Other findings of interest</h3>
<ul>
<li>Chopra and Haaland (<a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">2026</a>) reported faster discovery of themes for their AI moderators than single or multiple open-ended survey questions (their Figure 8), roughly equivalent to discovery rates with human participants.</li>
<li>Cuevas et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3710947">2025</a>) reported their AI moderator had better follow-up than a hard-coded bot but no practical difference in richness of the responses.</li>
<li>Kuric et al. (<a href="https://www.tandfonline.com/doi/pdf/10.1080/10447318.2024.2427978">2025</a>) reported poorer results for their AI moderator compared to static pre-written follow-up questions.</li>
</ul>
<p><strong>Bottom line:</strong> The evidence is mixed regarding whether AI moderators can achieve human levels of response depth from participants, but there is no compelling evidence that responses to AI moderators are shallower or deeper.</p>
<h2>Does AI moderate as competently as a human (protocol fidelity, probing, task outcomes)?</h2>
<p>Focusing on comparison with human moderators, once again the relevant studies were Wuttke et al. (<a href="https://aclanthology.org/2025.latechclfl-1.17.pdf">2025</a>) and Zhu et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3772318.3791653">2026</a>).</p>
<p>The findings in Wuttke et al. were about evenly distributed among pro AI, neutral, and anti AI. AI had the advantage in listening/paraphrasing participant responses and in being less likely to lead participants (<strong>pro AI</strong>). On the other hand, AI often failed to follow up on unclear/surprising answers and had a potentially biasing tendency to praise respondents (<strong>anti AI</strong>). Overall, both AI and humans “faithfully followed the provided questionnaire” and operated at similar levels of competence (<strong>neutral</strong>). When interpreting these findings, keep in mind that this finding is complicated, as the respondents in each condition were students who also played other roles (interviewer, observer) across the two conditions.</p>
<p>Zhu et al. reported no practical difference for procedural adherence or think-aloud guidance quality (<strong>neutral</strong>). AI moderation was affected by occasional task fixation and flow disruption and had significantly poorer context-aware follow-up and trust building (Table 8, <strong>anti AI</strong>).</p>
<h3>Other findings of interest</h3>
<ul>
<li>Kuric et al. (<a href="https://www.tandfonline.com/doi/pdf/10.1080/10447318.2024.2427978">2025</a>) reported poorer performance of their AI moderator relative to static pre-written follow-up for surfacing new usability issues, leading participants, and perceived reasonability of questions.</li>
<li>Cuevas et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3710947">2025</a>) reported better performance of their AI moderator relative to a hard-coded baseline chatbot for follow-up quality and continuity, but no practical difference for relevance, clarity, or specificity.</li>
<li>Panfilova et al. (<a href="https://www.nature.com/articles/s41598-026-46517-7.pdf">2026</a>) measured absolute protocol compliance and probing judgment scores across six different LLMs, excluding DeepSeek early in the study for compliance failures and noting that Grok over-probed already complete answers.</li>
<li>Chopra and Haaland (<a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">2026</a>) reported high protocol fidelity for their AI moderator.</li>
<li>Jacobsen and colleagues (<a href="https://dl.acm.org/doi/10.1145/3706598.3714128">2025</a>) tested four different probe types (descriptive, idiographic, clarifying, and explanatory) embedded directly in an online survey (<em>n</em> = 64, 16 participants per probe type condition). They found that the idiographic strategy was generally the most competent and explanatory the least (based on human-coded metrics of relevance, specificity, and clarity).</li>
<li>Rosala (<a href="https://www.nngroup.com/articles/ai-interviewers/">2026</a>) reported Nielsen Norman Group’s hands-on test of AI-moderated interview tools, Marvin and UserFlix, with ten research leaders and ResearchOps professionals across eight countries evaluating the tools. They described rigid script-following, timing problems, and sycophancy with the AI interviewing products they evaluated.</li>
</ul>
<p><strong>Bottom line:</strong> Evidence relative to the competence of AI versus human moderation currently rests on two studies (Wuttke et al. with five AI-led and six student-led interviews; Zhu et al. with 30 each AI- and human-led interviews). The findings indicate that AI matches (but does not exceed) human skill for mechanical, procedural aspects of moderation, might do a little better at properly paraphrasing without leading, but is inferior to humans on judgment-dependent competence.</p>
<h2>Do participants want to talk to AI moderators (rapport, trust, willingness)?</h2>
<p><a href="https://measuringu.com/five-ways-to-use-rapport-when-moderating/">Rapport</a> is more than just breaking the ice; it’s a tool to help establish trust and elicit deeper insights. How well does AI do this compared to humans?</p>
<p>There were five relevant studies for this question:</p>
<ul>
<li>Chopra and Haaland (<a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">2026</a>) reported, “a majority of participants would prefer an AI interviewer over a human interviewer” (<strong>pro AI</strong>).</li>
<li>Zhu et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3772318.3791653">2026</a>) found, for most subgroups, a strong preference for human moderation (<strong>anti AI</strong>) but a preference for AI from their subgroup of introverts (<strong>pro AI</strong>).</li>
<li>Wuttke et al. (<a href="https://aclanthology.org/2025.latechclfl-1.17.pdf">2025</a>) reported no difference in overall satisfaction between AI- and human-led moderation (<strong>neutral</strong>) but better Interestingness and Repeatability (willingness to do it again) ratings for human moderation (<strong>anti AI</strong>).</li>
<li>Jacobsen et al. (<a href="https://dl.acm.org/doi/10.1145/3706598.3714128">2025</a>) found a slight preference for disclosing to an AI chatbot relative to a human (14 AI, 21%; 6 human, 9%; <strong>pro AI</strong>), but most participants had no preference (44, 67%; <strong>neutral</strong>).</li>
<li>Cuevas et al. (<a href="https://dl.acm.org/doi/pdf/10.1145/3710947">2025</a>) reported “no significant differences … [in] preference for a human versus an AI interviewer” (<strong>neutral</strong>).</li>
<li>Danó and colleagues (<a href="https://tge.sze.hu/tge/article/download/413/213">2025</a>) studied Hungarian attitudes toward AI interviewers, finding 49% of respondents were reluctant to engage at all with a virtual interviewer (<strong>anti AI</strong>).</li>
</ul>
<p><strong>Bottom line:</strong> Results were mixed regarding participant preference for AI or human moderation. For this question, however, even if only a substantial minority of people are reluctant to engage with AI moderators, for many types of research this could be a problem for recruiting and increased likelihood of study abandonment.</p>
<h2>Do findings generalize across models or prompts?</h2>
<p>LLMs are changing weekly and are also probabilistic. If changing prompts and models changes results (for better or worse), it’s difficult to generalize the findings to practitioners.</p>
<p>For this question, we classify evidence of generalizability as pro AI and evidence against it as anti AI. There were four relevant studies:</p>
<ul>
<li>Chopra and Haaland (<a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">2026</a>) reported using the same prompt across multiple studies and models with “no application-specific instructions, making it portable across domains” (<strong>pro AI</strong>).</li>
<li>Panfilova et al. (<a href="https://www.nature.com/articles/s41598-026-46517-7.pdf">2026</a>) ran six different LLMs on an identical task, finding meaningfully different rankings and failure modes (<strong>anti AI</strong>).</li>
<li>Jacobsen et al. (<a href="https://dl.acm.org/doi/10.1145/3706598.3714128">2025</a>) reported different probing behaviors as a function of manipulating prompts for probing strategies (<strong>anti AI</strong>).</li>
<li>Rosala (<a href="https://www.nngroup.com/articles/ai-interviewers/">2026</a>) reported that two commercial AI moderation tools (Marvin, UserFlix) had different strengths and weaknesses (<strong>anti AI</strong>).</li>
</ul>
<p><strong>Bottom line:</strong> Many authors noted the impermanence of rapidly changing models and different results for different prompts as a characteristic of generative AI. The moving target of models and prompts makes it difficult to conduct research on these topics. The model often credited with kicking off the AI craze in 2020, GPT-3, ran on 175 billion parameters with a context window of 2,048 tokens. Kimi K3, released in July 2026, carries 2.8 trillion parameters and a one-million-token context window. When the underlying technology changes that fast, any result you attribute to the model of the moment necessarily has a short shelf life.</p>
<h2>Summary and Discussion</h2>
<p>Is AI a suitable replacement for human moderators? A review of the (mostly) published literature found mixed results. Table 1 summarizes the results for the five questions, pro AI (6), neutral (6), and anti AI (9). Across the sources, pro- and anti-AI findings were about balanced.</p>

<table id="tablepress-1074" class="tablepress tablepress-id-1074">
<thead>
<tr class="row-1">
	<th class="column-1">The Five Questions</th><th class="column-2">Pro AI</th><th class="column-3">Neutral</th><th class="column-4">Anti AI</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">Do participants say less to AI than to human moderators (number of words, engagement, elaboration)?</td><td class="column-2">1 (Wu)</td><td class="column-3">1 (Zh)</td><td class="column-4">1 (Wu)</td>
</tr>
<tr class="row-3">
	<td class="column-1">Do participants provide as much depth of insights to AI and human moderators (richness, novel insight)?</td><td class="column-2">0</td><td class="column-3">1 (Zh)</td><td class="column-4">1 (Wu)</td>
</tr>
<tr class="row-4">
	<td class="column-1">Does AI moderate as competently as a human (protocol fidelity, probing, task outcomes)?</td><td class="column-2">1 (Wu)</td><td class="column-3">1 (Wu, Zh)</td><td class="column-4">1 (Wu, Zh)</td>
</tr>
<tr class="row-5">
	<td class="column-1">Do participants want to talk to AI moderators (rapport, trust, willingness)?</td><td class="column-2">3 (Ch, Zh, Ja)</td><td class="column-3">3 (Cu, Ja, Wu)</td><td class="column-4">3 (Da, Wu, Zh)</td>
</tr>
<tr class="row-6">
	<td class="column-1">Do findings generalize across models or prompts?</td><td class="column-2">1 (Ch)</td><td class="column-3">0</td><td class="column-4">3 (Ja, Pa, Ro)</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left"><strong>Table 1:</strong> Findings across the relevant sources, grouped by question. The two-letter codes are the first two letters of the lead author’s last name (e.g., Ch for Chopra, Cu for Cuevas). The same author appears in multiple columns for a question when the study findings were mixed.</p>
<p>AI moderators achieved surface-level adherence to practices like asking open-ended, non-leading questions and staying on the topic guide. Their most consistent weakness was judgment about <strong>when and how much to probe</strong>. Thus, on the more desirable outcomes of depth and richness, our review finds them falling short.</p>
<p>A potentially problematic issue for researchers is the possibility of low willingness of many participants to engage with AI moderators, leading to reduced recruitment and increased survey abandonment. So, when a client asks what we think about AI moderators, our answer is, “It depends on the job.”</p>
<p>For structured, high-volume, consistency-driven data collection, an AI moderator seems like a reasonable tool today. For discovery, emotional nuance, and studies where the whole point is to find what you didn&#8217;t know to ask about, the human moderator isn&#8217;t going anywhere yet. AI seems proficient with the mechanical aspects of this work but lacks in the key area of judgment.</p>
<p>Little of the work above is UX research, and none of it is ours. So, we decided to run our own study: a UX-focused evaluation of AI moderators against the work we do every day as human UX researchers. Stay tuned for that article.</p>
<h2>Appendix: References and Summaries</h2>
<p><a href="https://www.researchgate.net/profile/Felix-Chopra-2/publication/374633338_Conducting_Qualitative_Interviews_with_AI/links/6a01b3a4e93d46191595bfc2/Conducting-Qualitative-Interviews-with-AI.pdf">Chopra, F., &amp; Haaland, I. (2026). Conducting qualitative interviews with AI. CESifo Working Paper No. 10666.</a></p>
<p>Chopra and Haaland ran two large-sample open-ended qualitative interviews (one on stock market nonparticipation in 2023, <em>n</em> = 381; one on attitudes toward U.S. tariffs in 2025, <em>n</em> = 385) with a multi-agent AI chatbot at a scale and marginal cost no human team could match. The comparisons were between conditions with different levels of probing (a single open-ended question, multiple open-ended questions, AI-led interview with no probing, and AI-led interview with probing), claiming five times as many unique themes discovered for AI-led probing relative to a single open-ended survey item. Adding more open-ended questions or AI without probing closed about half that gap. Human coders judged the transcripts to hold substantive qualitative content, participants rated the experience favorably, and most participants approved of the AI-generated summaries of their statements presented after the interview. AI and human review of transcripts found the AI interviewers consistently adhered to core methodological guidelines for interviewing (open-ended, relevant, non-leading) as guided by the prompts (95% in the first survey, near 100% in the second). Most participants indicated a preference for AI over human moderation and a willingness to take future AI-led interviews.</p>
<p><a href="https://dl.acm.org/doi/10.1145/3710947">Cuevas, A., Scurrell, J. V., Brown, E. M., Entenmann, J., &amp; Daepp, M. I. G. (2025). Collecting qualitative data at scale with large language models: A case study. <em>Proceedings of the ACM on Human-Computer Interaction</em>, <em>9</em>.</a></p>
<p>Cuevas and colleagues ran a large user study (<em>n</em> = 399) pitting two LLM-based interview chatbots using GPT-3.5-turbo against a hard-coded-question baseline. They scored the transcripts in two ways: on established communication-quality metrics, and on a purpose-built &#8220;richness&#8221; scale: how well a response captured the complexity and specificity of the respondent&#8217;s actual situation. The chatbots scored well on the standard metrics but no better than the hard-coded baseline except for follow-up quality. Furthermore, the responses rarely surfaced a participant&#8217;s specific motives or personal examples, so they scored poorly on richness. Surface fluency does not necessarily lead to genuine qualitative depth.</p>
<p><a href="https://tge.sze.hu/tge/article/view/413">Danó, G., Kovács, S., &amp; Surman, V. (2025). Challenges and opportunities of AI in market research: Virtual interviewers. <em>Tér-Gazdaság-Ember/Journal of Region, Economy and Society</em>, <em>13</em>.</a></p>
<p>Danó and colleagues studied Hungarian attitudes toward AI interviewers. They conducted a large-scale survey (<em>n</em> = 1077) in June 2024 that asked people how they would feel about being interviewed by AI with a human voice, finding that 49% of respondents were reluctant to engage at all with a virtual interviewer. An AI interviewer that half of your sample doesn&#8217;t want to use could create a data-quality problem.</p>
<p><a href="https://www.tandfonline.com/doi/full/10.1080/10447318.2024.2427978">Kuric, E., Demcak, P., &amp; Krajcovic, M. (2025). Unmoderated usability studies evolved: Can GPT ask useful follow-up questions? <em>International Journal of Human–Computer Interaction</em>, <em>41</em>(15), 9752–9769.</a></p>
<p>Kuric and colleagues ran a between-subjects experiment (<em>n</em> = 60) comparing unmoderated usability debriefs either with or without real-time GPT-4 follow-ups, then compared four ways of asking follow-up questions within that data (none, researcher-authored static, GPT-4-generated, and a blend of the last two). The follow-ups added depth to known issues but did not surface any new ones. Participants also rated the questions as significantly less reasonable when the AI was probing, complaining they felt repetitive, and their answers to the primary seed question got worse when they knew AI follow-ups were coming.</p>
<p><a href="https://dl.acm.org/doi/10.1145/3706598.3714128">Jacobsen, R. M., Cox, S. R., Griggio, C. F., &amp; van Berkel, N. (2025). Chatbots for data collection in surveys: A comparison of four theory-based interview probes. <em>Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems</em>, 228.</a></p>
<p>Jacobsen and colleagues tested four different probe types (descriptive, idiographic, clarifying, and explanatory) embedded directly in an online survey (<em>n</em> = 64). In response to a participant’s statement, a descriptive probe asks what you were doing, feeling, and thinking; an idiographic probe asks for a specific example; a clarifying probe asks what it means to you; and an explanatory probe asks why you believe what you said. The idiographic version won overall and was the only one still working well by the evaluation stage. The explanatory &#8220;why&#8221; questions never came out on top.</p>
<p><a href="https://www.nature.com/articles/s41598-026-46517-7">Panfilova, A., Bolshev, V., Mozikov, M., Latynov, V., Vanin, A., Nestik, T., Nourkova, V., Vlasova, A., Kozin, M., Serohvostov, A., Tarasova, E., &amp; Nikolenko, S. (2026). The AI interviewer: Multi-faceted evaluation of adaptive questioning by large language models. <em>Scientific Reports</em>.</a></p>
<p>Panfilova et al. ran a controlled protocol across six frontier models at the time of data collection. AI interviewers created with five of the models completed the protocol (Claude Sonnet 4, Gemini 2.5 Pro, GPT-5, Grok 4, Qwen3), and one (DeepSeek) was removed due to serious performance issues. The researchers started by collecting ten baseline interviews in which the moderator and respondents were human, getting responses to 54 questions with no follow-up. The AI interviewers reviewed the transcripts question by question, decided whether follow-up was warranted and, if so, asked the additional question. Questions were answered by an AI agent conditioned with the human respondent’s Big Five personality profile, with the possibility of additional probing. The appropriateness of following up and quality of the additional questions were assessed by human evaluators on five binary criteria (benevolence, necessity, context-awareness, openness, and justified skip). Interviewing competence varied substantially by model but with no clear winner. For example, Gemini was rated as the most empathetic model, while Grok produced the most follow-ups but tended to over-probe already complete responses. The best-performing models on quality were not the cheapest or fastest.</p>
<p><a href="https://www.nngroup.com/articles/ai-interviewers/">Rosala, M. (2026, January 30). AI-moderated interviews: If, when, and how to use them. Nielsen Norman Group.</a></p>
<p>The Nielsen Norman Group&#8217;s AI Interviewers article reports a hands-on test of two AI-moderated interview tools, Marvin and UserFlix, with ten research leaders and ResearchOps professionals across eight countries (tool evaluators rather than recruited respondents). Their verdict: AI interviewers suit a bounded set of uses (product-feedback collection, recruitment screening, translated interviews, and teams without a dedicated researcher), but not exploratory research, high-stakes decisions, or work that demands domain expertise and real-time judgment.</p>
<p><a href="https://dl.acm.org/doi/10.1145/3637364">Wei, J., Kim, S., Jung, H., &amp; Kim, Y.-H. (2024). Leveraging large language models to power chatbots for collecting user self-reported data. <em>Proceedings of the ACM on Human-Computer Interaction</em>, <em>8</em>(CSCW1).</a></p>
<p>Wei and colleagues built chatbots that took on four roles: a sleep expert, a dietitian, a life coach, and a fitness coach. They manipulated prompts to create four AI interviewer types by crossing two formats (structured list and descriptive narrative) and two personality modifiers (with and without “who always shows empathy and engages my customer in conversations” in the prompt), abbreviated SP, SN, DP, and DN, with 12 participants assigned to each version (<em>n</em> = 48). Participants spent less than 20 minutes completing eight conversations with their assigned interviewer type (four roles by two scenario paths, one positive and one negative). A key dependent measure was the percentage of 18 predefined information slots (facts) the AIs were directed to gather during the conversations. Using nothing but the LLM itself (no training data and no fine-tuning), on average the chatbots captured about 79% of the targeted facts (SP: 83%, SN: 77%, DP: 72%, DN: 83%). How the prompt was written mattered. The researchers reported that the personality modifier improved performance for the structured format but hurt the descriptive format, even though both formats contained the same factual information.</p>
<p><a href="https://aclanthology.org/2025.latechclfl-1.17/">Wuttke, A., Aßenmacher, M., Klamm, C., Lang, M. M., Würschinger, Q., &amp; Kreuter, F. (2025). AI conversational interviewing: Transforming surveys with LLMs as adaptive interviewers. <em>Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature </em>(LaTeCH-CLfL 2025).</a></p>
<p>Wuttke and colleagues ran a small classroom pilot study in which student pairs each completed one AI-led (using GPT-4 Turbo) and one human-led interview in randomized order, rotating roles. The researchers analyzed six human-led and five AI-led sessions, using an identical questionnaire on politics and democracy. The AI matched the student interviewers’ overall rate of guideline violations, but the errors differed in kind. The AI consistently missed chances to ask a follow-up on unexpected or unclear answers (88% of the violations of this rule). The students, on the other hand, often missed engaging in active listening, which includes restating what the participants had just said (94% of the violations of this rule). Both students and AI interviewers behaved in ways that potentially biased respondents’ answers. Participants rated the AI-led interviews as less interesting and were less willing to repeat them, even though task-level measures like response quality and understanding came out similarly. This finding is complicated by the respondents in each condition having also played other roles (interviewer, observer) across the two conditions.</p>
<p><a href="https://dl.acm.org/doi/10.1145/3772318.3791653">Zhu, W., Chen, G., Wang, Y., An, P., Du, J., &amp; Li, C. (2026). Agentic audio moderator vs human moderator in think-aloud usability testing: Results from a randomized controlled trial. <em>Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems</em>.</a></p>
<p>Zhu and colleagues ran a randomized controlled trial of an agentic voice moderator (using Chinese ByteDance technologies like Doubao-1.5-pro-32k-250115) against a human in think-aloud testing of a note-taking app (<em>n</em> = 60). They found no significant difference in participants’ task performance or verbalization behavior, but significantly lower social-perception ratings such as anthropomorphism, animacy, likeability, intelligence, and social presence for the AI. The verbalization behavior of AI and human moderators differed, with the AI moderator speaking more, faster, and more frequently than the humans during the think-aloud phase, behaviors that likely lowered the AI’s rapport score.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Nine Design Fixes for Your UX Research Reports</title>
		<link>https://measuringu.com/nine-design-fixes-for-your-ux-research-reports/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=nine-design-fixes-for-your-ux-research-reports</link>
		
		<dc:creator><![CDATA[Fernanda Villalobos, MS • Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Wed, 02 Sep 2026 01:24:05 +0000</pubDate>
				<category><![CDATA[Usability]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Report]]></category>
		<category><![CDATA[Usability Testing]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48303</guid>

					<description><![CDATA[We’re known in the industry for metrics, for methods, and for defining UX research measurement. But while we believe in using data to drive decisions, we also know that changing minds is not something you can just enter into an Excel formula. Findings are only as effective as the way they are communicated. If the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48385 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-300x169.jpg" alt="Feature image showing nine design fixes for your UX research reports" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>We’re known in the industry for metrics, for methods, and for defining UX research measurement.</p>
<p>But while we believe in using data to drive decisions, we also know that changing minds is not something you can just enter into an Excel formula.</p>
<p>Findings are only as effective as the way they are communicated. If the data-analysis tree falls in the forest, will it really make a sound if no one is listening?</p>
<p>While the data are crucial, your manner of presentation determines whether your insights are absorbed, acted upon, or ignored.</p>
<p>Some of that happens when displaying the data in graphs, and we’ve previously discussed best practices for <a href="https://measuringu.com/graphing-displaying-data/">graphing and displaying data</a>. But we also recognize that the design of the deliverable itself, not just the graphs, makes a big difference.</p>
<p>The UX research report is one of the <a href="https://measuringu.com/what-are-ux-deliverables/">core deliverables</a> in our field. It can be from a usability test, an in-depth interview, or a survey. Will AI put an end to the much-maligned slide deck? Maybe. Certainly, AI is making it easier to create presentations. For at least the next few months, we don’t see the presentation going away, even if AI generates it.</p>
<p>And one thing that won’t change is the limited attention span of the readers. They won’t spend as much time reviewing your report as you put into building it. When time is limited and attention is divided (when isn’t it?), visual displays aid in <a href="https://journals.sagepub.com/doi/10.1080/14640747308400340">comprehension and retention</a>.</p>
<p>Here are nine ways to make your UX deliverable more digestible.</p>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2>1. Start with a report structure.</h2>
<p>A sturdy building starts with a solid foundation. A clear structure makes the report stronger and is essential for guiding the narrative and helping readers follow the information, particularly since a UX researcher will not always be present to explain the content.</p>
<p>A common and effective layout for a UX research report includes:</p>
<ul>
<li><strong>Table of Contents:</strong> Helps with navigating the report and introduces the sections covered.</li>
<li><strong>Background:</strong> Provides necessary context, as not everyone will be familiar with the study’s goals.</li>
<li><strong>Methodology:</strong> Briefly outlines the research methods and participants involved.</li>
<li><strong>Executive Summary/Key Findings:</strong> Offers a concise summary of the main takeaways for clients with limited time.</li>
<li><strong>Detailed Findings:</strong> Presents the core data, including images, videos, participant quotes, and metrics (tables/charts) to support your key findings.</li>
<li><strong>Recommendations:</strong> Incorporated alongside detailed findings or in a dedicated section. Prioritize what needs immediate attention by utilizing urgency rankings.</li>
<li><strong>Next Steps:</strong> Highlights actionable items or future suggestions.</li>
<li><strong>Appendix:</strong> Include to keep the main report focused. Link to this section for additional supporting details, ensuring the document remains concise while still being comprehensive.</li>
</ul>
<p>This structure doesn’t work for every deliverable, but it’s a good place to start, and you can remove, add, or modify at will. You don’t need to always follow this structure or include every section. You also don’t need to create a separate slide or section for each part of the suggested structure.</p>
<h2>2. Break up text with high-resolution imagery.</h2>
<p>Visuals are processed faster than text, helping your reader grasp and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7618346/">remember</a> complex insights. This approach reduces cognitive load and keeps your report engaging, ensuring that key findings aren&#8217;t buried in dense paragraphs.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png" rel="attachment wp-att-48308"><img loading="lazy" decoding="async" class="alignnone wp-image-48352 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png" alt="Example slides of mostly text and bullets by a high-resolution image." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>3. Implement clear information hierarchy by varying heading sizes and incorporating two or three colors.</h2>
<p>A clear visual hierarchy guides the reader&#8217;s eye across the page to the most critical insights and data points without being overwhelming. Color helps highlight specific information, but do so sparingly, as too many colors make it harder to discern the hierarchy, diminishing their purpose.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png" rel="attachment wp-att-48307"><img loading="lazy" decoding="async" class="alignnone wp-image-48353 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png" alt="Example slides, one lacking hierarchy one one with headings, subheadings, and color coding." width="1200" height="505" srcset="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-1024x431.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-600x253.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>4. Avoid very small text size and low contrast.</h2>
<p>Prioritizing readability ensures that your deliverable is accessible and inclusive for all. Use adequate text size and high-contrast colors to prevent eye strain and frustration, making it easier to read all the information.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png" rel="attachment wp-att-48306"><img loading="lazy" decoding="async" class="alignnone wp-image-48355 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png" alt="Example slides of low-contrast text and high-contrast text." width="1200" height="466" srcset="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-300x117.png 300w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-1024x398.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-768x298.png 768w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-600x233.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>5. Maintain sufficient spacing within margins and between various components.</h2>
<p>Generous whitespace prevents visual clutter and focuses the reader&#8217;s attention on the most important content, making the overall presentation cleaner and less overwhelming.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png" rel="attachment wp-att-48309"><img loading="lazy" decoding="async" class="alignnone wp-image-48356 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png" alt="Example slides, one with no margins and one with margins." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>6. Left-aligned text is easier to read than centered text.</h2>
<p>Consistent left alignment creates a predictable reading pattern, reducing cognitive load and helping the reader process information more efficiently.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png" rel="attachment wp-att-48310"><img loading="lazy" decoding="async" class="alignnone wp-image-48357 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png" alt="Example slides of centered text and left-aligned text." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>7. Match charts and tables to the design of the slide deck.</h2>
<p>Consistency in visual elements like charts and tables reinforces your overall branding and ensures that your report feels like a cohesive, professional document. When data visualizations align with the aesthetic of your slide deck, they are less jarring and help the reader focus on the insights you are presenting.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png" rel="attachment wp-att-48311"><img loading="lazy" decoding="async" class="alignnone wp-image-48358 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png" alt="Example slides showing default chart styling and cohesive chart styling." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>8. Use a template.</h2>
<p>Don’t reinvent the deliverable wheel each time you present. If something already works, reuse it (see below on usability testing your template). To keep a consistent look across deliverables shared by a team, pair it with clear usage instructions and design principles. No two projects are identical, so it needs to stay flexible: fully editable and able to flex for different content lengths and project needs.</p>
<p>MeasuringU&#8217;s template includes placeholder text and components across the layouts we use most, built for each common report section. Uniform fonts, color palettes, imagery, icons, and title treatments balance consistency and variety. We hope to make it engaging without feeling repetitive. Even slides with distinct layouts still read as one cohesive presentation because they follow the same style guide.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48359 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png" alt="image side-by-side comparison of templates" width="1375" height="1359" srcset="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png 1375w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-300x297.png 300w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-1024x1012.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-768x759.png 768w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-70x70.png 70w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-600x593.png 600w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-100x100.png 100w" sizes="auto, (max-width: 1375px) 100vw, 1375px" /></a></p>
<h2>9. Usability test your template and iterate.</h2>
<p>A UX research report is an interface. As such, it has users: the researcher building the report and that short-attention-spanned stakeholder. Test the template with the researcher and reader. Our current template at MeasuringU went through four iterations with the researcher team. All too often, new templates get the visual approval of a team (it looks great!). But can you get the charts in? Do the call-outs allow for enough room? Is the contrast with the text poor? Are those changes working for your reader, or are they adding to more missed points and rework?</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48361" src="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg" alt="image of UX researchers testing the template" width="512" height="506" srcset="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg 741w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-300x297.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-70x70.jpg 70w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-600x594.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-100x100.jpg 100w" sizes="auto, (max-width: 512px) 100vw, 512px" /></a></p>
<h2>Summary and Discussion</h2>
<p>Good UX research doesn&#8217;t speak for itself. The UX research report has to do that work, often without you in the room. The design recommendations here (structure, imagery, hierarchy, contrast, spacing, alignment) aren&#8217;t decoration; they&#8217;re what determines whether a stakeholder reads your key finding or skims past it. Start with the weakest aspect of your current template, fix it, then usability test it on someone who wasn&#8217;t in the room for the research. If your report needs you there to explain it, it isn&#8217;t finished yet.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Five Ways to Use Rapport When Moderating</title>
		<link>https://measuringu.com/five-ways-to-use-rapport-when-moderating/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=five-ways-to-use-rapport-when-moderating</link>
		
		<dc:creator><![CDATA[Jenna Herring, PsyD • Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 21:18:05 +0000</pubDate>
				<category><![CDATA[Methods]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Facilitation]]></category>
		<category><![CDATA[moderated research]]></category>
		<category><![CDATA[Moderating]]></category>
		<category><![CDATA[Moderation]]></category>
		<category><![CDATA[moderator]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48270</guid>

					<description><![CDATA[What makes a good moderator? There is more to it than reading a script. While AI moderation is getting a lot of headlines, humans will still be moderating sessions. Some teams (including ours) are experimenting with AI moderation for lower-stakes, higher-volume studies (verbal surveys). But when it&#8217;s hard to get a participant, and you need [&#8230;]]]></description>
										<content:encoded><![CDATA[<style>
@media (max-width: 782px) {
  .mu-nofloat-mobile img { float: none !important; display: block; margin: 0 auto 1.25em; }
}
</style>
<p><span class="mu-nofloat-mobile"><a href="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48291 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-300x169.jpg" alt="Feature image showing five ways to use rapport when moderating" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a></span>What makes a good moderator? There is more to it than reading a script.</p>
<p>While AI moderation is getting a lot of headlines, humans will still be moderating sessions. Some teams (including ours) are experimenting with AI moderation for lower-stakes, higher-volume studies (verbal surveys). But when it&#8217;s hard to get a participant, and you need to get as much as you can out of their time and yours, the quality of a good moderator pays dividends.</p>
<p>In &#8220;<a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">What Makes a Good UX Research Moderator?</a>&#8221; we described how a good moderator knows the research questions and can differentiate between study types. A summative evaluation requires different moderation than an in-depth interview. In either context, however, a good moderator modulates their style and judiciously uses probing and time management. <strong>One of the fundamental moderator skills is establishing rapport.</strong></p>
<p>Establishing rapport with a participant helps build trust. That trust leads to more revealing insights. Research moderation is not therapy, though it has <a href="https://measuringu.com/thinking-aloud/">its roots there</a>. A participant who doesn&#8217;t trust the moderator may be more likely to say what they <em>think</em> the moderator wants to hear. In this article, we&#8217;ll dig deeper into how to establish a good rapport in your moderation.</p>
<h2>1. Prepare Before the Session, Not During</h2>
<p><a href="https://measuringu.com/facilitation-styles/">Good rapport starts in your prep</a>, not your opening line. Know who you&#8217;re talking to. You wouldn&#8217;t talk to a teenager the way you&#8217;d talk to an IT decision-maker, and you wouldn&#8217;t talk to a brand-new user the way you&#8217;d talk to someone who&#8217;s used the product for years.</p>
<p>Be familiar with the lingo your audience is likely to use. When you hit a term you don&#8217;t recognize, say so directly: “I&#8217;m not familiar with that. Can you explain it?” That&#8217;s not a weakness in the interview; it&#8217;s rapport in action: it tells the participant you&#8217;re actually listening, not just running a script.</p>
<h2>2. Break the Ice Before You Ask for Anything</h2>
<p>Once the session starts, the job is to lower the stakes before you raise a single task. Introduce yourself, tell participants the study is evaluating the product and not them, and remind them there are no wrong answers. If you work for an independent research firm rather than the study&#8217;s company itself, say so. That’s another small way to make it safe to provide critical information about a product or company.</p>
<p>When we break the ice, we like to open with something neutral: How long have you worked in this field? What does a typical day look like? This small talk is the on-ramp that gets a guarded participant comfortable talking openly once you get to the parts that matter. Keep in mind that session length and research goals change what rapport should look like; a 90-minute session may be able to absorb more tangents than a 30-minute session. Break the ice, but don&#8217;t let the ice-breaking break your schedule.</p>
<h2>3. Use a Verbal and Non-Verbal Toolkit</h2>
<p>A good moderator also uses a few additional techniques to carry rapport through the rest of the session:</p>
<ul>
<li><strong>The purposeful pause. </strong><a href="https://measuringu.com/10-golden-rules-of-facilitation/">Don&#8217;t rush to fill the silence</a>. Give participants room to keep thinking; the answer that comes after the pause is often the more honest one.</li>
<li><strong>Stay neutral, not encouraging. </strong>“Great job!” feels supportive, but it teaches the participant to perform for you instead of reacting honestly to the product. (An exception can be a participant who experiences multiple failures during a task-based study; you may choose to offer encouragement to keep them engaged.)</li>
<li><strong>Neutral, friendly expressions throughout. </strong>A face that only warms up at success reads as evaluation, which is exactly what you told them this isn&#8217;t.</li>
<li><strong>On camera, look at the lens, not the face. </strong>In a remote session, eye contact with the participant&#8217;s video feed doesn&#8217;t register as eye contact on their end. Looking at the camera does.</li>
</ul>
<h2>4. Watch for Rapport&#8217;s Failure Mode: Mutual People-Pleasing</h2>
<p>Rapport is supposed to make honesty easier, but push it too far and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8535677/">it does the opposite</a>. Managing the session means balancing participant authenticity with your own attempts to connect, which is easier said than done. <a href="https://www.amazon.com/Moderating-Usability-Tests-Interacting-Technologies/dp/0123739330"><em>Moderating Usability Tests</em></a> (Dumas) boils the guidance down to two directives: be professional, be genuine.</p>
<p>Watch for it in yourself first: You like this participant. They&#8217;re articulate, friendly, and easy to talk to. So you soften a follow-up you&#8217;d normally push on, or let a shaky answer slide because pressing feels rude. That&#8217;s not rapport anymore. That&#8217;s you people-pleasing the participant.</p>
<p>It runs the other way, too. An interviewer who&#8217;s a little too warm signals something the participant picks up on fast: <em>I like you, and I want this to go well.</em> The participant hears that as permission to perform rather than react; to keep telling you what a great job the product is doing, because that&#8217;s what a good, agreeable participant does for a researcher they like.</p>
<p>Same warmth. Same instinct. Opposite direction. Both land in the same place: a session that feels great and tells you almost nothing.</p>
<p>The fix isn&#8217;t less rapport. It&#8217;s checking, session by session, <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">which kind of warmth you&#8217;re building</a>. Is it the kind that makes honesty safer, or the kind that makes disagreement feel impolite? Stay friendly. Stay neutral about outcomes.</p>
<h2>5. Use Rapport as a Final Screening Step</h2>
<p>Suppose you’re conducting a usability study of an app for master plumbers who manage other plumbers&#8217; schedules. If you skip the warm-up and jump straight into tasks, you&#8217;re assuming the person is, in fact, a master plumber.</p>
<p>But say you notice they&#8217;re parroting your task language back at you instead of using the shop terms other master plumbers used in the same task. Say they rate everything at the top of the scale, or brush past usability issues that other participants flagged as real workflow problems. A few minutes of background questions at the start (What does a typical day look like? How long have you been in the trade?) would have surfaced that this participant went to trade school but never actually worked as a plumber. That&#8217;s not a rapport failure. That&#8217;s a screening failure that rapport-building would have caught.</p>
<h2>Summary and Discussion</h2>
<p>Using rapport is a core technique of a successful moderator. Five ways to make rapport work:</p>
<ul>
<li><strong>Prepare before the session. </strong>Know your audience and their vocabulary before you say hello.</li>
<li><strong>Break the ice deliberately. </strong>Lower the stakes with neutral small talk before you ask for anything.</li>
<li><strong>Use a verbal and non-verbal toolkit. </strong>Purposeful pauses, neutral reactions, and eye contact on camera keep the session honest.</li>
<li><strong>Watch for people-pleasing on both sides. </strong>Rapport should make disagreement easier, not harder.</li>
<li><strong>Treat rapport as a screening step. </strong>A few warm-up questions can catch a participant who never should have qualified.</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Streamlined Measurement of the UX of AI</title>
		<link>https://measuringu.com/streamlined-measurement-of-the-ux-of-ai/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=streamlined-measurement-of-the-ux-of-ai</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 22:21:12 +0000</pubDate>
				<category><![CDATA[Metrics]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48188</guid>

					<description><![CDATA[There’s nothing quite the same as chatting with AI. But it&#8217;s just like every other interface in one important way: we experience it. And what we experience, we can measure. Measuring the user experience means assessing both what people do and what people think (a mix of action and attitudinal metrics). For attitudes, standardized questionnaires [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48262 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-300x169.jpg" alt="" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>There’s nothing quite the same as chatting with AI. But it&#8217;s just like every other interface in one important way: we experience it. And what we experience, we can measure.</p>
<p>Measuring the user experience means assessing both what people do and what people think (a mix of action and <a href="https://measuringu.com/ux-attitudes/">attitudinal</a> metrics). For attitudes, standardized questionnaires like the UX-Lite<sup>®</sup> are a good place to start, but they&#8217;re not diagnostic on their own and won&#8217;t tell you why Trust is low or what&#8217;s driving Anxiety.</p>
<p>Building a standardized questionnaire involves assessing <a href="https://www.researchgate.net/publication/200085994_IBM_Computer_Usability_Satisfaction_Questionnaires_Psychometric_Evaluation_and_Instructions_for_Use">validity, reliability, and sensitivity</a>. Practically, this involves six broad steps:</p>
<ol>
<li>Identifying the items to ask participants</li>
<li>Collecting data on the candidate items</li>
<li>Checking that the items group the way you expect</li>
<li>Selecting the best items</li>
<li>Confirming a shorter version is still reliable</li>
<li>Demonstrating the items can tell good experiences from bad ones</li>
</ol>
<p>We did step one in our <a href="https://measuringu.com/measuring-the-ux-of-ai/">previous article</a>, where we identified 34 candidate items across six constructs (AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption) based on our reading and experience with these products.</p>
<p>In this article, we&#8217;ll cover the next five steps: collecting data from real users of real products and seeing how well those 34 items perform.</p>
<h2><span lang="EN-US">AI-Based Chat Software Benchmark Study</span></h2>
<p>In May 2026, we conducted a <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">retrospective study of four AI-based chat software products</a> with 420 U.S.-based panel participants. This study included the metrics we typically collect in our <a href="https://measuringu.com/consumer-software-ux-2025/">standard UX and NPS study of consumer software</a>.</p>
<p>There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several existing questionnaires, including the <a href="https://measuringu.com/nps-ux/">NPS</a>, <a href="https://measuringu.com/10-things-sus/">SUS</a>, <a href="https://measuringu.com/from-umux-lite-to-ux-lite/">UX-Lite</a>, and <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" />. In addition to these existing questionnaires, respondents completed 34 new items focused on different aspects of the UX of AI. The AI-based chat products and sample sizes were:</p>
<ul>
<li>ChatGPT: 113</li>
<li>Claude: 103</li>
<li>Gemini: 101</li>
<li>Grok: 103</li>
</ul>
<p>The sample sizes are modest but adequate to <a href="https://measuringu.com/might-not-be-a-magic-number-but-there-are-magic-ranges/">establish baselines and identify medium-sized differences </a>relative to each other and other software products we measure (e.g., ± 5 for 0–100-point rating scales). With a combined <em>n</em> = 420, the sample size is also large enough to support <a href="https://measuringu.com/advanced-stats/">advanced analyses</a> (e.g., factor analysis, regression analysis, reliability analysis, ANOVA).</p>
<h2><span lang="EN-US">Factor Analysis: Items (Almost) Perfectly Aligned with Target Constructs</span></h2>
<p>We used the multivariate technique of factor analysis to determine how well the candidate items map to their intended construct. The output of a factor analysis is <em>factor loadings</em>, numbers that range from −1 to +1, which are interpreted like a correlation. To review the factor loadings for all 34 items, see Appendix Figure 1.</p>
<p>All but one of the items (“I often rely on AI chatbots to perform tasks that I would otherwise do myself”) didn’t map well to its intended construct of Trust (it loaded as high on Productivity), so this was a good candidate to exclude. The remaining 33 items strongly loaded on their intended constructs—<strong>solid evidence of construct validity</strong>.</p>
<h2><span lang="EN-US">Item Analyses: Selecting the Best Items for a Streamlined Questionnaire</span></h2>
<p>A common approach to item selection in the psychometric process of developing a standardized questionnaire is to retain the items with the highest loadings on their associated factors. In our current practice, we enhance that method by examining item means and beta weights from regression models with a key outcome variable.</p>
<p>While still paying attention to the magnitude of item loadings, we also try to select items that vary in their observed means to better differentiate low and high levels of the construct. We also created six regression models, one for each construct, to see by examining beta weights which items accounted for larger amounts of variation in a key outcome metric: the likelihood to continue using the product.</p>
<p>In the following sections, we present the three criteria for all the items for each target construct to guide the items we selected for inclusion in a final, streamlined questionnaire. For easier interpretation, the means in the tables were <a href="https://measuringu.com/converting-scales-to-100-points/">converted from their original five-point scale to a scale ranging from 0 to 100</a>. Our typical target for measuring a construct at this stage of questionnaire development is to select two to three items per construct with the goal of achieving scale reliability (measured with coefficient alpha) of at least 0.70 for each construct.</p>
<p>The selected items are at the top of each table. Table cells are highlighted for the three largest loadings, the lowest and highest means, and the three largest beta weights. We present all the items and their scores, as you may choose to try out alternative combinations of items for your own questionnaire (<a href="https://measuringu.com/contact/">call us</a> if you need to talk it through!).</p>
<h3><span lang="EN-US">AI Productivity</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48078" src="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png" alt="Image showing man working at computer" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>For AI Productivity, we selected two of the highest loading items, one of which also had the highest beta weight. The third selected item balanced an acceptably high loading and beta weight plus a relatively low mean.</p>

<table id="tablepress-1066" class="tablepress tablepress-id-1066">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Using this AI chatbot greatly improves my productivity.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.959</p></td><td class="column-4"><center>69.2</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .259</p></td>
</tr>
<tr class="row-3">
	<td class="column-1">Using this AI chatbot makes me feel more capable in my work or studies.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.904</p></td><td class="column-4"><center>65.1</td><td class="column-5"><center>−.091</td>
</tr>
<tr class="row-4">
	<td class="column-1">I feel comfortable being accountable for work that used this AI chatbot.</td><td class="column-2"><center>x</td><td class="column-3"><center>.568</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">63.3</p></td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .183</p></td>
</tr>
<tr class="row-5">
	<td class="column-1">This AI chatbot adds substantial value to my personal tasks.</td><td class="column-2"></td><td class="column-3"><center>.701</td><td class="column-4"><center>64.9</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .245</p></td>
</tr>
<tr class="row-6">
	<td class="column-1">This AI chatbot adds substantial value to my professional tasks.</td><td class="column-2"></td><td class="column-3"><center>.884</td><td class="column-4"><center>63.7</td><td class="column-5"><center> .041</td>
</tr>
<tr class="row-7">
	<td class="column-1">Using this AI chatbot helps me achieve my goals.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.959</p></td><td class="column-4"><center>67.5</td><td class="column-5"><center> .079</td>
</tr>
<tr class="row-8">
	<td class="column-1">The amount of time it takes for this AI chatbot to respond is acceptable.</td><td class="column-2"></td><td class="column-3"><center>.524</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">78.5</p></td><td class="column-5"><center> .002</td>
</tr>
<tr class="row-9">
	<td class="column-1">This AI chatbot’s responses efficiently tell me the information I need.</td><td class="column-2"></td><td class="column-3"><center>.618</td><td class="column-4"><center>72.8</td><td class="column-5"><center> .144</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>AI Productivity items.</p>
<h3><span lang="EN-US">AI Trust</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48080" src="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png" alt="Image showing trust in AI" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items selected for AI Trust all had acceptably high loadings. The top two had impressive beta weights but little difference in their means (64.6, 62.7), so the third item was included to extend the lower range of the item means to 49.4.</p>

<table id="tablepress-1067" class="tablepress tablepress-id-1067">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I trust this AI chatbot to provide reliable information.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.654</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">64.6</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .344</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel confident relying on responses from this AI chatbot when making decisions.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.606</td><td class="column-4"><center>62.7</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .295</td>
</tr>
<tr class="row-4">
	<td class="column-1">It’s easy to understand what happens to the information I share with this AI chatbot.</td><td class="column-2"><center>x</td><td class="column-3"><center>.570</td><td class="column-4"><center>49.4</td><td class="column-5"><center> .008</td>
</tr>
<tr class="row-5">
	<td class="column-1">This AI chatbot always provides accurate responses.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.813</td><td class="column-4"><center>56.4</td><td class="column-5"><center> .003</td>
</tr>
<tr class="row-6">
	<td class="column-1">When this AI chatbot makes mistakes, they are usually easy to detect.</td><td class="column-2"></td><td class="column-3"><center>.417</td><td class="column-4"><center>58.0</td><td class="column-5"><center>−.019</td>
</tr>
<tr class="row-7">
	<td class="column-1">I don’t worry about how my data is used when interacting with this AI chatbot.</td><td class="column-2"></td><td class="column-3"><center>.391</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">44.8</td><td class="column-5"><center> .011</td>
</tr>
<tr class="row-8">
	<td class="column-1">My professional value is not affected by products like this AI chatbot.</td><td class="column-2"></td><td class="column-3"><center>.372</td><td class="column-4"><center>61.1</td><td class="column-5"><center> .076</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>AI Trust items.</p>
<h3><span lang="EN-US">AI Dependency</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48079" src="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png" alt="Image showing a man lounging while an AI does the work" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>There were only three items developed for AI Dependency, and we excluded one of them, “I often rely on AI chatbots to perform tasks that I would otherwise do myself,” because it loaded on both Trust and Productivity. There were no issues warranting exclusion of the other two items, so we kept them both.</p>

<table id="tablepress-1068" class="tablepress tablepress-id-1068">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I tend to accept answers from AI chatbots without verifying their accuracy.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.983</td><td class="column-4"><center>37.3</td><td class="column-5"><center>.034</td>
</tr>
<tr class="row-3">
	<td class="column-1">I rarely double-check information provided by AI chatbots.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.852</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">35.1</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">.108</td>
</tr>
<tr class="row-4">
	<td class="column-1">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td><td class="column-2"></td><td class="column-3"><center>.363</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">49.2</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">.250</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 3: </strong>AI Dependency items.</p>
<h3><span lang="EN-US">AI Anxiety</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48081" src="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png" alt="Image showing nervous man watching computer." width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items retained for AI Anxiety had the highest loadings of the set, reasonably impactful beta weights, and a reasonable range of means.</p>

<table id="tablepress-1069" class="tablepress tablepress-id-1069">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">The increasing use of AI makes me uneasy.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.822</td><td class="column-4"><center>55.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.259</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am often concerned that AI could cause serious harm to society.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.899</td><td class="column-4"><center>58.9</td><td class="column-5"><center>−.095</td>
</tr>
<tr class="row-4">
	<td class="column-1">AI development feels difficult to control.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.836</td><td class="column-4"><center>60.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .126</td>
</tr>
<tr class="row-5">
	<td class="column-1">I often worry about the environmental impact of AI.</td><td class="column-2"></td><td class="column-3"><center>.780</td><td class="column-4"><center>61.8</td><td class="column-5"><center>−.044</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI development feels risky.</td><td class="column-2"></td><td class="column-3"><center>.821</td><td class="column-4"><center>57.0</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.192</td>
</tr>
<tr class="row-7">
	<td class="column-1">There should be more government regulation for AI development.</td><td class="column-2"></td><td class="column-3"><center>.794</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">66.5</td><td class="column-5"><center> .009</td>
</tr>
<tr class="row-8">
	<td class="column-1">Using AI chatbots for work or school feels unethical.</td><td class="column-2"></td><td class="column-3"><center>.603</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">47.1</td><td class="column-5"><center>−.021</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 4: </strong>AI Anxiety items.</p>
<h3><span lang="EN-US">AI Personification</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48082" src="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png" alt="Image showing man speaking to an angelic AI incarnation." width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items selected for AI Personification had acceptably high loadings, significant beta weights, and a reasonable range of means.</p>

<table id="tablepress-1070" class="tablepress tablepress-id-1070">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Interacting with this AI chatbot feels like communicating with a human.</td><td class="column-2"><center>x</td><td class="column-3"><center>.593</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">45.1</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .249</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel like AI chatbots understand me well.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.777</td><td class="column-4"><center>43.5</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .162</td>
</tr>
<tr class="row-4">
	<td class="column-1">I tend to feel a sense of connection when interacting with AI chatbots.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.985</td><td class="column-4"><center>35.5</td><td class="column-5"><center> .145</td>
</tr>
<tr class="row-5">
	<td class="column-1">I tend to feel like I’m socializing when I interact with AI chatbots.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">1.050</td><td class="column-4"><center>34.5</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.299</td>
</tr>
<tr class="row-6">
	<td class="column-1">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td><td class="column-2"></td><td class="column-3"><center>.712</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">33.9</td><td class="column-5"><center> .151</td>
</tr>
<tr class="row-7">
	<td class="column-1">I’m more likely to share personal information with AI chatbots than with other people.</td><td class="column-2"></td><td class="column-3"><center>.684</td><td class="column-4"><center>36.6</td><td class="column-5"><center> .081</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 5: </strong>AI Personification items.</p>
<h3><span lang="EN-US">Early Adoption</span></h3>
<p>We did not see any issues that warranted excluding any of these items, so we kept all three.</p>

<table id="tablepress-1071" class="tablepress tablepress-id-1071">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I like to experiment with new technologies before most people do.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.961</td><td class="column-4"><center>59.9</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .118</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am usually among the first to try new digital tools.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.942</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">55.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .171</td>
</tr>
<tr class="row-4">
	<td class="column-1">I actively seek out new technologies to try.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.920</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">62.4</td><td class="column-5"><center>−.087</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 6: </strong>Early Adoption items.</p>
<h2><span lang="EN-US">Reliability Analysis: All Streamlined Scale Reliabilities Exceeded 0.80</span></h2>
<p>Table 7 shows the coefficient alpha values for each measure for all items and for the streamlined item set. For research, the typical reliability goal is to exceed 0.70. There was little reduction in reliability for the streamlined versions and for AI Dependency; eliminating its one problematic item increased its reliability even though only two items were retained. The reliabilities for all the streamlined versions of the questionnaires not only met the typical research goal but exceeded 0.80—<strong>strong evidence of reliability for these new scales and statistical justification for the selection of their constituent items</strong>.</p>

<table id="tablepress-1072" class="tablepress tablepress-id-1072">
<thead>
<tr class="row-1">
	<th class="column-1">Reliability (Coefficient Alpha)</th><th class="column-2"><center>All Items</th><th class="column-3"><center>Streamlined</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">AI Productivity</td><td class="column-2"><center>0.92</td><td class="column-3"><center>0.85</td>
</tr>
<tr class="row-3">
	<td class="column-1">AI Trust</td><td class="column-2"><center>0.82</td><td class="column-3"><center>0.81</td>
</tr>
<tr class="row-4">
	<td class="column-1">AI Dependency</td><td class="column-2"><center>0.76</td><td class="column-3"><center>0.84</td>
</tr>
<tr class="row-5">
	<td class="column-1">AI Anxiety</td><td class="column-2"><center>0.91</td><td class="column-3"><center>0.87</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI Personification</td><td class="column-2"><center>0.92</td><td class="column-3"><center>0.86</td>
</tr>
<tr class="row-7">
	<td class="column-1">Early Adoption</td><td class="column-2"><center>0.94</td><td class="column-3"><center>NA</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 7: </strong>Scale reliabilities (coefficient alpha) for the six new metrics.</p>
<h2><span lang="EN-US">Profile Analysis: Claude Leads in Productivity, ChatGPT Lags in Trust</span></h2>
<p>We created two different visualizations of profiles for the streamlined scales, showing how the AI assistants compare. The line graph in Figure 1 makes it easy to see at a glance which scales differentiate among the products, while the column chart in Figure 2 makes it easy to compare the confidence intervals around the means.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48264 size-large" src="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1024x361.jpg" alt="Line graph of AI scale scores for four generative AI chatbots." width="1024" height="361" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1024x361.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-300x106.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-768x271.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1536x541.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-2048x721.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-600x211.jpg 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Line graph of AI scale scores for four generative AI chatbots.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48275 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg" alt="Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals." width="1525" height="590" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg 1525w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-300x116.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-1024x396.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-768x297.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-600x232.jpg 600w" sizes="auto, (max-width: 1525px) 100vw, 1525px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals.</p>
<p>The results show Claude leading in AI Productivity, ChatGPT lagging in AI Trust, and little difference among the products for AI Dependency. Grok scored the lowest in AI Anxiety and the highest in AI Personification, and all four products scored different levels of Early Adoption (highest for Grok, lowest for ChatGPT).</p>
<p>A mixed ANOVA of the ratings indicated a significant main effect of scale (<em>F</em>(5, 2080) = 85.4, <em>p</em> &lt; .0001), a significant main effect of product (<em>F</em>(3, 416) = 2.8, <em>p</em> = .039), and most importantly, a highly significant scale by product interaction (<em>F</em>(15, 2080) = 3.6, <em>p </em>&lt; .0001). The interaction is important because it is <strong>strong statistical evidence of the sensitivity of these new scales</strong>.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>We collected data from 420 respondents for 34 items designed to measure six constructs related to attitudes toward four generative AI chatbots (ChatGPT, Claude, Gemini, and Grok). We then conducted analyses to complete the psychometric measurement goals of construct validity (factor analysis), measurement efficiency (item selection), scale reliability (coefficient alpha), and scale sensitivity (ANOVA).</p>
<p>The key points are:</p>
<p><strong>We have solid evidence of construct validity for the new questionnaires. </strong>Our factor analysis of the data demonstrated almost perfect alignment of items with their intended constructs for the measurement of AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. One item loaded on two factors and was thus not retained in the streamlined versions of the questionnaires.</p>
<p><strong>The items selected for streamlined versions of the questionnaires produced reliable measurement. </strong>For each of the six constructs, we retained two to three items to balance item loadings, item mean ranges, and strong beta weights for their relationships with the likelihood to continue using the product. For all the streamlined versions of the questionnaires, coefficient alpha exceeded 0.80 (ranging from 0.81 to 0.87). The common criterion for acceptable reliability is &gt; 0.70.</p>
<p><strong>The statistical evidence for scale sensitivity is strong. </strong>Profile analysis of the mean ratings of the new scales by product indicated a significant main effect of scale, a significant main effect of product, and a highly significant product-by-scale interaction. Claude led in AI Productivity, and ChatGPT lagged in AI Trust. There was little difference among the products for AI Dependency, Grok scored the lowest for AI Anxiety and the highest for AI Personification, and there were different levels of Early Adoption for all four products.</p>
<h2><span lang="EN-US">Appendix: Factor Structure and Item Key</span></h2>
<p>Appendix Figure 1 shows the pattern matrix from the factor analysis for the original 34 items. The numbers in the figure are item loadings, which indicate the degree of connection of the item with the target constructs (with values ranging from -1 to +1, interpreted like a correlation). The usual criterion for a meaningfully large loading is anything more extreme than ±0.3. A maximum likelihood factor analysis with Promax rotation using SPSS 23 was used following a <a href="https://en.wikipedia.org/wiki/Parallel_analysis">parallel analysis</a> that indicated, as expected, retention of six factors. A key pattern to look for when identifying problematic items is any item with strong loading on more than one factor. This happened to OftenRelyOnChatbots (“I often rely on AI chatbots to perform tasks that I would otherwise do myself,” highlighted in yellow), leading to its exclusion from the final questionnaires.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48206" src="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png" alt="Alignment of items with constructs (values greater than 0.3 are highlighted in green). To see the complete item text, refer to the appendix. " width="624" height="533" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png 624w, https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1-300x256.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1-600x513.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 1: </strong>Alignment of items with constructs (values greater than 0.3 are highlighted in green). <strong>To see the complete item text, refer to Appendix Table 1. </strong></p>
<p>Appendix Table 1 documents the short labels for each item. Items highlighted in green are the ones retained for the streamlined versions of these questionnaires.</p>

<table id="tablepress-1073" class="tablepress tablepress-id-1073 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<th class="column-1">Short Label</th><th class="column-2">Item</th>
</tr>
</thead>
<tbody class="row-hover">
<tr class="row-2">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI PRODUCTIVITY</td>
</tr>
<tr class="row-3">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">ImprovedProductivity</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Using this AI chatbot greatly improves my productivity.</td>
</tr>
<tr class="row-4">
	<td class="column-1">AddsValuePersonal</td><td class="column-2">This AI chatbot adds substantial value to my personal tasks.</td>
</tr>
<tr class="row-5">
	<td class="column-1">AddsValueProfessional</td><td class="column-2">This AI chatbot adds substantial value to my professional tasks.</td>
</tr>
<tr class="row-6">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIMakesMeFeelMoreCapable</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Using this AI chatbot makes me feel more capable in my work or studies.</td>
</tr>
<tr class="row-7">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">IFeelAccountableForMyAIAssistedWork</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel comfortable being accountable for work that used this AI chatbot.</td>
</tr>
<tr class="row-8">
	<td class="column-1">AIHelpsMeAchieveMyGoals</td><td class="column-2">Using this AI chatbot helps me achieve my goals.</td>
</tr>
<tr class="row-9">
	<td class="column-1">AIResponseTimeIsAcceptable</td><td class="column-2">The amount of time it takes for this AI chatbot to respond is acceptable.</td>
</tr>
<tr class="row-10">
	<td class="column-1">AIResponsesAreEfficient</td><td class="column-2">This AI chatbot’s responses efficiently tell me the information I need.</td>
</tr>
<tr class="row-11">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI TRUST</td>
</tr>
<tr class="row-12">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">TrustReliableInfo</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I trust this AI chatbot to provide reliable information.</td>
</tr>
<tr class="row-13">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SupportsConfidentDecisions</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel confident relying on responses from this AI chatbot when making decisions.</td>
</tr>
<tr class="row-14">
	<td class="column-1">AlwaysAccurate</td><td class="column-2">This AI chatbot always provides accurate responses.</td>
</tr>
<tr class="row-15">
	<td class="column-1">EasyToDetectMistakes</td><td class="column-2">When this AI chatbot makes mistakes, they are usually easy to detect.</td>
</tr>
<tr class="row-16">
	<td class="column-1">NotWorriedAboutDataUse</td><td class="column-2">I don’t worry about how my data is used when interacting with this AI chatbot.</td>
</tr>
<tr class="row-17">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">EasyToUnderstandHowInfoIsShared</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">It’s easy to understand what happens to the information I share with this AI chatbot.</td>
</tr>
<tr class="row-18">
	<td class="column-1">ProfValueNotAffected</td><td class="column-2">My professional value is not affected by products like this AI chatbot.</td>
</tr>
<tr class="row-19">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI DEPENDENCE</td>
</tr>
<tr class="row-20">
	<td class="column-1">OftenRelyOnChatbots</td><td class="column-2">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td>
</tr>
<tr class="row-21">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AcceptAnswersWithoutVerification</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I tend to accept answers from AI chatbots without verifying their accuracy.</td>
</tr>
<tr class="row-22">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">RarelyDoubleCheck</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I rarely double-check information provided by AI chatbots.</td>
</tr>
<tr class="row-23">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI ANXIETY</td>
</tr>
<tr class="row-24">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">UneasyWithIncreasingUse</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">The increasing use of AI makes me uneasy.</td>
</tr>
<tr class="row-25">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SeriousSocietalHarm</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I am often concerned that AI could cause serious harm to society.</td>
</tr>
<tr class="row-26">
	<td class="column-1">EnvironmentalImpact</td><td class="column-2">I often worry about the environmental impact of AI.</td>
</tr>
<tr class="row-27">
	<td class="column-1">AIDevelopmentRisky</td><td class="column-2">AI development feels risky.</td>
</tr>
<tr class="row-28">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIDevelopmentHardToControl</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">AI development feels difficult to control.</td>
</tr>
<tr class="row-29">
	<td class="column-1">NeedMoreGovRegulation</td><td class="column-2">There should be more government regulation for AI development.</td>
</tr>
<tr class="row-30">
	<td class="column-1">UseForWorkOrSchoolUnethical</td><td class="column-2">Using AI chatbots for work or school feels unethical.</td>
</tr>
<tr class="row-31">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI PERSONIFICATION</td>
</tr>
<tr class="row-32">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIWorkFeelsLikeHumanCommunication</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Interacting with this AI chatbot feels like communicating with a human.</td>
</tr>
<tr class="row-33">
	<td class="column-1">SometimesAIFeelsLikeAFriend</td><td class="column-2">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td>
</tr>
<tr class="row-34">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">ChatbotsUnderstandMeWell</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel like AI chatbots understand me well.</td>
</tr>
<tr class="row-35">
	<td class="column-1">FeelSenseOfConnection</td><td class="column-2">I tend to feel a sense of connection when interacting with AI chatbots.</td>
</tr>
<tr class="row-36">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">FeelsLikeSocializing</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I tend to feel like I’m socializing when I interact with AI chatbots.</td>
</tr>
<tr class="row-37">
	<td class="column-1">MoreLikelyToSharePersonalInfo</td><td class="column-2">I’m more likely to share personal information with AI chatbots than with other people.</td>
</tr>
<tr class="row-38">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">EARLY ADOPTION</td>
</tr>
<tr class="row-39">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">IExperimentBeforeOthers</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I like to experiment with new technologies before most people do.</td>
</tr>
<tr class="row-40">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">FirstToTryNewDigitalTools</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I am usually among the first to try new digital tools.</td>
</tr>
<tr class="row-41">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SeekNewTechnologiesToTry</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I actively seek out new technologies to try.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 1: </strong>Short labels and full text for each item.</p>
<p>&nbsp;</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Using AI To Find Usability Problems: A Replication</title>
		<link>https://measuringu.com/does-ai-find-real-usability-problems-a-replication/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=does-ai-find-real-usability-problems-a-replication</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 22:36:28 +0000</pubDate>
				<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Problem Discovery]]></category>
		<category><![CDATA[Usability Problem]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48115</guid>

					<description><![CDATA[If AI can code problems and potentially moderate a structured interview, can it also reliably find usability issues from watching a video? We can talk about hypotheticals, or we can actually try it out. We tried it out. In our first study, we started small. We compared human and AI reviews of a video taken [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48154 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-300x169.jpg" alt="Feature image showing AI finding usability problems" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>If AI can <a href="https://measuringu.com/classification-agreement-between-ux-researchers-and-chatgpt/">code problems</a> and potentially <a href="https://measuringu.com/can-we-trust-ai-to-moderate-ux-interviews/">moderate a structured interview</a>, can it also reliably find usability issues from watching a video?</p>
<p>We can talk about hypotheticals, or we can actually try it out. We tried it out.</p>
<p>In our <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">first study</a>, we started small. We compared human and AI reviews of a video taken from a usability test on making dinner reservations with OpenTable. We used the standard ChatGPT and Gemini interfaces (standard Claude doesn’t support video review).</p>
<p><strong>The good:</strong> AI identified roughly <strong>half</strong> the usability problems identified by human UX researchers. It even uncovered one genuine problem not found by any of the four human evaluators.</p>
<p><strong>The bad:</strong> AI identified seven false alarms. They were true statements but not connected to the participants’ actions or utterances.</p>
<p><strong>The ugly:</strong> Three problems were straight-up hallucinations. AI described events that just didn’t happen.</p>
<p>So, of the eleven problems the two AIs reported that no human flagged, only one was a genuine problem. That&#8217;s a useful number to keep in mind: 91% of the AI-only problems in this study required either correction or dismissal.</p>
<p>A case study like that, however, is limited. There was one video, two AIs (ChatGPT, Gemini), and one prompt. Replication helps with generalization.</p>
<p>To move beyond that first case study, we kept all experimental variables the same and conducted the same type of analysis on a similar video from a different usability study.</p>
<h2><span lang="EN-US">Experimental Design: One Human Researcher, Two AIs, and One Video</span></h2>
<p>For this research, one UX researcher with decades of experience (Jim, the lead researcher from the previous study) reviewed a similar video from a previous usability benchmark study of online pet websites multiple times, creating a timeline of key events and a list of observed usability problems for this participant (referred to as Participant A; see Appendix A for the details of the timeline).</p>
<p>The characteristics the new video shared with the first one were:</p>
<ul>
<li>A similar task (steps toward booking a specified appointment)</li>
<li>A similar outcome (the user experienced several usability problems but was ultimately successful).</li>
</ul>
<p>Using the same prompt each time, we then ran the video four times each through the same two AIs used in the previous case study (ChatGPT-5.4 Thinking and Gemini 3 Flash Thinking).</p>
<p>So, in this study, we held constant the videos, the key elements of the prompt, and the AI versions/settings—variables that we eventually plan to vary. This time, as in the first case study, we only varied the type of analyst: human, ChatGPT, and Gemini.</p>
<h3><span lang="EN-US">The Task</span></h3>
<p>During the previous usability benchmark study conducted at MeasuringU in 2019, participants used the PetSmart website to start the process of booking a grooming appointment in Glendale, CO, for a bath and full haircut for an English Springer Spaniel older than six months. The task was successfully completed if the participant found that specific grooming option and reported the listed price of $61. For the step-by-step details of the &#8220;happy path&#8221; to complete this task, see Appendix B.</p>
<h3><span lang="EN-US">The Prompt</span></h3>
<p>The prompt we used for this study was:</p>
<blockquote><p><em>During a usability test, the facilitator must keep track of participant behaviors as they navigate through tasks on a website, mobile app, software program, etc. We’d like you to watch a video of a usability test where participants were asked to book a grooming reservation for their dog. As you&#8217;re watching, please look for problems the participant has while attempting to complete the task. For example, you can document the path users take, describe issues they encounter as well as what on the website might be causing problems. The task has been successfully completed if the participant finds the target service (“Bath &amp; Full Haircut” which costs $61; not “Bath &amp; Full Haircut with FURminator” which costs $74). If you understand these instructions, let me know and I&#8217;ll drag the video in for you to review. Are you ready for the video?</em></p></blockquote>
<p>This was based on the prompt we used for the OpenTable case study with slight modifications. The task details are necessarily different because there was only one correct choice for this new task (compared to many correct choices for the OpenTable task), so we specified the end task details required for successful completion.</p>
<h2><span lang="EN-US">Major Findings</span></h2>
<p>Even though the participant successfully completed the task in the video, a total of ten usability problems were identified by the AIs and the human researcher.</p>
<p>Usability problems aren’t like observing a visual defect in a product. They require judgement, so it’s worth digging into what these problems are because they&#8217;re at the crux of how AI may or may not be able to effectively emulate UX researcher judgement (for now). So, let’s dig into what we saw.</p>
<p>Participant A started with a few clicks not on the happy path, but on the third click got the task started, so these two early clicks could be considered minor usability problems that were quickly corrected.</p>
<p>There were two more impactful usability problems, both associated with breed selection (Figure 1). First, there was no visible indication that the dropdown list could be filtered by typing over the word &#8220;breed&#8221;; the participant scrolled through the entire list. Second, the order of presentation of the breeds in the dropdown list was alphabetically inconsistent (e.g., &#8220;English Toy Spaniel&#8221; started with E; &#8220;Springer Spaniel &#8211; English&#8221; started with S). Despite this, the participant successfully completed the task, just not on the most efficient path.</p>
<p><strong>a:</strong> Breed list at boundary of D and E—&#8221;English Toy Spaniel&#8221; but no English Springer Spaniel</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48116" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png" alt="Image showing breed list with English Toy Spaniel but not English Springer Spaniel" width="507" height="342" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png 507w, https://measuringu.com/wp-content/uploads/2026/08/081126-F1a-300x202.png 300w" sizes="auto, (max-width: 507px) 100vw, 507px" /></a></p>
<p><strong>b</strong>: Location of &#8220;Springer Spaniel &#8211; English&#8221;</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48117" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png" alt="Breed list showing Springer Spaniel - English" width="531" height="315" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png 531w, https://measuringu.com/wp-content/uploads/2026/08/081126-F1b-300x178.png 300w" sizes="auto, (max-width: 531px) 100vw, 531px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Participant A&#8217;s usability issues with breed selection.</p>
<h3><span lang="EN-US">Some Agreement Between Human and AI, No Hallucinations, Six False Alarms</span></h3>
<p>Table 1 shows the four problems discovered by the human researcher (three of which were also identified by the AIs) and the six additional problems reported only by the AIs. We looked to see if the problems identified only by the AIs were false alarms (an event happened but not really a usability problem) or hallucinations (the event just didn&#8217;t happen).</p>
<p>The good news was that none of the six problems were hallucinations. The bad news is that they were all determined to be false alarms. Either the participant never actually noticed the issue flagged by the AI or wasn&#8217;t affected by it (e.g., the location prompt, the below-the-fold item), or it was just normal, expected system behavior rather than a flaw (e.g., the ZIP search returning multiple locations, the menu reloading).</p>

<table id="tablepress-1061" class="tablepress tablepress-id-1061">
<thead>
<tr class="row-1">
	<th class="column-1">#</th><th class="column-2">Problem description</th><th class="column-3">Human</th><th class="column-4">ChatGPT</th><th class="column-5">Gemini</th><th class="column-6">Why (if false alarm)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"><center>Y</td><td class="column-4"><center>--</td><td class="column-5"><center>--</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-4">
	<td class="column-1">H3</td><td class="column-2">No attempt to filter breed list by typing; scrolled instead</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-5">
	<td class="column-1">H4</td><td class="column-2">Searched breed list for Springer Spaniel, couldn't find it alphabetically</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-6">
	<td class="column-1">C1</td><td class="column-2">Blank/loading state after selecting Grooming Salon</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Lasted ~2 seconds; participant had already moved on</td>
</tr>
<tr class="row-7">
	<td class="column-1">C2/G1</td><td class="column-2">Browser location prompt appears alongside site's own location modal</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>N</td><td class="column-6">No sign participant noticed it; used site's own ZIP entry instead</td>
</tr>
<tr class="row-8">
	<td class="column-1">C3</td><td class="column-2">Entering ZIP returns multiple grooming locations</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Accurate, but that's how the control is designed to work</td>
</tr>
<tr class="row-9">
	<td class="column-1">C4</td><td class="column-2">Menu reloads after selecting breed and age</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Expected behavior after Check Prices &amp; Book Now; no impact</td>
</tr>
<tr class="row-10">
	<td class="column-1">C5/G2</td><td class="column-2">Target service is at the bottom of the list, below the fold</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>N</td><td class="column-6">True, but not actually a problem for this participant</td>
</tr>
<tr class="row-11">
	<td class="column-1">C6</td><td class="column-2">Promotional content crowds out the service menu</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">True, but not an obvious problem for this participant</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>Summary of usability problem discovery by the human researcher and the AIs. In the # column, H indicates a problem identified by the human researcher, C indicates a problem identified by ChatGPT, and G indicates a problem identified by Gemini. In the Human, ChatGPT, and Gemini columns, Y indicates the discovery of a verified usability problem, and N indicates the reporting of a false alarm. The full runs of ChatGPT and Gemini are shown in Appendix A.</p>
<p>As we did in our first study, we ran the videos four times through the AIs (because of the <a href="https://measuringu.com/can-ai-detect-usability-problems/">probabilistic nature of how they work</a>). Problems identified in any of the four runs were included in this analysis. We used the mean <a href="https://measuringu.com/ai-usability-problem-analysis-of-a-video/">any-2 agreement</a> to assess overlap.</p>
<p><strong><em>Technical note</em></strong><em><strong>:</strong> Our preferred method for quantifying the correspondence between two lists of usability issues is </em><em>any-2 agreement</em><em>. Any-2 agreement is the ratio of the intersection of the two sets divided by their union. Historically, we’ve found an any-2 agreement of 50% to be average (typical), around 25% to be low, and around 75% to be high.</em></p>
<h4><span lang="EN-US">ChatGPT Agreement: 35%</span></h4>
<p>The mean any-2 agreement of the four ChatGPT runs and the UX researcher was 35%. ChatGPT identified (at least once) three of the four problems reported by the researcher but also produced six false alarms.</p>
<h4><span lang="EN-US">Gemini Agreement: 38%</span></h4>
<p>The mean any-2 agreement of the four Gemini runs and the UX researcher was 38%. Gemini identified (at least once) three of the four problems reported by the researcher but also produced two false alarms (matching two of the false alarms produced by ChatGPT).</p>
<p>The mean any-2 agreement between the four runs of the AIs was 40%. Figure 2 shows the Venn diagram for the problem discovery results for the UX researcher and the AIs.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48120" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png" alt="Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini." width="991" height="743" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png 991w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-300x225.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-768x576.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-600x450.png 600w" sizes="auto, (max-width: 991px) 100vw, 991px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini.</p>
<p>The Venn diagram illustrates the relationship between the human reviewer and AI analyses of Participant A. The human reviewer identified four usability issues, three of which were identified at least once by an AI. However, there was <strong>one usability problem that was not caught by either AI</strong>, while the <strong>AIs produced six issues that were not legitimate usability problems</strong> (all false alarms, no hallucinations). The AIs did not discover any real problems that the human reviewer failed to identify.</p>
<h2><span lang="EN-US">Comparison with the OpenTable Case Study Results</span></h2>
<p>Figure 3 shows the Venn diagram from our <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">OpenTable case study</a>.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48121" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png" alt="Venn diagram of usability problem discovery from our OpenTable case study." width="965" height="724" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png 965w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-300x225.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-768x576.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-600x450.png 600w" sizes="auto, (max-width: 965px) 100vw, 965px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3: </strong>Venn diagram of usability problem discovery from our OpenTable case study.</p>
<p>At a glance, the diagrams in Figures 2 and 3 have some similarities and some differences. To help with comparison, Table 2 shows the side-by-side comparisons for different aspects of the results (graphed in Figure 4).</p>

<table id="tablepress-1062" class="tablepress tablepress-id-1062">
<thead>
<tr class="row-1">
	<th class="column-1">Comparison</th><th class="column-2">Participant A</th><th class="column-3">OpenTable</th><th class="column-4">Abs. Diff.</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">% human identified</td><td class="column-2">40%</td><td class="column-3">45%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-3">
	<td class="column-1">% AI &amp; human overlap</td><td class="column-2">30%</td><td class="column-3">30%</td><td class="column-4"> 0%</td>
</tr>
<tr class="row-4">
	<td class="column-1">% AI-only identified</td><td class="column-2">60%</td><td class="column-3">55%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-5">
	<td class="column-1">% AI errors</td><td class="column-2">60%</td><td class="column-3">50%</td><td class="column-4">10%</td>
</tr>
<tr class="row-6">
	<td class="column-1">% AI false alarms</td><td class="column-2">60%</td><td class="column-3">35%</td><td class="column-4">25%</td>
</tr>
<tr class="row-7">
	<td class="column-1">% AI hallucinations</td><td class="column-2"> 0%</td><td class="column-3">15%</td><td class="column-4">15%</td>
</tr>
<tr class="row-8">
	<td class="column-1">% AI-only discovery</td><td class="column-2"> 0%</td><td class="column-3"> 5%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-9">
	<td class="column-1">Total unique problems</td><td class="column-2">10</td><td class="column-3">20</td><td class="column-4">10</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>Comparison of PetSmart Participant A and OpenTable problem identification rates.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48157 size-large" src="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1024x525.jpg" alt="" width="1024" height="525" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1024x525.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-300x154.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-768x393.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1536x787.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-2048x1049.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-600x307.jpg 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4: </strong>PetSmart Participant A and OpenTable problem identification rates.</p>
<p>In most respects, <strong>this second evaluation of AI problem discovery has replicated the first case study</strong> (OpenTable).</p>
<p>The usability problem identification rates were similar (within 10 percentage points) for the percentages of usability problems identified by the human UX researchers, AIs, both (the AI/Human overlap), and AI-only discovery.</p>
<p>Observed differences in the patterns were due to the incidence of AI hallucinations in the OpenTable case study (3) compared to none in the AI outputs for Participant A.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Our key findings were:</p>
<p><strong>Successful replication of the OpenTable case study.</strong> Most metrics comparing the two videos were within 10 percentage points of each other (e.g., 40% vs. 45% of human-identified problems found by AI). The one real divergence was in the <em>type</em> of AI-only errors: this PetSmart video had more false alarms (60% vs. 35%) but zero hallucinations.</p>
<p><strong>A troubling number of false alarms.</strong> Focusing on the new data for PetSmart Participant A, the total number of unique usability problems was ten, of which only four were identified by the UX researcher (which we treat as ground truth). Of the six unique problems the AIs reported that the human researcher did not flag, all were false alarms (no hallucinations).</p>
<p><strong>AI adds potential value as an overly enthusiastic junior researcher, not a trusted expert.</strong> In the analysis of these videos, the AIs discovered three of the four usability problems reported by the UX researcher. Relying only on these multiple runs of the AIs would have missed a quarter of the real usability problems. Unlike our earlier research with the restaurant reservation video, where the <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">AIs found one problem that UX researchers missed</a>, the AIs in this study did not identify any valid usability problems that the UX researcher failed to discover.</p>
<p><strong>Like humans, AI usability reviews of videos are prone to the “evaluator effect.”</strong> Just like human evaluators, multiple runs of AI usability evaluations of videos are not perfectly consistent, so it’s good practice to run these evaluations multiple times for consistency checks. Running multiple evaluations and looking for consistency across runs is a practical filter before any human review.</p>
<p><strong>Bottom line—AI usability reviews of videos require human oversight.</strong> In their current form (what we tested), these AI products can add some value to this type of UX research, but more as junior researchers whose actions and conclusions require expert human oversight rather than as trusted experts themselves.</p>
<p><strong>Future research:</strong> Our next step in this research program is to perform the same analyses on two PetSmart videos in which the user experiences were different from PetSmart Participant A and OpenTable regarding task success and number of problems identified by the human UX researcher (one who did not complete the task successfully and one who experienced no problems completing the task).</p>
<h2><span lang="EN-US">Appendix A: Detailed Timeline and Problem-by-Problem Tables</span></h2>
<h3><span lang="EN-US">Key Events Timeline for Participant A</span></h3>
<p>Appendix Table 1 summarizes the key events in the video (compiled by the UX researcher), identifying four problematic events deviating from the “happy” path (see Appendix B).</p>

<table id="tablepress-1063" class="tablepress tablepress-id-1063">
<thead>
<tr class="row-1">
	<th class="column-1">Event #</th><th class="column-2">Timestamp</th><th class="column-3">Summary of key user actions</th><th class="column-4">Notes</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">1</td><td class="column-2">0:00:13</td><td class="column-3">Clicked Shop by Pet from top menu</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-3">
	<td class="column-1">2</td><td class="column-2">0:00:21</td><td class="column-3">Clicked search field</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-4">
	<td class="column-1">3</td><td class="column-2">0:00:24</td><td class="column-3">Clicked Pet Services</td><td class="column-4"></td>
</tr>
<tr class="row-5">
	<td class="column-1">4</td><td class="column-2">0:00:44</td><td class="column-3">Grooming form begins to appear</td><td class="column-4"></td>
</tr>
<tr class="row-6">
	<td class="column-1">5</td><td class="column-2">0:00:45</td><td class="column-3">All form elements except dog/cat buttons appear</td><td class="column-4"></td>
</tr>
<tr class="row-7">
	<td class="column-1">6</td><td class="column-2">0:00:46</td><td class="column-3">Dog/cat buttons appear—participant cursor was headed for breed but when these buttons appeared he changed course toward the dog/cat buttons</td><td class="column-4"></td>
</tr>
<tr class="row-8">
	<td class="column-1">7</td><td class="column-2">0:00:49</td><td class="column-3">Clicked Dog button</td><td class="column-4"></td>
</tr>
<tr class="row-9">
	<td class="column-1">8</td><td class="column-2">0:00:56</td><td class="column-3">Clicked "select" link by Select a Store</td><td class="column-4"></td>
</tr>
<tr class="row-10">
	<td class="column-1">9</td><td class="column-2">0:00:58</td><td class="column-3">Find a Grooming Salon Near You pops up (one field: "Zip Code, City or State" and Search button)</td><td class="column-4"></td>
</tr>
<tr class="row-11">
	<td class="column-1">10</td><td class="column-2">0:00:59</td><td class="column-3">Location permission prompt appeared at top of screen apparently triggered by presentation of the Use My Current Location link in Find a Grooming Salon Near You</td><td class="column-4"></td>
</tr>
<tr class="row-12">
	<td class="column-1">11</td><td class="column-2">0:01:09</td><td class="column-3">Typed zip code from task instructions and clicked Search</td><td class="column-4"></td>
</tr>
<tr class="row-13">
	<td class="column-1">12</td><td class="column-2">0:01:13</td><td class="column-3">List of locations appears with target Glendale at top</td><td class="column-4"></td>
</tr>
<tr class="row-14">
	<td class="column-1">13</td><td class="column-2">0:01:19</td><td class="column-3">Clicked Glendale</td><td class="column-4"></td>
</tr>
<tr class="row-15">
	<td class="column-1">14</td><td class="column-2">0:01:21</td><td class="column-3">Clicked x to clear the location permission prompt</td><td class="column-4"></td>
</tr>
<tr class="row-16">
	<td class="column-1">15</td><td class="column-2">0:01:29</td><td class="column-3">Clicked Breed dropdown, list of breeds appears</td><td class="column-4"></td>
</tr>
<tr class="row-17">
	<td class="column-1">16</td><td class="column-2">0:01:30</td><td class="column-3">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-18">
	<td class="column-1">17</td><td class="column-2">0:01:35</td><td class="column-3">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-19">
	<td class="column-1">18</td><td class="column-2">0:01:56</td><td class="column-3">Scrolling farther down the list found "Springer Spaniel - English"</td><td class="column-4"></td>
</tr>
<tr class="row-20">
	<td class="column-1">19</td><td class="column-2">0:02:00</td><td class="column-3">Clicked Age dropdown</td><td class="column-4"></td>
</tr>
<tr class="row-21">
	<td class="column-1">20</td><td class="column-2">0:02:04</td><td class="column-3">Selected "6 months or older"</td><td class="column-4"></td>
</tr>
<tr class="row-22">
	<td class="column-1">21</td><td class="column-2">0:02:07</td><td class="column-3">Clicked Check Prices &amp; Book Now button</td><td class="column-4"></td>
</tr>
<tr class="row-23">
	<td class="column-1">22</td><td class="column-2">0:02:13</td><td class="column-3">Grooming Salon Menu appeared</td><td class="column-4"></td>
</tr>
<tr class="row-24">
	<td class="column-1">23</td><td class="column-2">0:02:32</td><td class="column-3">Scrolled through list of services to the bottom</td><td class="column-4"></td>
</tr>
<tr class="row-25">
	<td class="column-1">24</td><td class="column-2">0:02:35</td><td class="column-3">Did not click service but said, "Just bath and full haircut, $61, let me write that down."</td><td class="column-4"></td>
</tr>
<tr class="row-26">
	<td class="column-1">25</td><td class="column-2">0:02:36</td><td class="column-3">Task successfully completed (correct location and service)</td><td class="column-4"></td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 1: </strong>Timeline for Participant A.</p>
<h3><span lang="EN-US">ChatGPT Problem-by-Problem Results</span></h3>
<p>Appendix Table 2 shows the usability problems reported by ChatGPT and the UX researcher, run by run.</p>

<table id="tablepress-1064" class="tablepress tablepress-id-1064">
<thead>
<tr class="row-1">
	<th class="column-1"><center>Prob #</th><th class="column-2">Description</th><th class="column-3"><center>Run 1</th><th class="column-4"><center>Run 2</th><th class="column-5"><center>Run 3</th><th class="column-6"><center>Run 4</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-4">
	<td class="column-1">C1</td><td class="column-2">Page shows a mostly blank/loading state after selecting Grooming Salon (FALSE ALARM—this happened but only lasted about two seconds within which the participant started to select Breed but diverted to Dog when that button appeared)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"><center>1</td>
</tr>
<tr class="row-5">
	<td class="column-1">C2</td><td class="column-2">Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-6">
	<td class="column-1">C3</td><td class="column-2">Participant enters 80246 and gets multiple grooming locations (FALSE ALARM—true but this is how this control is supposed to work)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-7">
	<td class="column-1">H3</td><td class="column-2">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-8">
	<td class="column-1">H4</td><td class="column-2">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-9">
	<td class="column-1">C4</td><td class="column-2">After selecting breed and age, the menu reloads. (FALSE ALARM—true but happens quickly and is the expected action after clicking Check Prices &amp; Book Now—no impact on participant behavior)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-10">
	<td class="column-1">C5</td><td class="column-2">The target option is at the bottom of the list below the fold/below the Bath&amp; Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)</td><td class="column-3"></td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-11">
	<td class="column-1">C6</td><td class="column-2">Large promotional grooming content and video tiles take up much of the page while the actual service menu is constrained to the right side (FALSE ALARM—true but not an obvious problem for this participant)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 2: </strong>Usability problems reported by ChatGPT 5.4 Thinking and the UX researcher for Participant A.</p>
<h3><span lang="EN-US">Gemini Problem-by-Problem Results</span></h3>
<p>Appendix Table 3 shows the usability problems reported by Gemini and the UX researcher, run by run.</p>

<table id="tablepress-1065" class="tablepress tablepress-id-1065">
<thead>
<tr class="row-1">
	<th class="column-1"><center>Prob #</th><th class="column-2">Description</th><th class="column-3"><center>Run 1</th><th class="column-4"><center>Run 2</th><th class="column-5"><center>Run 3</th><th class="column-6"><center>Run 4</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"></td><td class="column-4"><center>1</td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-4">
	<td class="column-1">G1</td><td class="column-2">Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"></td>
</tr>
<tr class="row-5">
	<td class="column-1">H3</td><td class="column-2">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-6">
	<td class="column-1">H4</td><td class="column-2">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-7">
	<td class="column-1">G2</td><td class="column-2">The target option is at the bottom of the list below the fold/below the Bath&amp; Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 3: </strong>Usability problems reported by Gemini 3 Flash Thinking and the UX researcher for Participant A.</p>
<h2><span lang="EN-US">Appendix B: The PetSmart Reservation Task</span></h2>
<p>For a <a href="https://measuringu.com/ux-pets/">UX benchmark study conducted in 2019</a>, one of the participants’ tasks was to use the PetSmart website to start booking a grooming appointment in Glendale, CO for a one-year-old English Springer Spaniel, then stop after determining the cost of a bath and full haircut. In this section, we review the steps through the “happy path” and speculate about possible user behaviors that would be reasonable to track to provide background knowledge for understanding the problem lists presented later.</p>
<p>Appendix Figure 1 shows the home page. Before continuing, ask yourself, where would you start?</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48131" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png" alt="PetSmart home page for the pet grooming task." width="624" height="306" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png 624w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF1-300x147.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF1-600x294.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 1: </strong>Home page for the pet grooming task.</p>
<p>For this task, the best first choice is to click “pet services” from the horizontal navigation menu close to the top of the page. From here, there are two paths to grooming, shown in Appendix Figure 2. Appendix Figure 2a shows the dropdown from which a user could drag the cursor down and release the button to select Grooming Salon. Appendix Figure 2b shows the pet services menu that appears after clicking “pet services” but releasing the mouse button without dragging, from which the user would click Grooming.</p>
<p><strong>Appendix Figure 2a: Click and drag path</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48132" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png" alt="Click and drag path to grooming" width="780" height="328" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-600x252.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><strong>Appendix Figure 2b: Click without dragging path</strong><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48132" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png" alt="Click and drag path to grooming" width="780" height="328" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-600x252.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48133" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png" alt="Click without dragging path." width="780" height="308" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-300x118.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-768x303.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-600x237.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 2: </strong>Two paths to the grooming menu.</p>
<p>Appendix Figure 3 shows the grooming form. This is where users who are not in Glendale can change the location to Glendale, select dog, select the breed, select the age, then click the button to check prices.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48135" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png" alt="The grooming form." width="780" height="93" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-300x36.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-768x92.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-600x72.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 3: </strong>The grooming form.</p>
<p>Before checking prices, users needed to select a breed and age for their dog. As shown in Appendix Figure 4, clicking breed produced a searchable breed dropdown. Appendix Figure 4a shows the initial appearance of the dropdown; Appendix Figure 4b shows its appearance after typing “english&#8221; over the placeholder text &#8220;breed&#8221; in the combobox.</p>
<p><strong>Appendix Figure 4a: Initial appearance of the breed dropdown</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48136" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png" alt="Initial appearance of the breed dropdown." width="780" height="217" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-300x83.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-768x214.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-600x167.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><strong>Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48137" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png" alt="Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”" width="780" height="177" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-300x68.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-768x174.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-600x136.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 4: </strong>The searchable breed dropdown.</p>
<p>For the happy path, a user should type “english” into the combobox, so one of the potential problems we anticipated was users not realizing the dropdown list could be filtered. Compounding the complexity of this step in the process is that the location of the word “English” for various breeds was inconsistent. For example, after filtering, the list in Figure 4b included English Toy Spaniel, Old English Sheepdog, and Springer Spaniel &#8211; English. That’s less of a problem after filtering but could be more problematic if scrolling through the unfiltered list.</p>
<p>The age dropdown, shown in Appendix Figure 5, was relatively straightforward with only two choices.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48138" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png" alt="The age dropdown." width="780" height="128" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-300x49.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-768x126.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-600x98.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 5: </strong>The age dropdown.</p>
<p>With the grooming form completed, the next step is to click “check prices &amp; book now” to get the list of grooming options shown in Appendix Figure 6. Because the target option was the last one in the list and below the fold, we anticipated that some users might select an earlier option.</p>
<p><strong>Appendix Figure 6a: Completed grooming menu and first option in Grooming Salon Menu (above the fold)<a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48139 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png" alt="Completed grooming menu and first option in Grooming Salon Menu (above the fold)" width="780" height="359" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-300x138.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-768x353.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-600x276.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></strong></p>
<p><strong>Appendix Figure 6b: The other grooming options (below the fold; the last option is the target)</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48140" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png" alt="The other grooming options (below the fold, last option is the target)" width="780" height="304" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-300x117.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-768x299.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-600x234.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 6: </strong>Grooming options above and below the fold, showing the target Bath &amp; Full Haircut for $61.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Can We Trust AI to Moderate UX Interviews?</title>
		<link>https://measuringu.com/can-we-trust-ai-to-moderate-ux-interviews/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=can-we-trust-ai-to-moderate-ux-interviews</link>
		
		<dc:creator><![CDATA[Jeff Sauro, PhD • Lucas Plabst, PhD • Jim Lewis, PhD]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 23:01:15 +0000</pubDate>
				<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[mod]]></category>
		<category><![CDATA[moderated research]]></category>
		<category><![CDATA[Moderating]]></category>
		<category><![CDATA[Moderation]]></category>
		<category><![CDATA[moderator]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48092</guid>

					<description><![CDATA[Can AI Replace UX Researchers? Some criticized us for using that headline in 2023, calling it clickbait. Fair criticism, given the inflammatory nature of AI (for the record, we did have a subtitle). But here we are three years later, and we’re not talking about just the tedious task of having AI code comments—a task [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48106 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-300x169.jpg" alt="Feature image showing an AI moderator" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a><em>Can AI Replace UX Researchers?</em></p>
<p>Some criticized us for using that <a href="https://measuringu.com/classification-agreement-between-ux-researchers-and-chatgpt/">headline in 2023</a>, calling it clickbait. Fair criticism, given the inflammatory nature of AI (for the record, we did have a subtitle). But here we are three years later, and we’re not talking about just the tedious task of having AI code comments—a task that most researchers are happy to offload. We’re talking about AI being able to do some of the core jobs of the UX researcher, in particular, moderating participant sessions.</p>
<p>When we wrote our article three years ago about AI doing comment coding, the idea of replacing a UX moderator seemed about as probable as replacing a waiter in a restaurant. Both jobs seemed immune to AI. Yet here we are. AI isn’t just crunching numbers; it’s being tasked with doing qualitative work. <span data-preserver-spaces="true">It’s</span><span data-preserver-spaces="true"> a serious enough threat that </span><a class="editor-rtfLink" href="https://journals.sagepub.com/doi/full/10.1177/10778004251401851" target="_blank" rel="noopener"><span data-preserver-spaces="true">419 professionals </span></a><span data-preserver-spaces="true">have signed letters against the use of AI for qual research.</span></p>
<p>And just like we did with <a href="https://measuringu.com/review-of-experiments-with-synthetic-users/">synthetic users</a>, we need to separate the hype from the data. That means asking what the claim actually is, who it comes from, and how strong the evidence behind it is. Is it anecdotal? Is it peer-reviewed? Is it from a company selling AI moderators?</p>
<p>Before we dig into the claims about AI and moderation and start letter-writing campaigns, it’s important to understand what moderating is, what an AI moderator is, and what one can do.</p>
<h2><span lang="EN">What Is Moderating in UX Research?</span></h2>
<p>Moderating is a general term that’s not to be confused with content moderation on social networks. In UX research, <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">moderating is a broad term</a> that spans several UX research methods: <a href="https://measuringu.com/three-goals/">usability testing</a>, interviews, <a href="https://measuringu.com/contextual-inquiry/">contextual inquiry</a>, and occasionally <a href="https://measuringu.com/f-word/">focus groups</a>. The moderator&#8217;s job changes depending on the method and goals.</p>
<p>An unstructured in-depth interview where you need to uncover problems and opportunities for a new product requires a different approach than a large-scale summative evaluation where you need minimal interaction to collect metrics (yes, you can moderate and collect metrics).</p>
<h2>What Makes a Moderator Good?</h2>
<p>Beyond knowing the study type, the practical skills that separate strong moderators from weak ones are things like building enough rapport that participants give honest reactions, probing appropriately (asking a follow-up once or twice to get past a surface answer, without dragging out a dead end), catching participants who are misrepresenting themselves or reading from a script, using silence to let someone think rather than rushing to fill it, and knowing when to go off-script versus when to stick to the guide.</p>
<p>None of this is really a checklist. Instead, it&#8217;s judgment applied in real time to what a specific person just said. Even with that list, there is no agreed-upon way to score whether a moderator did a good job. What makes a good moderator isn’t even clearly defined, let alone measured objectively.</p>
<h2><span lang="EN">What Is an AI Moderator?</span></h2>
<p>An AI moderator, as the name suggests, is software that uses AI to interactively interview participants using voice and even video to replicate the interactivity of a human. But it’s more than software reading a script and asking questions. An AI can be prompted to probe and follow up based on what the participants say. The technology has advanced considerably since Dragon NaturallySpeaking in the 1990s and Google’s cutting-edge transcription <a href="https://technologizer.com/2010/08/22/worst-google-voice-transcription-errors/index.html">failures of the 2010s</a>. AI conversations now flow with high accuracy and little latency.</p>
<h2><span lang="EN">AI Moderators (Not Quite HAL 9000)</span></h2>
<p>While the technology has come a long way, AI moderators are not like the diabolical AI in <em>2001: A Space Odyssey</em>. They are limited by the script you feed them and in their interactive ability.</p>
<p>An AI moderator is software (prompts) built on top of the same LLM models that make headlines (Claude, ChatGPT, Gemini). The moderators can be voice-only or have a range of appearances, from simple visualizations to more life-like avatars.</p>
<p><a href="https://chatgpt.com/features/voice/">ChatGPT</a>, <a href="https://gemini.google/overview/gemini-live/">Gemini</a>, and <a href="https://support.claude.com/en/articles/11101966-use-voice-mode">Claude</a> all ship voice modes that hold unstructured back-and-forth conversations, interruptions included, with no command list behind them. Figure 1 shows the appearance of ChatGPT&#8217;s voice interface compared with HAL 9000 from <em>2001: A Space Odyssey</em>.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1a.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48093" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1a.png" alt="ChatGPT’s visual placeholder for its voice." width="198" height="216" /></a><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48094" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1b.png" alt="HAL 9000" width="288" height="216" /> </a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>ChatGPT’s visual placeholder for its voice (left) and HAL 9000 (right).</p>
<p>An example of a simplified prompt for the LLM could be something like the following for doing research on a new sleep-aid product:</p>
<blockquote><p>You are a UX researcher running semi-structured interviews with participants in a research study. The goal of the study is to find out more about their sleeping behavior and to find potential market opportunities for wake-up devices. Follow the discussion guide you are given directly, covering all required questions, but follow up on interesting threads.</p></blockquote>
<p>AI moderators can also have avatars to go with the voice, so respondents aren’t talking to a circle or blank screen. Figure 2 shows an example of one used for job interviews.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48095" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png" alt="Example of a visual AI moderator used for job interviews from Humanly.io." width="624" height="301" srcset="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png 624w, https://measuringu.com/wp-content/uploads/2026/08/08042026-F2-300x145.png 300w, https://measuringu.com/wp-content/uploads/2026/08/08042026-F2-600x289.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Example of a visual AI moderator used for job interviews from Humanly.io.</p>
<h2><span lang="EN">AI Moderation in Job Interviews</span></h2>
<p>We can talk about AI moderators as a concept, but are they really being used at scale? One clear application: interviewing for a job, where services like <a href="http://codesignal.com">CodeSignal</a>, <a href="http://humanly.io">Humanly</a>, or <a href="https://eightfold.ai/">Eightfold</a> offer AI interviewers.</p>
<p>One estimate suggests over <a href="https://www.greenhouse.com/newsroom/63-of-job-seekers-have-faced-an-ai-interview-most-havent-had-a-good-one-yet">60% of job seekers</a> have had exposure to an AI interview. Note that this estimate comes from a company selling AI moderating technology, so it’s not necessarily objective.</p>
<p>However, a large <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5395709">pre-registered experiment from 2025</a> with over 70,000 job applicants in the Philippines provides a more objective data point. Participants were randomly assigned to either a human interviewer (20% of sample), an AI job interviewer (60% of sample), or allowed to choose between the two (20% of sample) after a brief intro to the AI interviewer.</p>
<p>Candidates interviewed by the AI were 12% more likely to get an offer, 18% more likely to start, and 17% more likely to still be employed after a month. Of the subset that had a choice, a surprisingly high 78% picked the AI.</p>
<p>These numbers are hard to ignore. Consider a few caveats, though, before we fire all the job interviewers and UX moderators out there. First, this was for entry-level customer service jobs (call center) in the Philippines. Second, humans made the hiring decision, not AI. Probably most importantly, the value that was derived was in the perception and reality of a consistently delivered interviewer. The type of questions being asked were mostly closed, factual, verification-style questions, things like commute time, salary fit, availability, and contact info.</p>
<p>Job interviews are notorious for both explicit and implicit bias. There does seem to be a benefit if those biases can be significantly reduced (or even eliminated).</p>
<p>Interviewing a candidate for a job isn’t the same as interviewing a participant to understand usage patterns. While we are concerned about <a href="https://measuringu.com/ut-bias/">identifying and reducing bias</a> in UX research, UX moderation is often best deployed for unstructured and unknown problems.</p>
<p>Job interviews, especially screening interviews, are usually very structured with a clear set of questions. They are designed to be consistent. A structured UX interview is really more akin to a verbalized survey (something we’ll revisit). Where interviews really matter is when you don’t know what you don’t know.</p>
<h2><span lang="EN">AI in Qual Research</span></h2>
<p>Job interviewing is qualitatively different from the sort of inquiry in moderated research. In one of the best-known AI interviewer studies, in December 2025, Anthropic conducted a large study of 80k Claude users across 159 countries using their new Anthropic Interviewer tool. Their interview questions were a bit closer to the types of topics investigated in UX moderated interviews:</p>
<ol>
<li>What&#8217;s the last thing you used an AI chatbot for?</li>
<li>If you could wave a magic wand, what would AI do for you?</li>
<li>Has AI ever taken a step towards that vision for you?</li>
<li>Are there ways AI might be developed that would be contrary to your vision or what you value?</li>
</ol>
<p>Anthropic Interviewer followed up on each, probing for the underlying values and experiences behind people&#8217;s answers. Claude was then used to synthesize the transcripts for themes (much to the chagrin of the 419 professionals who rejected such usage).</p>
<p>Claude’s analysis determined that 88% to 98% of the open-ended comments were “substantive.” There aren’t many details to assess the quality relative to a human, but the authors reported that a subset of comments were validated as having at least 90% agreement with human coders on 25 labels. The authors were surprised by how candid some people were. Respondents shared things like grief, mental health crises, financial precarity, and relationship failures, all of which our human user researchers rarely encounter in traditional interviews. It’s unclear if those same comments would have been shared with a human, but it’s certainly a possible benefit worth investigating more when a topic may elicit more sensitive comments from participants.</p>
<p>What is clear is that this is a huge sample that would almost certainly never happen with human moderation. It’s less clear how well the insights from a smaller sample conducted by humans would have performed.</p>
<p>Anthropic isn’t the only one in the AI interviewer game, and others have noticed the self-disclosure. <a href="https://heymarvin.com/product/ai-moderated-interviewer">Marvin’s</a> AI Interviewer “leads human-like conversations, searches for the why behind responses, and gathers insights faster than ever before.” Or, from <a href="https://getperspective.ai/agents/interviewer">Perspective</a>, “customers share things in these conversations they&#8217;d never put in a form and would rarely say on a Zoom call with a stranger. Not a chatbot, not a survey, not a junior researcher reading from a script—the best interviewer you&#8217;ve ever seen, available at any hour, in any language, for every single customer.”</p>
<p>Bold vendor claims like these are worth investigating. <a href="https://www.nngroup.com/articles/ai-interviewers/">NN/Group’s</a> analysis of ten experienced researchers using two AI moderators (including <a href="https://heymarvin.com/product/ai-moderated-interviewer">Marvin</a> from above and <a href="https://userflix.de/">UserFlix</a>) found that AI interviewers can handle structured, scripted interviews (at scale). However, they didn’t think AI interviewers were adequate yet for semi-structured interviews.</p>
<h2>Are We Ready to Deploy the AI Moderators?</h2>
<p>So, we have some data suggesting that AI can be used at scale to interview job candidates. We have at least a proof of concept that AI can moderate a massive number of sessions for a handful of more open-ended research questions. And there’s some evidence that people may be more willing to disclose more to an AI moderator than to a human.</p>
<p>But if you recruit a dozen IT decision makers, or Chief Product Officers, and want to build them a better product by conducting a semi-structured interview that requires a moderator to go off script, would you use an AI moderator? NN/Group certainly suggests we aren’t there.</p>
<p>But this raises the question of how well an AI moderator would do compared to a human. That’s a question we’ll take up in our upcoming articles. First, we’ll review the published literature, and then we&#8217;ll report on the results of our own controlled experiment, where we compare an AI moderator with a human moderator in a UX research context.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Measuring the UX of AI</title>
		<link>https://measuringu.com/measuring-the-ux-of-ai/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=measuring-the-ux-of-ai</link>
		
		<dc:creator><![CDATA[Jeff Sauro, PhD • Jim Lewis, PhD]]></dc:creator>
		<pubDate>Tue, 28 Jul 2026 21:47:36 +0000</pubDate>
				<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48046</guid>

					<description><![CDATA[AI is everywhere and getting embedded in all of our products. If you ask the typical person in 2026 what AI is, they’ll probably say it’s a generative chat product like ChatGPT, Copilot, Claude, or Gemini. Of course, these are only the frontier Large Language Models (LLMs). That&#8217;s not all AI is. For example, the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48076 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-300x169.jpg" alt="Feature image showing a researcher measuring the UX of AI" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>AI is everywhere and getting embedded in all of our products.</p>
<p>If you ask the <a href="https://hbr.org/2026/06/how-people-are-really-using-ai-in-2026">typical person in 2026 what AI is</a>, they’ll probably say it’s a generative chat product like ChatGPT, Copilot, Claude, or Gemini. Of course, these are only the <a href="https://www.iguazio.com/glossary/frontier-model/">frontier Large Language Models</a> (LLMs). That&#8217;s not all AI is.</p>
<p>For example, the algorithm recommending your next Netflix show, the AI drafting a formula in Excel, or the model flagging fraud on your credit card are all types of AI (none of which have chat interfaces).</p>
<p>But when most people say they&#8217;re “using AI,” they mean typing into a chat box, and that&#8217;s a good place to start when thinking about how to measure the UX of AI as it’s popularly understood.</p>
<p>How should we measure the quality of these experiences?</p>
<h2><span lang="EN-US">Measuring UX in General</span></h2>
<p>Measuring the user experience in general involves assessing <a href="https://measuringu.com/ux-measurement-purpose/">what people think and feel, and what people do</a>. That means using a mix of <a href="https://measuringu.com/get-comfortable-with-four-ux-metrics/">attitudinal and action measures</a>.</p>
<p>Action (behavioral) measures are more straightforward to interpret. A typical suite of action metrics includes a combination of effectiveness (<a href="https://measuringu.com/completion-rates/">completion rates</a>, <a href="https://measuringu.com/errors-ux/">errors</a>) and efficiency (<a href="https://measuringu.com/task-times/">time on task</a>). But they’re relatively hard to collect because you have to set up <a href="https://measuringu.com/task-based-metrics/">task scenarios</a> and record or observe behaviors.</p>
<p>Attitudinal measures are easier to collect, but you need to be sure you’re measuring the right thing.</p>
<p>For attitudes, we’ve found that standardized metrics like the <a href="https://measuringu.com/how-to-score-and-interpret-the-ux-lite/">UX-Lite</a><sup>®</sup> provide a good measure of overall attitudes toward a product’s usefulness (capabilities/features) and usability (ease of use). For example, see our <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">2026 retrospective benchmarks for ChatGPT, Claude, Gemini, and Grok</a>.</p>
<p>Having an assessment of usefulness and usability from the UX-Lite provides good high-level measures that can be compared to historical benchmarks. But even though it discriminates at a high level between usefulness and usability, it doesn’t provide very granular or diagnostic measures. Consequently, we’ll also want to explore more specific attitudes to see how they affect the way people think about their use of AI chatbots.</p>
<p>When measuring specific interfaces, it can be helpful to identify additional constructs, features, or interactions that participants can rate so you have a better idea of which aspects of the UX are perceived as good or bad. That gives you a more diagnostic set of items.</p>
<p>How do you do that for AI chat interfaces like ChatGPT, Claude, Gemini, and Grok? You follow the process for creating standardized measures (for example, see our <a href="https://measuringu.com/article/measuring-the-perceived-clutter-of-websites/">IJHCI paper on measuring the perceived clutter of websites</a>). You need items, data, and validation.</p>
<h2><span lang="EN-US">Picking The Items: What Matters When Interacting with Generative AI Chat Software?</span></h2>
<p>When building a standardized measure, the first step is to pick a set of items. We developed an initial set of 34 items based on input from the MeasuringU research team, drawing upon the existing literature and their experiences using these products to measure constructs like AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. Consistent with psychometric practice, we created at least three items per construct.</p>
<h3><span lang="EN-US">AI Productivity</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48078" src="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png" alt="A researcher using an AI for productivity" width="514" height="247" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-600x288.png 600w" sizes="auto, (max-width: 514px) 100vw, 514px" /></p>
<p>One of the most attractive capabilities of generative AI chatbots is the potential for enhanced productivity, making this an important construct to measure. A survey conducted by Microsoft and LinkedIn in 2024 found that <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part">90% of respondents said using AI helped them save time</a>.</p>
<p>In 2026, results from an Anthropic survey indicated <a href="https://www.anthropic.com/research/economic-index-june-2026-report">86% of respondents reported improvements in the speed</a> of their work. Table 1 shows the initial set of items we developed for this construct of increased productivity.</p>

<table id="tablepress-1055" class="tablepress tablepress-id-1055">
<thead>
<tr class="row-1">
	<th class="column-1">AI Productivity (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Using this AI chatbot greatly improves my productivity.</td>
</tr>
<tr class="row-3">
	<td class="column-1">This AI chatbot adds substantial value to my personal tasks.</td>
</tr>
<tr class="row-4">
	<td class="column-1">This AI chatbot adds substantial value to my professional tasks.</td>
</tr>
<tr class="row-5">
	<td class="column-1">Using this AI chatbot makes me feel more capable in my work or studies.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I feel comfortable being accountable for work that used this AI chatbot.</td>
</tr>
<tr class="row-7">
	<td class="column-1">Using this AI chatbot helps me achieve my goals.</td>
</tr>
<tr class="row-8">
	<td class="column-1">The amount of time it takes for this AI chatbot to respond is acceptable.</td>
</tr>
<tr class="row-9">
	<td class="column-1">This AI chatbot’s responses efficiently tell me the information I need.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>Initial item set for AI Productivity.</p>
<h3><span lang="EN-US">AI Trust</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48080" src="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png" alt="A researcher with computer screen showing &quot;AI trust&quot;" width="510" height="245" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-600x288.png 600w" sizes="auto, (max-width: 510px) 100vw, 510px" /></p>
<p>The flip side of excitement about increased productivity is distrust in AI output and data security. Even recent models <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">hallucinate</a> usability problems after reviewing videos of usability test sessions. A 2025 Melbourne-KPMG survey of 48,000 people across 47 countries found that <a href="https://fbe.unimelb.edu.au/newsroom/media-release-global-study-reveals-trust-of-ai-remains-a-critical-challenge-reflecting-tension-between-benefits-and-risks">less than half of the people regularly using AI were willing to trust it</a>. Table 2 shows our initial set of AI Trust items.</p>

<table id="tablepress-1056" class="tablepress tablepress-id-1056">
<thead>
<tr class="row-1">
	<th class="column-1">AI Trust (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I trust this AI chatbot to provide reliable information.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel confident relying on responses from this AI chatbot when making decisions.</td>
</tr>
<tr class="row-4">
	<td class="column-1">This AI chatbot always provides accurate responses.</td>
</tr>
<tr class="row-5">
	<td class="column-1">When this AI chatbot makes mistakes, they are usually easy to detect.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I don’t worry about how my data is used when interacting with this AI chatbot.</td>
</tr>
<tr class="row-7">
	<td class="column-1">It’s easy to understand what happens to the information I share with this AI chatbot.</td>
</tr>
<tr class="row-8">
	<td class="column-1">My professional value is not affected by products like this AI chatbot.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>Initial item set for AI Trust.</p>
<h3><span lang="EN-US">AI Dependency</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48079" src="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png" alt="A relaxed researcher watching an AI assistant doing the work" width="518" height="249" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-600x288.png 600w" sizes="auto, (max-width: 518px) 100vw, 518px" /></p>
<p>As with Trust, there are particular concerns about AI users being overly dependent and uncritical of AI outputs. In the Melbourne-KPMG survey, 66% of respondents reported relying on AI output without evaluating its accuracy, and <a href="https://assets.kpmg.com/content/dam/kpmgsites/xx/pdf/2025/05/trust-attitudes-and-use-of-ai-global-report.pdf">56% reported making mistakes in their work due to uncritical acceptance of an AI output</a>. See Table 3 for the AI Dependency items.</p>

<table id="tablepress-1057" class="tablepress tablepress-id-1057">
<thead>
<tr class="row-1">
	<th class="column-1">AI Dependency (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I tend to accept answers from AI chatbots without verifying their accuracy.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I rarely double-check information provided by AI chatbots.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 3: </strong>Initial item set for AI Dependency.</p>
<h3><span lang="EN-US">AI Anxiety</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48081" src="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png" alt="Stressed researcher looking at a line chart" width="512" height="246" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-600x288.png 600w" sizes="auto, (max-width: 512px) 100vw, 512px" /></p>
<p>Conversations about AI often turn to anxiety about the potential negative effects of generative AI chatbots on society, the environment, and ethics. A 2025 University of Chicago AP-NORC survey found <a href="https://epic.uchicago.edu/wp-content/uploads/sites/5/2025/10/EPIC-AP-NORC-Poll_AI_2025_Fact-Sheet.pdf">44% believed AI would do more to hurt than help society</a>, compared with 22% who expected it to do more good, and 41% were extremely or very concerned about AI’s environmental impact. Table 4 shows our initial item set for AI Anxiety.</p>

<table id="tablepress-1058" class="tablepress tablepress-id-1058">
<thead>
<tr class="row-1">
	<th class="column-1">AI Anxiety (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">The increasing use of AI makes me uneasy.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am often concerned that AI could cause serious harm to society.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I often worry about the environmental impact of AI.</td>
</tr>
<tr class="row-5">
	<td class="column-1">AI development feels risky.</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI development feels difficult to control.</td>
</tr>
<tr class="row-7">
	<td class="column-1">There should be more government regulation for AI development.</td>
</tr>
<tr class="row-8">
	<td class="column-1">Using AI chatbots for work or school feels unethical.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 4: </strong>Initial item set for AI Anxiety.</p>
<h3><span lang="EN-US">AI Personification</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48082" src="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png" alt="A researcher interacting with an AI in the form of an angelic woman" width="507" height="244" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-600x288.png 600w" sizes="auto, (max-width: 507px) 100vw, 507px" /></p>
<p>Another aspect of interaction with generative AI chatbots that interests us is the extent to which users feel a personal relationship with the AI. This could range from feeling like you&#8217;re communicating with a human-like entity to feelings of friendship. <a href="https://www.nature.com/articles/s41598-025-19212-2">Not everyone has the same emotional reaction to generative AI chatbot products</a>, but we are interested in how products may differ in the extent to which they lead to social connection with their users. Our initial set of AI Personification items is listed in Table 5.</p>

<table id="tablepress-1059" class="tablepress tablepress-id-1059">
<thead>
<tr class="row-1">
	<th class="column-1">AI Personification (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Interacting with this AI chatbot feels like communicating with a human.</td>
</tr>
<tr class="row-3">
	<td class="column-1">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I feel like AI chatbots understand me well.</td>
</tr>
<tr class="row-5">
	<td class="column-1">I tend to feel a sense of connection when interacting with AI chatbots.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I tend to feel like I’m socializing when I interact with AI chatbots.</td>
</tr>
<tr class="row-7">
	<td class="column-1">I’m more likely to share personal information with AI chatbots than with other people.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 5: </strong>Initial item set for AI Personification.</p>
<h3><span lang="EN-US">Early Adoption</span></h3>
<p>Although not related solely to AI, we included three items to assess respondents’ tendencies to be early adopters of new technologies (Table 6).</p>

<table id="tablepress-1060" class="tablepress tablepress-id-1060">
<thead>
<tr class="row-1">
	<th class="column-1">Early Adoption (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I like to experiment with new technologies before most people do.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am usually among the first to try new digital tools.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I actively seek out new technologies to try.</td>
</tr>
</tbody>
</table>

<p class="wp-caption-text" style="text-align: left;"><strong>Table 6: </strong>Initial item set for Early Adoption.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>In this article, we discussed how to measure the UX of AI in general and, specifically, generative AI chatbots like ChatGPT, Claude, Gemini, or Grok. The key points in the article were:</p>
<p><strong>Measuring AI UX Starts with Measuring UX. </strong>Measuring UX in general means using a mix of attitudinal and action metrics. Action metrics are more rigorous but harder to collect—you need task scenarios and observation. Attitudinal metrics are easier to collect and often take the form of standardized questionnaires. Standardized metrics like the UX-Lite can discriminate among products based on their perceived usefulness and usability and are a good place to start. But if you want more diagnostic insight—to know more than just that <em>something</em> is off—you need a deeper, more specialized set of measures.</p>
<p><strong>What Constructs Comprise the UX of AI? </strong>A deeper dive into the UX of generative AI chatbots requires investigation of specialized constructs. Based on our reading and experience with these types of products, we&#8217;ve proposed items for measuring AI Productivity, AI Trust, AI Dependency, AI Anxiety, and AI Personification.</p>
<p><strong>To Validate Items, You Need Data from Real People. </strong>Creating an initial set of items is an important first step to develop standardized metrics, but it is just a first step. Items that look sensible on paper don&#8217;t always hold up once real people respond to them. In future articles, we&#8217;ll report the results of psychometric evaluation to (1) determine if the initial items, as we expect, group into statistical factors, (2) examine item quality to determine which items to retain for a final streamlined instrument, and (3) explore the connection between the new constructs and higher-level constructs like brand attitude, intention to continue use, and intention to recommend. Stay tuned!<strong><br />
</strong></p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!-- plugin=object-cache-pro client=phpredis metric#hits=13120 metric#misses=659 metric#hit-ratio=95.2 metric#bytes=2747230 metric#prefetches=0 metric#store-reads=367 metric#store-writes=380 metric#store-hits=88 metric#store-misses=639 metric#sql-queries=185 metric#ms-total=1153.31 metric#ms-cache=96.67 metric#ms-cache-avg=0.1296 metric#ms-cache-ratio=8.4 sample#redis-hits=32761290 sample#redis-misses=10097138 sample#redis-hit-ratio=76.4 sample#redis-ops-per-sec=528 sample#redis-evicted-keys=0 sample#redis-used-memory=62483920 sample#redis-used-memory-rss=63016960 sample#redis-memory-fragmentation-ratio=1.0 sample#redis-connected-clients=1 sample#redis-tracking-clients=0 sample#redis-rejected-connections=0 sample#redis-keys=798 -->
