<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>MeasuringU</title>
	<atom:link href="https://measuringu.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://measuringu.com</link>
	<description>UX Research and Software</description>
	<lastBuildDate>Wed, 02 Sep 2026 02:04:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.4</generator>

<image>
	<url>https://measuringu.com/wp-content/uploads/2020/11/site-icon.png</url>
	<title>MeasuringU</title>
	<link>https://measuringu.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Nine Design Fixes for Your UX Research Reports</title>
		<link>https://measuringu.com/nine-design-fixes-for-your-ux-research-reports/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=nine-design-fixes-for-your-ux-research-reports</link>
		
		<dc:creator><![CDATA[Fernanda Villalobos, MS&nbsp;•&nbsp;Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Wed, 02 Sep 2026 01:24:05 +0000</pubDate>
				<category><![CDATA[Usability]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Report]]></category>
		<category><![CDATA[Usability Testing]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48303</guid>

					<description><![CDATA[We’re known in the industry for metrics, for methods, and for defining UX research measurement. But while we believe in using data to drive decisions, we also know that changing minds is not something you can just enter into an Excel formula. Findings are only as effective as the way they are communicated. If the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1.jpg"><img fetchpriority="high" decoding="async" class="alignleft wp-image-48385 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-300x169.jpg" alt="Feature image showing nine design fixes for your UX research reports" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/090126-FeatureImage-1.jpg 2000w" sizes="(max-width: 300px) 100vw, 300px" /></a>We’re known in the industry for metrics, for methods, and for defining UX research measurement.</p>
<p>But while we believe in using data to drive decisions, we also know that changing minds is not something you can just enter into an Excel formula.</p>
<p>Findings are only as effective as the way they are communicated. If the data-analysis tree falls in the forest, will it really make a sound if no one is listening?</p>
<p>While the data are crucial, your manner of presentation determines whether your insights are absorbed, acted upon, or ignored.</p>
<p>Some of that happens when displaying the data in graphs, and we’ve previously discussed best practices for <a href="https://measuringu.com/graphing-displaying-data/">graphing and displaying data</a>. But we also recognize that the design of the deliverable itself, not just the graphs, makes a big difference.</p>
<p>The UX research report is one of the <a href="https://measuringu.com/what-are-ux-deliverables/">core deliverables</a> in our field. It can be from a usability test, an in-depth interview, or a survey. Will AI put an end to the much-maligned slide deck? Maybe. Certainly, AI is making it easier to create presentations. For at least the next few months, we don’t see the presentation going away, even if AI generates it.</p>
<p>And one thing that won’t change is the limited attention span of the readers. They won’t spend as much time reviewing your report as you put into building it. When time is limited and attention is divided (when isn’t it?), visual displays aid in <a href="https://journals.sagepub.com/doi/10.1080/14640747308400340">comprehension and retention</a>.</p>
<p>Here are nine ways to make your UX deliverable more digestible.</p>
		<div data-elementor-type="section" data-elementor-id="48413" class="elementor elementor-48413" data-elementor-post-type="elementor_library">
					<section class="elementor-section elementor-inner-section elementor-element elementor-element-443ec2f1 elementor-section-boxed elementor-section-height-default elementor-section-height-default" data-id="443ec2f1" data-element_type="section" data-e-type="section" data-settings="{&quot;background_background&quot;:&quot;gradient&quot;}">
						<div class="elementor-container elementor-column-gap-default">
					<div class="elementor-column elementor-col-100 elementor-inner-column elementor-element elementor-element-6bb96814" data-id="6bb96814" data-element_type="column" data-e-type="column">
			<div class="elementor-widget-wrap elementor-element-populated">
						<div class="elementor-element elementor-element-a0f71f8 elementor-widget elementor-widget-heading" data-id="a0f71f8" data-element_type="widget" data-e-type="widget" data-widget_type="heading.default">
				<div class="elementor-widget-container">
					<h4 class="elementor-heading-title elementor-size-default">Stay informed with MeasuringU.</h4>				</div>
				</div>
				<div class="elementor-element elementor-element-3c8e6208 elementor-widget elementor-widget-text-editor" data-id="3c8e6208" data-element_type="widget" data-e-type="widget" data-widget_type="text-editor.default">
				<div class="elementor-widget-container">
									<p>Get the latest research insights delivered weekly to your inbox.</p>								</div>
				</div>
				<div class="elementor-element elementor-element-609c9c32 elementor-mobile-button-align-start elementor-button-align-stretch elementor-widget elementor-widget-form" data-id="609c9c32" data-element_type="widget" data-e-type="widget" data-settings="{&quot;button_width&quot;:&quot;40&quot;,&quot;step_next_label&quot;:&quot;Next&quot;,&quot;step_previous_label&quot;:&quot;Previous&quot;,&quot;step_type&quot;:&quot;number_text&quot;,&quot;step_icon_shape&quot;:&quot;circle&quot;}" data-widget_type="form.default">
				<div class="elementor-widget-container">
							<form class="elementor-form" method="post" id="newsletter1" name="Newsletter Blog Footer 1" aria-label="Newsletter Blog Footer 1">
			<input type="hidden" name="post_id" value="48413"/>
			<input type="hidden" name="form_id" value="609c9c32"/>
			<input type="hidden" name="referer_title" value="" />

			
			<div class="elementor-form-fields-wrapper elementor-labels-">
								<div class="elementor-field-type-email elementor-field-group elementor-column elementor-field-group-email elementor-col-60 elementor-field-required">
												<label for="form-field-email" class="elementor-field-label elementor-screen-only">
								Email							</label>
														<input size="1" type="email" name="form_fields[email]" id="form-field-email" class="elementor-field elementor-size-md  elementor-field-textual" placeholder="Enter your email..." required="required">
											</div>
								<div class="elementor-field-group elementor-column elementor-field-type-submit elementor-col-40 e-form__buttons">
					<button class="elementor-button elementor-size-md" type="submit">
						<span class="elementor-button-content-wrapper">
																						<span class="elementor-button-text">Subscribe</span>
													</span>
					</button>
				</div>
			</div>
		</form>
						</div>
				</div>
					</div>
		</div>
					</div>
		</section>
				</div>
		
<h2>1. Start with a report structure.</h2>
<p>A sturdy building starts with a solid foundation. A clear structure makes the report stronger and is essential for guiding the narrative and helping readers follow the information, particularly since a UX researcher will not always be present to explain the content.</p>
<p>A common and effective layout for a UX research report includes:</p>
<ul>
<li><strong>Table of Contents:</strong> Helps with navigating the report and introduces the sections covered.</li>
<li><strong>Background:</strong> Provides necessary context, as not everyone will be familiar with the study’s goals.</li>
<li><strong>Methodology:</strong> Briefly outlines the research methods and participants involved.</li>
<li><strong>Executive Summary/Key Findings:</strong> Offers a concise summary of the main takeaways for clients with limited time.</li>
<li><strong>Detailed Findings:</strong> Presents the core data, including images, videos, participant quotes, and metrics (tables/charts) to support your key findings.</li>
<li><strong>Recommendations:</strong> Incorporated alongside detailed findings or in a dedicated section. Prioritize what needs immediate attention by utilizing urgency rankings.</li>
<li><strong>Next Steps:</strong> Highlights actionable items or future suggestions.</li>
<li><strong>Appendix:</strong> Include to keep the main report focused. Link to this section for additional supporting details, ensuring the document remains concise while still being comprehensive.</li>
</ul>
<p>This structure doesn’t work for every deliverable, but it’s a good place to start, and you can remove, add, or modify at will. You don’t need to always follow this structure or include every section. You also don’t need to create a separate slide or section for each part of the suggested structure.</p>
<h2>2. Break up text with high-resolution imagery.</h2>
<p>Visuals are processed faster than text, helping your reader grasp and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7618346/">remember</a> complex insights. This approach reduces cognitive load and keeps your report engaging, ensuring that key findings aren&#8217;t buried in dense paragraphs.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png" rel="attachment wp-att-48308"><img decoding="async" class="alignnone wp-image-48352 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png" alt="Example slides of mostly text and bullets by a high-resolution image." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/high-res-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/high-res-example-600x238.png 600w" sizes="(max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>3. Implement clear information hierarchy by varying heading sizes and incorporating two or three colors.</h2>
<p>A clear visual hierarchy guides the reader&#8217;s eye across the page to the most critical insights and data points without being overwhelming. Color helps highlight specific information, but do so sparingly, as too many colors make it harder to discern the hierarchy, diminishing their purpose.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png" rel="attachment wp-att-48307"><img decoding="async" class="alignnone wp-image-48353 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png" alt="Example slides, one lacking hierarchy one one with headings, subheadings, and color coding." width="1200" height="505" srcset="https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-1024x431.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/hierarchy-example-600x253.png 600w" sizes="(max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>4. Avoid very small text size and low contrast.</h2>
<p>Prioritizing readability ensures that your deliverable is accessible and inclusive for all. Use adequate text size and high-contrast colors to prevent eye strain and frustration, making it easier to read all the information.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png" rel="attachment wp-att-48306"><img loading="lazy" decoding="async" class="alignnone wp-image-48355 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png" alt="Example slides of low-contrast text and high-contrast text." width="1200" height="466" srcset="https://measuringu.com/wp-content/uploads/2026/08/contrast-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-300x117.png 300w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-1024x398.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-768x298.png 768w, https://measuringu.com/wp-content/uploads/2026/08/contrast-example-600x233.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>5. Maintain sufficient spacing within margins and between various components.</h2>
<p>Generous whitespace prevents visual clutter and focuses the reader&#8217;s attention on the most important content, making the overall presentation cleaner and less overwhelming.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png" rel="attachment wp-att-48309"><img loading="lazy" decoding="async" class="alignnone wp-image-48356 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png" alt="Example slides, one with no margins and one with margins." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/margins-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/margins-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>6. Left-aligned text is easier to read than centered text.</h2>
<p>Consistent left alignment creates a predictable reading pattern, reducing cognitive load and helping the reader process information more efficiently.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png" rel="attachment wp-att-48310"><img loading="lazy" decoding="async" class="alignnone wp-image-48357 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png" alt="Example slides of centered text and left-aligned text." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/aligned-text-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>7. Match charts and tables to the design of the slide deck.</h2>
<p>Consistency in visual elements like charts and tables reinforces your overall branding and ensures that your report feels like a cohesive, professional document. When data visualizations align with the aesthetic of your slide deck, they are less jarring and help the reader focus on the insights you are presenting.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png" rel="attachment wp-att-48311"><img loading="lazy" decoding="async" class="alignnone wp-image-48358 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png" alt="Example slides showing default chart styling and cohesive chart styling." width="1200" height="475" srcset="https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example.png 1200w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-300x119.png 300w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-1024x405.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-768x304.png 768w, https://measuringu.com/wp-content/uploads/2026/08/chart-styling-example-600x238.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<h2>8. Use a template.</h2>
<p>Don’t reinvent the deliverable wheel each time you present. If something already works, reuse it (see below on usability testing your template). To keep a consistent look across deliverables shared by a team, pair it with clear usage instructions and design principles. No two projects are identical, so it needs to stay flexible: fully editable and able to flex for different content lengths and project needs.</p>
<p>MeasuringU&#8217;s template includes placeholder text and components across the layouts we use most, built for each common report section. Uniform fonts, color palettes, imagery, icons, and title treatments balance consistency and variety. We hope to make it engaging without feeling repetitive. Even slides with distinct layouts still read as one cohesive presentation because they follow the same style guide.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48359 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png" alt="image side-by-side comparison of templates" width="1375" height="1359" srcset="https://measuringu.com/wp-content/uploads/2026/08/template-comparison.png 1375w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-300x297.png 300w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-1024x1012.png 1024w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-768x759.png 768w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-70x70.png 70w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-600x593.png 600w, https://measuringu.com/wp-content/uploads/2026/08/template-comparison-100x100.png 100w" sizes="auto, (max-width: 1375px) 100vw, 1375px" /></a></p>
<h2>9. Usability test your template and iterate.</h2>
<p>A UX research report is an interface. As such, it has users: the researcher building the report and that short-attention-spanned stakeholder. Test the template with the researcher and reader. Our current template at MeasuringU went through four iterations with the researcher team. All too often, new templates get the visual approval of a team (it looks great!). But can you get the charts in? Do the call-outs allow for enough room? Is the contrast with the text poor? Are those changes working for your reader, or are they adding to more missed points and rework?</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48361" src="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg" alt="image of UX researchers testing the template" width="512" height="506" srcset="https://measuringu.com/wp-content/uploads/2026/08/testingtemplate.jpg 741w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-300x297.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-70x70.jpg 70w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-600x594.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/testingtemplate-100x100.jpg 100w" sizes="auto, (max-width: 512px) 100vw, 512px" /></a></p>
<h2>Summary and Discussion</h2>
<p>Good UX research doesn&#8217;t speak for itself. The UX research report has to do that work, often without you in the room. The design recommendations here (structure, imagery, hierarchy, contrast, spacing, alignment) aren&#8217;t decoration; they&#8217;re what determines whether a stakeholder reads your key finding or skims past it. Start with the weakest aspect of your current template, fix it, then usability test it on someone who wasn&#8217;t in the room for the research. If your report needs you there to explain it, it isn&#8217;t finished yet.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Five Ways to Use Rapport When Moderating</title>
		<link>https://measuringu.com/five-ways-to-use-rapport-when-moderating/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=five-ways-to-use-rapport-when-moderating</link>
		
		<dc:creator><![CDATA[Jenna Herring, PsyD&nbsp;•&nbsp;Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 21:18:05 +0000</pubDate>
				<category><![CDATA[Methods]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Facilitation]]></category>
		<category><![CDATA[moderated research]]></category>
		<category><![CDATA[Moderating]]></category>
		<category><![CDATA[Moderation]]></category>
		<category><![CDATA[moderator]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48270</guid>

					<description><![CDATA[What makes a good moderator? There is more to it than reading a script. While AI moderation is getting a lot of headlines, humans will still be moderating sessions. Some teams (including ours) are experimenting with AI moderation for lower-stakes, higher-volume studies (verbal surveys). But when it&#8217;s hard to get a participant, and you need [&#8230;]]]></description>
										<content:encoded><![CDATA[<style>
@media (max-width: 782px) {
  .mu-nofloat-mobile img { float: none !important; display: block; margin: 0 auto 1.25em; }
}
</style>
<p><span class="mu-nofloat-mobile"><a href="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48291 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-300x169.jpg" alt="Feature image showing five ways to use rapport when moderating" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/082526-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a></span>What makes a good moderator? There is more to it than reading a script.</p>
<p>While AI moderation is getting a lot of headlines, humans will still be moderating sessions. Some teams (including ours) are experimenting with AI moderation for lower-stakes, higher-volume studies (verbal surveys). But when it&#8217;s hard to get a participant, and you need to get as much as you can out of their time and yours, the quality of a good moderator pays dividends.</p>
<p>In &#8220;<a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">What Makes a Good UX Research Moderator?</a>&#8221; we described how a good moderator knows the research questions and can differentiate between study types. A summative evaluation requires different moderation than an in-depth interview. In either context, however, a good moderator modulates their style and judiciously uses probing and time management. <strong>One of the fundamental moderator skills is establishing rapport.</strong></p>
<p>Establishing rapport with a participant helps build trust. That trust leads to more revealing insights. Research moderation is not therapy, though it has <a href="https://measuringu.com/thinking-aloud/">its roots there</a>. A participant who doesn&#8217;t trust the moderator may be more likely to say what they <em>think</em> the moderator wants to hear. In this article, we&#8217;ll dig deeper into how to establish a good rapport in your moderation.</p>
<h2>1. Prepare Before the Session, Not During</h2>
<p><a href="https://measuringu.com/facilitation-styles/">Good rapport starts in your prep</a>, not your opening line. Know who you&#8217;re talking to. You wouldn&#8217;t talk to a teenager the way you&#8217;d talk to an IT decision-maker, and you wouldn&#8217;t talk to a brand-new user the way you&#8217;d talk to someone who&#8217;s used the product for years.</p>
<p>Be familiar with the lingo your audience is likely to use. When you hit a term you don&#8217;t recognize, say so directly: “I&#8217;m not familiar with that. Can you explain it?” That&#8217;s not a weakness in the interview; it&#8217;s rapport in action: it tells the participant you&#8217;re actually listening, not just running a script.</p>
<h2>2. Break the Ice Before You Ask for Anything</h2>
<p>Once the session starts, the job is to lower the stakes before you raise a single task. Introduce yourself, tell participants the study is evaluating the product and not them, and remind them there are no wrong answers. If you work for an independent research firm rather than the study&#8217;s company itself, say so. That’s another small way to make it safe to provide critical information about a product or company.</p>
<p>When we break the ice, we like to open with something neutral: How long have you worked in this field? What does a typical day look like? This small talk is the on-ramp that gets a guarded participant comfortable talking openly once you get to the parts that matter. Keep in mind that session length and research goals change what rapport should look like; a 90-minute session may be able to absorb more tangents than a 30-minute session. Break the ice, but don&#8217;t let the ice-breaking break your schedule.</p>
<h2>3. Use a Verbal and Non-Verbal Toolkit</h2>
<p>A good moderator also uses a few additional techniques to carry rapport through the rest of the session:</p>
<ul>
<li><strong>The purposeful pause. </strong><a href="https://measuringu.com/10-golden-rules-of-facilitation/">Don&#8217;t rush to fill the silence</a>. Give participants room to keep thinking; the answer that comes after the pause is often the more honest one.</li>
<li><strong>Stay neutral, not encouraging. </strong>“Great job!” feels supportive, but it teaches the participant to perform for you instead of reacting honestly to the product. (An exception can be a participant who experiences multiple failures during a task-based study; you may choose to offer encouragement to keep them engaged.)</li>
<li><strong>Neutral, friendly expressions throughout. </strong>A face that only warms up at success reads as evaluation, which is exactly what you told them this isn&#8217;t.</li>
<li><strong>On camera, look at the lens, not the face. </strong>In a remote session, eye contact with the participant&#8217;s video feed doesn&#8217;t register as eye contact on their end. Looking at the camera does.</li>
</ul>
<h2>4. Watch for Rapport&#8217;s Failure Mode: Mutual People-Pleasing</h2>
<p>Rapport is supposed to make honesty easier, but push it too far and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8535677/">it does the opposite</a>. Managing the session means balancing participant authenticity with your own attempts to connect, which is easier said than done. <a href="https://www.amazon.com/Moderating-Usability-Tests-Interacting-Technologies/dp/0123739330"><em>Moderating Usability Tests</em></a> (Dumas) boils the guidance down to two directives: be professional, be genuine.</p>
<p>Watch for it in yourself first: You like this participant. They&#8217;re articulate, friendly, and easy to talk to. So you soften a follow-up you&#8217;d normally push on, or let a shaky answer slide because pressing feels rude. That&#8217;s not rapport anymore. That&#8217;s you people-pleasing the participant.</p>
<p>It runs the other way, too. An interviewer who&#8217;s a little too warm signals something the participant picks up on fast: <em>I like you, and I want this to go well.</em> The participant hears that as permission to perform rather than react; to keep telling you what a great job the product is doing, because that&#8217;s what a good, agreeable participant does for a researcher they like.</p>
<p>Same warmth. Same instinct. Opposite direction. Both land in the same place: a session that feels great and tells you almost nothing.</p>
<p>The fix isn&#8217;t less rapport. It&#8217;s checking, session by session, <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">which kind of warmth you&#8217;re building</a>. Is it the kind that makes honesty safer, or the kind that makes disagreement feel impolite? Stay friendly. Stay neutral about outcomes.</p>
<h2>5. Use Rapport as a Final Screening Step</h2>
<p>Suppose you’re conducting a usability study of an app for master plumbers who manage other plumbers&#8217; schedules. If you skip the warm-up and jump straight into tasks, you&#8217;re assuming the person is, in fact, a master plumber.</p>
<p>But say you notice they&#8217;re parroting your task language back at you instead of using the shop terms other master plumbers used in the same task. Say they rate everything at the top of the scale, or brush past usability issues that other participants flagged as real workflow problems. A few minutes of background questions at the start (What does a typical day look like? How long have you been in the trade?) would have surfaced that this participant went to trade school but never actually worked as a plumber. That&#8217;s not a rapport failure. That&#8217;s a screening failure that rapport-building would have caught.</p>
<h2>Summary and Discussion</h2>
<p>Using rapport is a core technique of a successful moderator. Five ways to make rapport work:</p>
<ul>
<li><strong>Prepare before the session. </strong>Know your audience and their vocabulary before you say hello.</li>
<li><strong>Break the ice deliberately. </strong>Lower the stakes with neutral small talk before you ask for anything.</li>
<li><strong>Use a verbal and non-verbal toolkit. </strong>Purposeful pauses, neutral reactions, and eye contact on camera keep the session honest.</li>
<li><strong>Watch for people-pleasing on both sides. </strong>Rapport should make disagreement easier, not harder.</li>
<li><strong>Treat rapport as a screening step. </strong>A few warm-up questions can catch a participant who never should have qualified.</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Streamlined Measurement of the UX of AI</title>
		<link>https://measuringu.com/streamlined-measurement-of-the-ux-of-ai/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=streamlined-measurement-of-the-ux-of-ai</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 22:21:12 +0000</pubDate>
				<category><![CDATA[Metrics]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48188</guid>

					<description><![CDATA[There’s nothing quite the same as chatting with AI. But it&#8217;s just like every other interface in one important way: we experience it. And what we experience, we can measure. Measuring the user experience means assessing both what people do and what people think (a mix of action and attitudinal metrics). For attitudes, standardized questionnaires [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48262 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-300x169.jpg" alt="" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/081926-FeatureImage-2.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>There’s nothing quite the same as chatting with AI. But it&#8217;s just like every other interface in one important way: we experience it. And what we experience, we can measure.</p>
<p>Measuring the user experience means assessing both what people do and what people think (a mix of action and <a href="https://measuringu.com/ux-attitudes/">attitudinal</a> metrics). For attitudes, standardized questionnaires like the UX-Lite<sup>®</sup> are a good place to start, but they&#8217;re not diagnostic on their own and won&#8217;t tell you why Trust is low or what&#8217;s driving Anxiety.</p>
<p>Building a standardized questionnaire involves assessing <a href="https://www.researchgate.net/publication/200085994_IBM_Computer_Usability_Satisfaction_Questionnaires_Psychometric_Evaluation_and_Instructions_for_Use">validity, reliability, and sensitivity</a>. Practically, this involves six broad steps:</p>
<ol>
<li>Identifying the items to ask participants</li>
<li>Collecting data on the candidate items</li>
<li>Checking that the items group the way you expect</li>
<li>Selecting the best items</li>
<li>Confirming a shorter version is still reliable</li>
<li>Demonstrating the items can tell good experiences from bad ones</li>
</ol>
<p>We did step one in our <a href="https://measuringu.com/measuring-the-ux-of-ai/">previous article</a>, where we identified 34 candidate items across six constructs (AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption) based on our reading and experience with these products.</p>
<p>In this article, we&#8217;ll cover the next five steps: collecting data from real users of real products and seeing how well those 34 items perform.</p>
<h2><span lang="EN-US">AI-Based Chat Software Benchmark Study</span></h2>
<p>In May 2026, we conducted a <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">retrospective study of four AI-based chat software products</a> with 420 U.S.-based panel participants. This study included the metrics we typically collect in our <a href="https://measuringu.com/consumer-software-ux-2025/">standard UX and NPS study of consumer software</a>.</p>
<p>There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several existing questionnaires, including the <a href="https://measuringu.com/nps-ux/">NPS</a>, <a href="https://measuringu.com/10-things-sus/">SUS</a>, <a href="https://measuringu.com/from-umux-lite-to-ux-lite/">UX-Lite</a>, and <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" />. In addition to these existing questionnaires, respondents completed 34 new items focused on different aspects of the UX of AI. The AI-based chat products and sample sizes were:</p>
<ul>
<li>ChatGPT: 113</li>
<li>Claude: 103</li>
<li>Gemini: 101</li>
<li>Grok: 103</li>
</ul>
<p>The sample sizes are modest but adequate to <a href="https://measuringu.com/might-not-be-a-magic-number-but-there-are-magic-ranges/">establish baselines and identify medium-sized differences </a>relative to each other and other software products we measure (e.g., ± 5 for 0–100-point rating scales). With a combined <em>n</em> = 420, the sample size is also large enough to support <a href="https://measuringu.com/advanced-stats/">advanced analyses</a> (e.g., factor analysis, regression analysis, reliability analysis, ANOVA).</p>
<h2><span lang="EN-US">Factor Analysis: Items (Almost) Perfectly Aligned with Target Constructs</span></h2>
<p>We used the multivariate technique of factor analysis to determine how well the candidate items map to their intended construct. The output of a factor analysis is <em>factor loadings</em>, numbers that range from −1 to +1, which are interpreted like a correlation. To review the factor loadings for all 34 items, see Appendix Figure 1.</p>
<p>All but one of the items (“I often rely on AI chatbots to perform tasks that I would otherwise do myself”) didn’t map well to its intended construct of Trust (it loaded as high on Productivity), so this was a good candidate to exclude. The remaining 33 items strongly loaded on their intended constructs—<strong>solid evidence of construct validity</strong>.</p>
<h2><span lang="EN-US">Item Analyses: Selecting the Best Items for a Streamlined Questionnaire</span></h2>
<p>A common approach to item selection in the psychometric process of developing a standardized questionnaire is to retain the items with the highest loadings on their associated factors. In our current practice, we enhance that method by examining item means and beta weights from regression models with a key outcome variable.</p>
<p>While still paying attention to the magnitude of item loadings, we also try to select items that vary in their observed means to better differentiate low and high levels of the construct. We also created six regression models, one for each construct, to see by examining beta weights which items accounted for larger amounts of variation in a key outcome metric: the likelihood to continue using the product.</p>
<p>In the following sections, we present the three criteria for all the items for each target construct to guide the items we selected for inclusion in a final, streamlined questionnaire. For easier interpretation, the means in the tables were <a href="https://measuringu.com/converting-scales-to-100-points/">converted from their original five-point scale to a scale ranging from 0 to 100</a>. Our typical target for measuring a construct at this stage of questionnaire development is to select two to three items per construct with the goal of achieving scale reliability (measured with coefficient alpha) of at least 0.70 for each construct.</p>
<p>The selected items are at the top of each table. Table cells are highlighted for the three largest loadings, the lowest and highest means, and the three largest beta weights. We present all the items and their scores, as you may choose to try out alternative combinations of items for your own questionnaire (<a href="https://measuringu.com/contact/">call us</a> if you need to talk it through!).</p>
<h3><span lang="EN-US">AI Productivity</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48078" src="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png" alt="Image showing man working at computer" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>For AI Productivity, we selected two of the highest loading items, one of which also had the highest beta weight. The third selected item balanced an acceptably high loading and beta weight plus a relatively low mean.</p>

<table id="tablepress-1066" class="tablepress tablepress-id-1066">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Using this AI chatbot greatly improves my productivity.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.959</p></td><td class="column-4"><center>69.2</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .259</p></td>
</tr>
<tr class="row-3">
	<td class="column-1">Using this AI chatbot makes me feel more capable in my work or studies.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.904</p></td><td class="column-4"><center>65.1</td><td class="column-5"><center>−.091</td>
</tr>
<tr class="row-4">
	<td class="column-1">I feel comfortable being accountable for work that used this AI chatbot.</td><td class="column-2"><center>x</td><td class="column-3"><center>.568</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">63.3</p></td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .183</p></td>
</tr>
<tr class="row-5">
	<td class="column-1">This AI chatbot adds substantial value to my personal tasks.</td><td class="column-2"></td><td class="column-3"><center>.701</td><td class="column-4"><center>64.9</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .245</p></td>
</tr>
<tr class="row-6">
	<td class="column-1">This AI chatbot adds substantial value to my professional tasks.</td><td class="column-2"></td><td class="column-3"><center>.884</td><td class="column-4"><center>63.7</td><td class="column-5"><center> .041</td>
</tr>
<tr class="row-7">
	<td class="column-1">Using this AI chatbot helps me achieve my goals.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.959</p></td><td class="column-4"><center>67.5</td><td class="column-5"><center> .079</td>
</tr>
<tr class="row-8">
	<td class="column-1">The amount of time it takes for this AI chatbot to respond is acceptable.</td><td class="column-2"></td><td class="column-3"><center>.524</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">78.5</p></td><td class="column-5"><center> .002</td>
</tr>
<tr class="row-9">
	<td class="column-1">This AI chatbot’s responses efficiently tell me the information I need.</td><td class="column-2"></td><td class="column-3"><center>.618</td><td class="column-4"><center>72.8</td><td class="column-5"><center> .144</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1066 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>AI Productivity items.</p>
<h3><span lang="EN-US">AI Trust</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48080" src="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png" alt="Image showing trust in AI" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items selected for AI Trust all had acceptably high loadings. The top two had impressive beta weights but little difference in their means (64.6, 62.7), so the third item was included to extend the lower range of the item means to 49.4.</p>

<table id="tablepress-1067" class="tablepress tablepress-id-1067">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I trust this AI chatbot to provide reliable information.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.654</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">64.6</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .344</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel confident relying on responses from this AI chatbot when making decisions.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.606</td><td class="column-4"><center>62.7</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .295</td>
</tr>
<tr class="row-4">
	<td class="column-1">It’s easy to understand what happens to the information I share with this AI chatbot.</td><td class="column-2"><center>x</td><td class="column-3"><center>.570</td><td class="column-4"><center>49.4</td><td class="column-5"><center> .008</td>
</tr>
<tr class="row-5">
	<td class="column-1">This AI chatbot always provides accurate responses.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.813</td><td class="column-4"><center>56.4</td><td class="column-5"><center> .003</td>
</tr>
<tr class="row-6">
	<td class="column-1">When this AI chatbot makes mistakes, they are usually easy to detect.</td><td class="column-2"></td><td class="column-3"><center>.417</td><td class="column-4"><center>58.0</td><td class="column-5"><center>−.019</td>
</tr>
<tr class="row-7">
	<td class="column-1">I don’t worry about how my data is used when interacting with this AI chatbot.</td><td class="column-2"></td><td class="column-3"><center>.391</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">44.8</td><td class="column-5"><center> .011</td>
</tr>
<tr class="row-8">
	<td class="column-1">My professional value is not affected by products like this AI chatbot.</td><td class="column-2"></td><td class="column-3"><center>.372</td><td class="column-4"><center>61.1</td><td class="column-5"><center> .076</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1067 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>AI Trust items.</p>
<h3><span lang="EN-US">AI Dependency</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48079" src="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png" alt="Image showing a man lounging while an AI does the work" width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>There were only three items developed for AI Dependency, and we excluded one of them, “I often rely on AI chatbots to perform tasks that I would otherwise do myself,” because it loaded on both Trust and Productivity. There were no issues warranting exclusion of the other two items, so we kept them both.</p>

<table id="tablepress-1068" class="tablepress tablepress-id-1068">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I tend to accept answers from AI chatbots without verifying their accuracy.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.983</td><td class="column-4"><center>37.3</td><td class="column-5"><center>.034</td>
</tr>
<tr class="row-3">
	<td class="column-1">I rarely double-check information provided by AI chatbots.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.852</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">35.1</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">.108</td>
</tr>
<tr class="row-4">
	<td class="column-1">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td><td class="column-2"></td><td class="column-3"><center>.363</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">49.2</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">.250</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1068 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 3: </strong>AI Dependency items.</p>
<h3><span lang="EN-US">AI Anxiety</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48081" src="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png" alt="Image showing nervous man watching computer." width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items retained for AI Anxiety had the highest loadings of the set, reasonably impactful beta weights, and a reasonable range of means.</p>

<table id="tablepress-1069" class="tablepress tablepress-id-1069">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">The increasing use of AI makes me uneasy.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.822</td><td class="column-4"><center>55.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.259</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am often concerned that AI could cause serious harm to society.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.899</td><td class="column-4"><center>58.9</td><td class="column-5"><center>−.095</td>
</tr>
<tr class="row-4">
	<td class="column-1">AI development feels difficult to control.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.836</td><td class="column-4"><center>60.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .126</td>
</tr>
<tr class="row-5">
	<td class="column-1">I often worry about the environmental impact of AI.</td><td class="column-2"></td><td class="column-3"><center>.780</td><td class="column-4"><center>61.8</td><td class="column-5"><center>−.044</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI development feels risky.</td><td class="column-2"></td><td class="column-3"><center>.821</td><td class="column-4"><center>57.0</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.192</td>
</tr>
<tr class="row-7">
	<td class="column-1">There should be more government regulation for AI development.</td><td class="column-2"></td><td class="column-3"><center>.794</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">66.5</td><td class="column-5"><center> .009</td>
</tr>
<tr class="row-8">
	<td class="column-1">Using AI chatbots for work or school feels unethical.</td><td class="column-2"></td><td class="column-3"><center>.603</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">47.1</td><td class="column-5"><center>−.021</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1069 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 4: </strong>AI Anxiety items.</p>
<h3><span lang="EN-US">AI Personification</span></h3>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48082" src="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png" alt="Image showing man speaking to an angelic AI incarnation." width="2560" height="1230" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-600x288.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p>The three items selected for AI Personification had acceptably high loadings, significant beta weights, and a reasonable range of means.</p>

<table id="tablepress-1070" class="tablepress tablepress-id-1070">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Interacting with this AI chatbot feels like communicating with a human.</td><td class="column-2"><center>x</td><td class="column-3"><center>.593</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">45.1</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .249</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel like AI chatbots understand me well.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.777</td><td class="column-4"><center>43.5</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .162</td>
</tr>
<tr class="row-4">
	<td class="column-1">I tend to feel a sense of connection when interacting with AI chatbots.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.985</td><td class="column-4"><center>35.5</td><td class="column-5"><center> .145</td>
</tr>
<tr class="row-5">
	<td class="column-1">I tend to feel like I’m socializing when I interact with AI chatbots.</td><td class="column-2"></td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">1.050</td><td class="column-4"><center>34.5</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;">−.299</td>
</tr>
<tr class="row-6">
	<td class="column-1">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td><td class="column-2"></td><td class="column-3"><center>.712</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">33.9</td><td class="column-5"><center> .151</td>
</tr>
<tr class="row-7">
	<td class="column-1">I’m more likely to share personal information with AI chatbots than with other people.</td><td class="column-2"></td><td class="column-3"><center>.684</td><td class="column-4"><center>36.6</td><td class="column-5"><center> .081</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1070 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 5: </strong>AI Personification items.</p>
<h3><span lang="EN-US">Early Adoption</span></h3>
<p>We did not see any issues that warranted excluding any of these items, so we kept all three.</p>

<table id="tablepress-1071" class="tablepress tablepress-id-1071">
<thead>
<tr class="row-1">
	<th class="column-1">Item</th><th class="column-2"><center>Selected</th><th class="column-3"><center>Loading</th><th class="column-4"><center>Mean</th><th class="column-5"><center>Beta</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I like to experiment with new technologies before most people do.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.961</td><td class="column-4"><center>59.9</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .118</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am usually among the first to try new digital tools.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.942</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">55.4</td><td class="column-5"><center><p style="margin-bottom:0; background-color: lightpink;"> .171</td>
</tr>
<tr class="row-4">
	<td class="column-1">I actively seek out new technologies to try.</td><td class="column-2"><center>x</td><td class="column-3"><center><p style="margin-bottom:0; background-color: lightgreen;">.920</td><td class="column-4"><center><p style="margin-bottom:0; background-color: lightblue;">62.4</td><td class="column-5"><center>−.087</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1071 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 6: </strong>Early Adoption items.</p>
<h2><span lang="EN-US">Reliability Analysis: All Streamlined Scale Reliabilities Exceeded 0.80</span></h2>
<p>Table 7 shows the coefficient alpha values for each measure for all items and for the streamlined item set. For research, the typical reliability goal is to exceed 0.70. There was little reduction in reliability for the streamlined versions and for AI Dependency; eliminating its one problematic item increased its reliability even though only two items were retained. The reliabilities for all the streamlined versions of the questionnaires not only met the typical research goal but exceeded 0.80—<strong>strong evidence of reliability for these new scales and statistical justification for the selection of their constituent items</strong>.</p>

<table id="tablepress-1072" class="tablepress tablepress-id-1072">
<thead>
<tr class="row-1">
	<th class="column-1">Reliability (Coefficient Alpha)</th><th class="column-2"><center>All Items</th><th class="column-3"><center>Streamlined</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">AI Productivity</td><td class="column-2"><center>0.92</td><td class="column-3"><center>0.85</td>
</tr>
<tr class="row-3">
	<td class="column-1">AI Trust</td><td class="column-2"><center>0.82</td><td class="column-3"><center>0.81</td>
</tr>
<tr class="row-4">
	<td class="column-1">AI Dependency</td><td class="column-2"><center>0.76</td><td class="column-3"><center>0.84</td>
</tr>
<tr class="row-5">
	<td class="column-1">AI Anxiety</td><td class="column-2"><center>0.91</td><td class="column-3"><center>0.87</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI Personification</td><td class="column-2"><center>0.92</td><td class="column-3"><center>0.86</td>
</tr>
<tr class="row-7">
	<td class="column-1">Early Adoption</td><td class="column-2"><center>0.94</td><td class="column-3"><center>NA</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1072 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 7: </strong>Scale reliabilities (coefficient alpha) for the six new metrics.</p>
<h2><span lang="EN-US">Profile Analysis: Claude Leads in Productivity, ChatGPT Lags in Trust</span></h2>
<p>We created two different visualizations of profiles for the streamlined scales, showing how the AI assistants compare. The line graph in Figure 1 makes it easy to see at a glance which scales differentiate among the products, while the column chart in Figure 2 makes it easy to compare the confidence intervals around the means.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48264 size-large" src="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1024x361.jpg" alt="Line graph of AI scale scores for four generative AI chatbots." width="1024" height="361" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1024x361.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-300x106.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-768x271.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-1536x541.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-2048x721.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081926-F1-600x211.jpg 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Line graph of AI scale scores for four generative AI chatbots.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48275 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg" alt="Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals." width="1525" height="590" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2.jpg 1525w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-300x116.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-1024x396.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-768x297.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081926-F2v2-scaled-2-600x232.jpg 600w" sizes="auto, (max-width: 1525px) 100vw, 1525px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals.</p>
<p>The results show Claude leading in AI Productivity, ChatGPT lagging in AI Trust, and little difference among the products for AI Dependency. Grok scored the lowest in AI Anxiety and the highest in AI Personification, and all four products scored different levels of Early Adoption (highest for Grok, lowest for ChatGPT).</p>
<p>A mixed ANOVA of the ratings indicated a significant main effect of scale (<em>F</em>(5, 2080) = 85.4, <em>p</em> &lt; .0001), a significant main effect of product (<em>F</em>(3, 416) = 2.8, <em>p</em> = .039), and most importantly, a highly significant scale by product interaction (<em>F</em>(15, 2080) = 3.6, <em>p </em>&lt; .0001). The interaction is important because it is <strong>strong statistical evidence of the sensitivity of these new scales</strong>.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>We collected data from 420 respondents for 34 items designed to measure six constructs related to attitudes toward four generative AI chatbots (ChatGPT, Claude, Gemini, and Grok). We then conducted analyses to complete the psychometric measurement goals of construct validity (factor analysis), measurement efficiency (item selection), scale reliability (coefficient alpha), and scale sensitivity (ANOVA).</p>
<p>The key points are:</p>
<p><strong>We have solid evidence of construct validity for the new questionnaires. </strong>Our factor analysis of the data demonstrated almost perfect alignment of items with their intended constructs for the measurement of AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. One item loaded on two factors and was thus not retained in the streamlined versions of the questionnaires.</p>
<p><strong>The items selected for streamlined versions of the questionnaires produced reliable measurement. </strong>For each of the six constructs, we retained two to three items to balance item loadings, item mean ranges, and strong beta weights for their relationships with the likelihood to continue using the product. For all the streamlined versions of the questionnaires, coefficient alpha exceeded 0.80 (ranging from 0.81 to 0.87). The common criterion for acceptable reliability is &gt; 0.70.</p>
<p><strong>The statistical evidence for scale sensitivity is strong. </strong>Profile analysis of the mean ratings of the new scales by product indicated a significant main effect of scale, a significant main effect of product, and a highly significant product-by-scale interaction. Claude led in AI Productivity, and ChatGPT lagged in AI Trust. There was little difference among the products for AI Dependency, Grok scored the lowest for AI Anxiety and the highest for AI Personification, and there were different levels of Early Adoption for all four products.</p>
<h2><span lang="EN-US">Appendix: Factor Structure and Item Key</span></h2>
<p>Appendix Figure 1 shows the pattern matrix from the factor analysis for the original 34 items. The numbers in the figure are item loadings, which indicate the degree of connection of the item with the target constructs (with values ranging from -1 to +1, interpreted like a correlation). The usual criterion for a meaningfully large loading is anything more extreme than ±0.3. A maximum likelihood factor analysis with Promax rotation using SPSS 23 was used following a <a href="https://en.wikipedia.org/wiki/Parallel_analysis">parallel analysis</a> that indicated, as expected, retention of six factors. A key pattern to look for when identifying problematic items is any item with strong loading on more than one factor. This happened to OftenRelyOnChatbots (“I often rely on AI chatbots to perform tasks that I would otherwise do myself,” highlighted in yellow), leading to its exclusion from the final questionnaires.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48206" src="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png" alt="Alignment of items with constructs (values greater than 0.3 are highlighted in green). To see the complete item text, refer to the appendix. " width="624" height="533" srcset="https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1.png 624w, https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1-300x256.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081926-AF1-1-600x513.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 1: </strong>Alignment of items with constructs (values greater than 0.3 are highlighted in green). <strong>To see the complete item text, refer to Appendix Table 1. </strong></p>
<p>Appendix Table 1 documents the short labels for each item. Items highlighted in green are the ones retained for the streamlined versions of these questionnaires.</p>

<table id="tablepress-1073" class="tablepress tablepress-id-1073 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<th class="column-1">Short Label</th><th class="column-2">Item</th>
</tr>
</thead>
<tbody class="row-hover">
<tr class="row-2">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI PRODUCTIVITY</td>
</tr>
<tr class="row-3">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">ImprovedProductivity</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Using this AI chatbot greatly improves my productivity.</td>
</tr>
<tr class="row-4">
	<td class="column-1">AddsValuePersonal</td><td class="column-2">This AI chatbot adds substantial value to my personal tasks.</td>
</tr>
<tr class="row-5">
	<td class="column-1">AddsValueProfessional</td><td class="column-2">This AI chatbot adds substantial value to my professional tasks.</td>
</tr>
<tr class="row-6">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIMakesMeFeelMoreCapable</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Using this AI chatbot makes me feel more capable in my work or studies.</td>
</tr>
<tr class="row-7">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">IFeelAccountableForMyAIAssistedWork</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel comfortable being accountable for work that used this AI chatbot.</td>
</tr>
<tr class="row-8">
	<td class="column-1">AIHelpsMeAchieveMyGoals</td><td class="column-2">Using this AI chatbot helps me achieve my goals.</td>
</tr>
<tr class="row-9">
	<td class="column-1">AIResponseTimeIsAcceptable</td><td class="column-2">The amount of time it takes for this AI chatbot to respond is acceptable.</td>
</tr>
<tr class="row-10">
	<td class="column-1">AIResponsesAreEfficient</td><td class="column-2">This AI chatbot’s responses efficiently tell me the information I need.</td>
</tr>
<tr class="row-11">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI TRUST</td>
</tr>
<tr class="row-12">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">TrustReliableInfo</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I trust this AI chatbot to provide reliable information.</td>
</tr>
<tr class="row-13">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SupportsConfidentDecisions</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel confident relying on responses from this AI chatbot when making decisions.</td>
</tr>
<tr class="row-14">
	<td class="column-1">AlwaysAccurate</td><td class="column-2">This AI chatbot always provides accurate responses.</td>
</tr>
<tr class="row-15">
	<td class="column-1">EasyToDetectMistakes</td><td class="column-2">When this AI chatbot makes mistakes, they are usually easy to detect.</td>
</tr>
<tr class="row-16">
	<td class="column-1">NotWorriedAboutDataUse</td><td class="column-2">I don’t worry about how my data is used when interacting with this AI chatbot.</td>
</tr>
<tr class="row-17">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">EasyToUnderstandHowInfoIsShared</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">It’s easy to understand what happens to the information I share with this AI chatbot.</td>
</tr>
<tr class="row-18">
	<td class="column-1">ProfValueNotAffected</td><td class="column-2">My professional value is not affected by products like this AI chatbot.</td>
</tr>
<tr class="row-19">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI DEPENDENCE</td>
</tr>
<tr class="row-20">
	<td class="column-1">OftenRelyOnChatbots</td><td class="column-2">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td>
</tr>
<tr class="row-21">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AcceptAnswersWithoutVerification</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I tend to accept answers from AI chatbots without verifying their accuracy.</td>
</tr>
<tr class="row-22">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">RarelyDoubleCheck</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I rarely double-check information provided by AI chatbots.</td>
</tr>
<tr class="row-23">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI ANXIETY</td>
</tr>
<tr class="row-24">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">UneasyWithIncreasingUse</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">The increasing use of AI makes me uneasy.</td>
</tr>
<tr class="row-25">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SeriousSocietalHarm</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I am often concerned that AI could cause serious harm to society.</td>
</tr>
<tr class="row-26">
	<td class="column-1">EnvironmentalImpact</td><td class="column-2">I often worry about the environmental impact of AI.</td>
</tr>
<tr class="row-27">
	<td class="column-1">AIDevelopmentRisky</td><td class="column-2">AI development feels risky.</td>
</tr>
<tr class="row-28">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIDevelopmentHardToControl</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">AI development feels difficult to control.</td>
</tr>
<tr class="row-29">
	<td class="column-1">NeedMoreGovRegulation</td><td class="column-2">There should be more government regulation for AI development.</td>
</tr>
<tr class="row-30">
	<td class="column-1">UseForWorkOrSchoolUnethical</td><td class="column-2">Using AI chatbots for work or school feels unethical.</td>
</tr>
<tr class="row-31">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">AI PERSONIFICATION</td>
</tr>
<tr class="row-32">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">AIWorkFeelsLikeHumanCommunication</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">Interacting with this AI chatbot feels like communicating with a human.</td>
</tr>
<tr class="row-33">
	<td class="column-1">SometimesAIFeelsLikeAFriend</td><td class="column-2">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td>
</tr>
<tr class="row-34">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">ChatbotsUnderstandMeWell</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I feel like AI chatbots understand me well.</td>
</tr>
<tr class="row-35">
	<td class="column-1">FeelSenseOfConnection</td><td class="column-2">I tend to feel a sense of connection when interacting with AI chatbots.</td>
</tr>
<tr class="row-36">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">FeelsLikeSocializing</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I tend to feel like I’m socializing when I interact with AI chatbots.</td>
</tr>
<tr class="row-37">
	<td class="column-1">MoreLikelyToSharePersonalInfo</td><td class="column-2">I’m more likely to share personal information with AI chatbots than with other people.</td>
</tr>
<tr class="row-38">
	<td colspan="2" class="column-1"><p style="margin-bottom:0; background-color: lightgrey;">EARLY ADOPTION</td>
</tr>
<tr class="row-39">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">IExperimentBeforeOthers</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I like to experiment with new technologies before most people do.</td>
</tr>
<tr class="row-40">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">FirstToTryNewDigitalTools</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I am usually among the first to try new digital tools.</td>
</tr>
<tr class="row-41">
	<td class="column-1"><p style="margin-bottom:0; background-color: lightgreen;">SeekNewTechnologiesToTry</td><td class="column-2"><p style="margin-bottom:0; background-color: lightgreen;">I actively seek out new technologies to try.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1073 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 1: </strong>Short labels and full text for each item.</p>
<p>&nbsp;</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Using AI To Find Usability Problems: A Replication</title>
		<link>https://measuringu.com/does-ai-find-real-usability-problems-a-replication/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=does-ai-find-real-usability-problems-a-replication</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 22:36:28 +0000</pubDate>
				<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Problem Discovery]]></category>
		<category><![CDATA[Usability Problem]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48115</guid>

					<description><![CDATA[If AI can code problems and potentially moderate a structured interview, can it also reliably find usability issues from watching a video? We can talk about hypotheticals, or we can actually try it out. We tried it out. In our first study, we started small. We compared human and AI reviews of a video taken [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48154 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-300x169.jpg" alt="Feature image showing AI finding usability problems" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081126-FeatureImage-1-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>If AI can <a href="https://measuringu.com/classification-agreement-between-ux-researchers-and-chatgpt/">code problems</a> and potentially <a href="https://measuringu.com/can-we-trust-ai-to-moderate-ux-interviews/">moderate a structured interview</a>, can it also reliably find usability issues from watching a video?</p>
<p>We can talk about hypotheticals, or we can actually try it out. We tried it out.</p>
<p>In our <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">first study</a>, we started small. We compared human and AI reviews of a video taken from a usability test on making dinner reservations with OpenTable. We used the standard ChatGPT and Gemini interfaces (standard Claude doesn’t support video review).</p>
<p><strong>The good:</strong> AI identified roughly <strong>half</strong> the usability problems identified by human UX researchers. It even uncovered one genuine problem not found by any of the four human evaluators.</p>
<p><strong>The bad:</strong> AI identified seven false alarms. They were true statements but not connected to the participants’ actions or utterances.</p>
<p><strong>The ugly:</strong> Three problems were straight-up hallucinations. AI described events that just didn’t happen.</p>
<p>So, of the eleven problems the two AIs reported that no human flagged, only one was a genuine problem. That&#8217;s a useful number to keep in mind: 91% of the AI-only problems in this study required either correction or dismissal.</p>
<p>A case study like that, however, is limited. There was one video, two AIs (ChatGPT, Gemini), and one prompt. Replication helps with generalization.</p>
<p>To move beyond that first case study, we kept all experimental variables the same and conducted the same type of analysis on a similar video from a different usability study.</p>
<h2><span lang="EN-US">Experimental Design: One Human Researcher, Two AIs, and One Video</span></h2>
<p>For this research, one UX researcher with decades of experience (Jim, the lead researcher from the previous study) reviewed a similar video from a previous usability benchmark study of online pet websites multiple times, creating a timeline of key events and a list of observed usability problems for this participant (referred to as Participant A; see Appendix A for the details of the timeline).</p>
<p>The characteristics the new video shared with the first one were:</p>
<ul>
<li>A similar task (steps toward booking a specified appointment)</li>
<li>A similar outcome (the user experienced several usability problems but was ultimately successful).</li>
</ul>
<p>Using the same prompt each time, we then ran the video four times each through the same two AIs used in the previous case study (ChatGPT-5.4 Thinking and Gemini 3 Flash Thinking).</p>
<p>So, in this study, we held constant the videos, the key elements of the prompt, and the AI versions/settings—variables that we eventually plan to vary. This time, as in the first case study, we only varied the type of analyst: human, ChatGPT, and Gemini.</p>
<h3><span lang="EN-US">The Task</span></h3>
<p>During the previous usability benchmark study conducted at MeasuringU in 2019, participants used the PetSmart website to start the process of booking a grooming appointment in Glendale, CO, for a bath and full haircut for an English Springer Spaniel older than six months. The task was successfully completed if the participant found that specific grooming option and reported the listed price of $61. For the step-by-step details of the &#8220;happy path&#8221; to complete this task, see Appendix B.</p>
<h3><span lang="EN-US">The Prompt</span></h3>
<p>The prompt we used for this study was:</p>
<blockquote><p><em>During a usability test, the facilitator must keep track of participant behaviors as they navigate through tasks on a website, mobile app, software program, etc. We’d like you to watch a video of a usability test where participants were asked to book a grooming reservation for their dog. As you&#8217;re watching, please look for problems the participant has while attempting to complete the task. For example, you can document the path users take, describe issues they encounter as well as what on the website might be causing problems. The task has been successfully completed if the participant finds the target service (“Bath &amp; Full Haircut” which costs $61; not “Bath &amp; Full Haircut with FURminator” which costs $74). If you understand these instructions, let me know and I&#8217;ll drag the video in for you to review. Are you ready for the video?</em></p></blockquote>
<p>This was based on the prompt we used for the OpenTable case study with slight modifications. The task details are necessarily different because there was only one correct choice for this new task (compared to many correct choices for the OpenTable task), so we specified the end task details required for successful completion.</p>
<h2><span lang="EN-US">Major Findings</span></h2>
<p>Even though the participant successfully completed the task in the video, a total of ten usability problems were identified by the AIs and the human researcher.</p>
<p>Usability problems aren’t like observing a visual defect in a product. They require judgement, so it’s worth digging into what these problems are because they&#8217;re at the crux of how AI may or may not be able to effectively emulate UX researcher judgement (for now). So, let’s dig into what we saw.</p>
<p>Participant A started with a few clicks not on the happy path, but on the third click got the task started, so these two early clicks could be considered minor usability problems that were quickly corrected.</p>
<p>There were two more impactful usability problems, both associated with breed selection (Figure 1). First, there was no visible indication that the dropdown list could be filtered by typing over the word &#8220;breed&#8221;; the participant scrolled through the entire list. Second, the order of presentation of the breeds in the dropdown list was alphabetically inconsistent (e.g., &#8220;English Toy Spaniel&#8221; started with E; &#8220;Springer Spaniel &#8211; English&#8221; started with S). Despite this, the participant successfully completed the task, just not on the most efficient path.</p>
<p><strong>a:</strong> Breed list at boundary of D and E—&#8221;English Toy Spaniel&#8221; but no English Springer Spaniel</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48116" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png" alt="Image showing breed list with English Toy Spaniel but not English Springer Spaniel" width="507" height="342" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F1a.png 507w, https://measuringu.com/wp-content/uploads/2026/08/081126-F1a-300x202.png 300w" sizes="auto, (max-width: 507px) 100vw, 507px" /></a></p>
<p><strong>b</strong>: Location of &#8220;Springer Spaniel &#8211; English&#8221;</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48117" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png" alt="Breed list showing Springer Spaniel - English" width="531" height="315" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F1b.png 531w, https://measuringu.com/wp-content/uploads/2026/08/081126-F1b-300x178.png 300w" sizes="auto, (max-width: 531px) 100vw, 531px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Participant A&#8217;s usability issues with breed selection.</p>
<h3><span lang="EN-US">Some Agreement Between Human and AI, No Hallucinations, Six False Alarms</span></h3>
<p>Table 1 shows the four problems discovered by the human researcher (three of which were also identified by the AIs) and the six additional problems reported only by the AIs. We looked to see if the problems identified only by the AIs were false alarms (an event happened but not really a usability problem) or hallucinations (the event just didn&#8217;t happen).</p>
<p>The good news was that none of the six problems were hallucinations. The bad news is that they were all determined to be false alarms. Either the participant never actually noticed the issue flagged by the AI or wasn&#8217;t affected by it (e.g., the location prompt, the below-the-fold item), or it was just normal, expected system behavior rather than a flaw (e.g., the ZIP search returning multiple locations, the menu reloading).</p>

<table id="tablepress-1061" class="tablepress tablepress-id-1061">
<thead>
<tr class="row-1">
	<th class="column-1">#</th><th class="column-2">Problem description</th><th class="column-3">Human</th><th class="column-4">ChatGPT</th><th class="column-5">Gemini</th><th class="column-6">Why (if false alarm)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"><center>Y</td><td class="column-4"><center>--</td><td class="column-5"><center>--</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-4">
	<td class="column-1">H3</td><td class="column-2">No attempt to filter breed list by typing; scrolled instead</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-5">
	<td class="column-1">H4</td><td class="column-2">Searched breed list for Springer Spaniel, couldn't find it alphabetically</td><td class="column-3"><center>Y</td><td class="column-4"><center>Y</td><td class="column-5"><center>Y</td><td class="column-6"><center>--</td>
</tr>
<tr class="row-6">
	<td class="column-1">C1</td><td class="column-2">Blank/loading state after selecting Grooming Salon</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Lasted ~2 seconds; participant had already moved on</td>
</tr>
<tr class="row-7">
	<td class="column-1">C2/G1</td><td class="column-2">Browser location prompt appears alongside site's own location modal</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>N</td><td class="column-6">No sign participant noticed it; used site's own ZIP entry instead</td>
</tr>
<tr class="row-8">
	<td class="column-1">C3</td><td class="column-2">Entering ZIP returns multiple grooming locations</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Accurate, but that's how the control is designed to work</td>
</tr>
<tr class="row-9">
	<td class="column-1">C4</td><td class="column-2">Menu reloads after selecting breed and age</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">Expected behavior after Check Prices &amp; Book Now; no impact</td>
</tr>
<tr class="row-10">
	<td class="column-1">C5/G2</td><td class="column-2">Target service is at the bottom of the list, below the fold</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>N</td><td class="column-6">True, but not actually a problem for this participant</td>
</tr>
<tr class="row-11">
	<td class="column-1">C6</td><td class="column-2">Promotional content crowds out the service menu</td><td class="column-3"><center>--</td><td class="column-4"><center>N</td><td class="column-5"><center>--</td><td class="column-6">True, but not an obvious problem for this participant</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1061 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>Summary of usability problem discovery by the human researcher and the AIs. In the # column, H indicates a problem identified by the human researcher, C indicates a problem identified by ChatGPT, and G indicates a problem identified by Gemini. In the Human, ChatGPT, and Gemini columns, Y indicates the discovery of a verified usability problem, and N indicates the reporting of a false alarm. The full runs of ChatGPT and Gemini are shown in Appendix A.</p>
<p>As we did in our first study, we ran the videos four times through the AIs (because of the <a href="https://measuringu.com/can-ai-detect-usability-problems/">probabilistic nature of how they work</a>). Problems identified in any of the four runs were included in this analysis. We used the mean <a href="https://measuringu.com/ai-usability-problem-analysis-of-a-video/">any-2 agreement</a> to assess overlap.</p>
<p><strong><em>Technical note</em></strong><em><strong>:</strong> Our preferred method for quantifying the correspondence between two lists of usability issues is </em><em>any-2 agreement</em><em>. Any-2 agreement is the ratio of the intersection of the two sets divided by their union. Historically, we’ve found an any-2 agreement of 50% to be average (typical), around 25% to be low, and around 75% to be high.</em></p>
<h4><span lang="EN-US">ChatGPT Agreement: 35%</span></h4>
<p>The mean any-2 agreement of the four ChatGPT runs and the UX researcher was 35%. ChatGPT identified (at least once) three of the four problems reported by the researcher but also produced six false alarms.</p>
<h4><span lang="EN-US">Gemini Agreement: 38%</span></h4>
<p>The mean any-2 agreement of the four Gemini runs and the UX researcher was 38%. Gemini identified (at least once) three of the four problems reported by the researcher but also produced two false alarms (matching two of the false alarms produced by ChatGPT).</p>
<p>The mean any-2 agreement between the four runs of the AIs was 40%. Figure 2 shows the Venn diagram for the problem discovery results for the UX researcher and the AIs.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48120" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png" alt="Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini." width="991" height="743" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F2.png 991w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-300x225.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-768x576.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-F2-600x450.png 600w" sizes="auto, (max-width: 991px) 100vw, 991px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini.</p>
<p>The Venn diagram illustrates the relationship between the human reviewer and AI analyses of Participant A. The human reviewer identified four usability issues, three of which were identified at least once by an AI. However, there was <strong>one usability problem that was not caught by either AI</strong>, while the <strong>AIs produced six issues that were not legitimate usability problems</strong> (all false alarms, no hallucinations). The AIs did not discover any real problems that the human reviewer failed to identify.</p>
<h2><span lang="EN-US">Comparison with the OpenTable Case Study Results</span></h2>
<p>Figure 3 shows the Venn diagram from our <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">OpenTable case study</a>.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48121" src="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png" alt="Venn diagram of usability problem discovery from our OpenTable case study." width="965" height="724" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-F3.png 965w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-300x225.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-768x576.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-F3-600x450.png 600w" sizes="auto, (max-width: 965px) 100vw, 965px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3: </strong>Venn diagram of usability problem discovery from our OpenTable case study.</p>
<p>At a glance, the diagrams in Figures 2 and 3 have some similarities and some differences. To help with comparison, Table 2 shows the side-by-side comparisons for different aspects of the results (graphed in Figure 4).</p>

<table id="tablepress-1062" class="tablepress tablepress-id-1062">
<thead>
<tr class="row-1">
	<th class="column-1">Comparison</th><th class="column-2">Participant A</th><th class="column-3">OpenTable</th><th class="column-4">Abs. Diff.</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">% human identified</td><td class="column-2">40%</td><td class="column-3">45%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-3">
	<td class="column-1">% AI &amp; human overlap</td><td class="column-2">30%</td><td class="column-3">30%</td><td class="column-4"> 0%</td>
</tr>
<tr class="row-4">
	<td class="column-1">% AI-only identified</td><td class="column-2">60%</td><td class="column-3">55%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-5">
	<td class="column-1">% AI errors</td><td class="column-2">60%</td><td class="column-3">50%</td><td class="column-4">10%</td>
</tr>
<tr class="row-6">
	<td class="column-1">% AI false alarms</td><td class="column-2">60%</td><td class="column-3">35%</td><td class="column-4">25%</td>
</tr>
<tr class="row-7">
	<td class="column-1">% AI hallucinations</td><td class="column-2"> 0%</td><td class="column-3">15%</td><td class="column-4">15%</td>
</tr>
<tr class="row-8">
	<td class="column-1">% AI-only discovery</td><td class="column-2"> 0%</td><td class="column-3"> 5%</td><td class="column-4"> 5%</td>
</tr>
<tr class="row-9">
	<td class="column-1">Total unique problems</td><td class="column-2">10</td><td class="column-3">20</td><td class="column-4">10</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1062 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>Comparison of PetSmart Participant A and OpenTable problem identification rates.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48157 size-large" src="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1024x525.jpg" alt="" width="1024" height="525" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1024x525.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-300x154.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-768x393.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-1536x787.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-2048x1049.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/08/081126-Figure-1-600x307.jpg 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4: </strong>PetSmart Participant A and OpenTable problem identification rates.</p>
<p>In most respects, <strong>this second evaluation of AI problem discovery has replicated the first case study</strong> (OpenTable).</p>
<p>The usability problem identification rates were similar (within 10 percentage points) for the percentages of usability problems identified by the human UX researchers, AIs, both (the AI/Human overlap), and AI-only discovery.</p>
<p>Observed differences in the patterns were due to the incidence of AI hallucinations in the OpenTable case study (3) compared to none in the AI outputs for Participant A.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Our key findings were:</p>
<p><strong>Successful replication of the OpenTable case study.</strong> Most metrics comparing the two videos were within 10 percentage points of each other (e.g., 40% vs. 45% of human-identified problems found by AI). The one real divergence was in the <em>type</em> of AI-only errors: this PetSmart video had more false alarms (60% vs. 35%) but zero hallucinations.</p>
<p><strong>A troubling number of false alarms.</strong> Focusing on the new data for PetSmart Participant A, the total number of unique usability problems was ten, of which only four were identified by the UX researcher (which we treat as ground truth). Of the six unique problems the AIs reported that the human researcher did not flag, all were false alarms (no hallucinations).</p>
<p><strong>AI adds potential value as an overly enthusiastic junior researcher, not a trusted expert.</strong> In the analysis of these videos, the AIs discovered three of the four usability problems reported by the UX researcher. Relying only on these multiple runs of the AIs would have missed a quarter of the real usability problems. Unlike our earlier research with the restaurant reservation video, where the <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">AIs found one problem that UX researchers missed</a>, the AIs in this study did not identify any valid usability problems that the UX researcher failed to discover.</p>
<p><strong>Like humans, AI usability reviews of videos are prone to the “evaluator effect.”</strong> Just like human evaluators, multiple runs of AI usability evaluations of videos are not perfectly consistent, so it’s good practice to run these evaluations multiple times for consistency checks. Running multiple evaluations and looking for consistency across runs is a practical filter before any human review.</p>
<p><strong>Bottom line—AI usability reviews of videos require human oversight.</strong> In their current form (what we tested), these AI products can add some value to this type of UX research, but more as junior researchers whose actions and conclusions require expert human oversight rather than as trusted experts themselves.</p>
<p><strong>Future research:</strong> Our next step in this research program is to perform the same analyses on two PetSmart videos in which the user experiences were different from PetSmart Participant A and OpenTable regarding task success and number of problems identified by the human UX researcher (one who did not complete the task successfully and one who experienced no problems completing the task).</p>
<h2><span lang="EN-US">Appendix A: Detailed Timeline and Problem-by-Problem Tables</span></h2>
<h3><span lang="EN-US">Key Events Timeline for Participant A</span></h3>
<p>Appendix Table 1 summarizes the key events in the video (compiled by the UX researcher), identifying four problematic events deviating from the “happy” path (see Appendix B).</p>

<table id="tablepress-1063" class="tablepress tablepress-id-1063">
<thead>
<tr class="row-1">
	<th class="column-1">Event #</th><th class="column-2">Timestamp</th><th class="column-3">Summary of key user actions</th><th class="column-4">Notes</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">1</td><td class="column-2">0:00:13</td><td class="column-3">Clicked Shop by Pet from top menu</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-3">
	<td class="column-1">2</td><td class="column-2">0:00:21</td><td class="column-3">Clicked search field</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-4">
	<td class="column-1">3</td><td class="column-2">0:00:24</td><td class="column-3">Clicked Pet Services</td><td class="column-4"></td>
</tr>
<tr class="row-5">
	<td class="column-1">4</td><td class="column-2">0:00:44</td><td class="column-3">Grooming form begins to appear</td><td class="column-4"></td>
</tr>
<tr class="row-6">
	<td class="column-1">5</td><td class="column-2">0:00:45</td><td class="column-3">All form elements except dog/cat buttons appear</td><td class="column-4"></td>
</tr>
<tr class="row-7">
	<td class="column-1">6</td><td class="column-2">0:00:46</td><td class="column-3">Dog/cat buttons appear—participant cursor was headed for breed but when these buttons appeared he changed course toward the dog/cat buttons</td><td class="column-4"></td>
</tr>
<tr class="row-8">
	<td class="column-1">7</td><td class="column-2">0:00:49</td><td class="column-3">Clicked Dog button</td><td class="column-4"></td>
</tr>
<tr class="row-9">
	<td class="column-1">8</td><td class="column-2">0:00:56</td><td class="column-3">Clicked "select" link by Select a Store</td><td class="column-4"></td>
</tr>
<tr class="row-10">
	<td class="column-1">9</td><td class="column-2">0:00:58</td><td class="column-3">Find a Grooming Salon Near You pops up (one field: "Zip Code, City or State" and Search button)</td><td class="column-4"></td>
</tr>
<tr class="row-11">
	<td class="column-1">10</td><td class="column-2">0:00:59</td><td class="column-3">Location permission prompt appeared at top of screen apparently triggered by presentation of the Use My Current Location link in Find a Grooming Salon Near You</td><td class="column-4"></td>
</tr>
<tr class="row-12">
	<td class="column-1">11</td><td class="column-2">0:01:09</td><td class="column-3">Typed zip code from task instructions and clicked Search</td><td class="column-4"></td>
</tr>
<tr class="row-13">
	<td class="column-1">12</td><td class="column-2">0:01:13</td><td class="column-3">List of locations appears with target Glendale at top</td><td class="column-4"></td>
</tr>
<tr class="row-14">
	<td class="column-1">13</td><td class="column-2">0:01:19</td><td class="column-3">Clicked Glendale</td><td class="column-4"></td>
</tr>
<tr class="row-15">
	<td class="column-1">14</td><td class="column-2">0:01:21</td><td class="column-3">Clicked x to clear the location permission prompt</td><td class="column-4"></td>
</tr>
<tr class="row-16">
	<td class="column-1">15</td><td class="column-2">0:01:29</td><td class="column-3">Clicked Breed dropdown, list of breeds appears</td><td class="column-4"></td>
</tr>
<tr class="row-17">
	<td class="column-1">16</td><td class="column-2">0:01:30</td><td class="column-3">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-18">
	<td class="column-1">17</td><td class="column-2">0:01:35</td><td class="column-3">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-4">Problematic event</td>
</tr>
<tr class="row-19">
	<td class="column-1">18</td><td class="column-2">0:01:56</td><td class="column-3">Scrolling farther down the list found "Springer Spaniel - English"</td><td class="column-4"></td>
</tr>
<tr class="row-20">
	<td class="column-1">19</td><td class="column-2">0:02:00</td><td class="column-3">Clicked Age dropdown</td><td class="column-4"></td>
</tr>
<tr class="row-21">
	<td class="column-1">20</td><td class="column-2">0:02:04</td><td class="column-3">Selected "6 months or older"</td><td class="column-4"></td>
</tr>
<tr class="row-22">
	<td class="column-1">21</td><td class="column-2">0:02:07</td><td class="column-3">Clicked Check Prices &amp; Book Now button</td><td class="column-4"></td>
</tr>
<tr class="row-23">
	<td class="column-1">22</td><td class="column-2">0:02:13</td><td class="column-3">Grooming Salon Menu appeared</td><td class="column-4"></td>
</tr>
<tr class="row-24">
	<td class="column-1">23</td><td class="column-2">0:02:32</td><td class="column-3">Scrolled through list of services to the bottom</td><td class="column-4"></td>
</tr>
<tr class="row-25">
	<td class="column-1">24</td><td class="column-2">0:02:35</td><td class="column-3">Did not click service but said, "Just bath and full haircut, $61, let me write that down."</td><td class="column-4"></td>
</tr>
<tr class="row-26">
	<td class="column-1">25</td><td class="column-2">0:02:36</td><td class="column-3">Task successfully completed (correct location and service)</td><td class="column-4"></td>
</tr>
</tbody>
</table>
<!-- #tablepress-1063 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 1: </strong>Timeline for Participant A.</p>
<h3><span lang="EN-US">ChatGPT Problem-by-Problem Results</span></h3>
<p>Appendix Table 2 shows the usability problems reported by ChatGPT and the UX researcher, run by run.</p>

<table id="tablepress-1064" class="tablepress tablepress-id-1064">
<thead>
<tr class="row-1">
	<th class="column-1"><center>Prob #</th><th class="column-2">Description</th><th class="column-3"><center>Run 1</th><th class="column-4"><center>Run 2</th><th class="column-5"><center>Run 3</th><th class="column-6"><center>Run 4</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-4">
	<td class="column-1">C1</td><td class="column-2">Page shows a mostly blank/loading state after selecting Grooming Salon (FALSE ALARM—this happened but only lasted about two seconds within which the participant started to select Breed but diverted to Dog when that button appeared)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"><center>1</td>
</tr>
<tr class="row-5">
	<td class="column-1">C2</td><td class="column-2">Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-6">
	<td class="column-1">C3</td><td class="column-2">Participant enters 80246 and gets multiple grooming locations (FALSE ALARM—true but this is how this control is supposed to work)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-7">
	<td class="column-1">H3</td><td class="column-2">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-8">
	<td class="column-1">H4</td><td class="column-2">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-9">
	<td class="column-1">C4</td><td class="column-2">After selecting breed and age, the menu reloads. (FALSE ALARM—true but happens quickly and is the expected action after clicking Check Prices &amp; Book Now—no impact on participant behavior)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-10">
	<td class="column-1">C5</td><td class="column-2">The target option is at the bottom of the list below the fold/below the Bath&amp; Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)</td><td class="column-3"></td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-11">
	<td class="column-1">C6</td><td class="column-2">Large promotional grooming content and video tiles take up much of the page while the actual service menu is constrained to the right side (FALSE ALARM—true but not an obvious problem for this participant)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1064 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 2: </strong>Usability problems reported by ChatGPT 5.4 Thinking and the UX researcher for Participant A.</p>
<h3><span lang="EN-US">Gemini Problem-by-Problem Results</span></h3>
<p>Appendix Table 3 shows the usability problems reported by Gemini and the UX researcher, run by run.</p>

<table id="tablepress-1065" class="tablepress tablepress-id-1065">
<thead>
<tr class="row-1">
	<th class="column-1"><center>Prob #</th><th class="column-2">Description</th><th class="column-3"><center>Run 1</th><th class="column-4"><center>Run 2</th><th class="column-5"><center>Run 3</th><th class="column-6"><center>Run 4</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">H1</td><td class="column-2">Clicked Shop by Pet from top menu (incorrect first click)</td><td class="column-3"></td><td class="column-4"><center>1</td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-3">
	<td class="column-1">H2</td><td class="column-2">Clicked search field (incorrect second click)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-4">
	<td class="column-1">G1</td><td class="column-2">Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)</td><td class="column-3"><center>1</td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"></td>
</tr>
<tr class="row-5">
	<td class="column-1">H3</td><td class="column-2">Made no attempt to filter list by typing; instead started scrolling through the long list</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"></td><td class="column-6"></td>
</tr>
<tr class="row-6">
	<td class="column-1">H4</td><td class="column-2">Searching list for English Springer Spaniel, scrolling up and down between the D's and E's but couldn't find the breed there</td><td class="column-3"><center>1</td><td class="column-4"><center>1</td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
<tr class="row-7">
	<td class="column-1">G2</td><td class="column-2">The target option is at the bottom of the list below the fold/below the Bath&amp; Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)</td><td class="column-3"></td><td class="column-4"></td><td class="column-5"><center>1</td><td class="column-6"><center>1</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1065 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Table 3: </strong>Usability problems reported by Gemini 3 Flash Thinking and the UX researcher for Participant A.</p>
<h2><span lang="EN-US">Appendix B: The PetSmart Reservation Task</span></h2>
<p>For a <a href="https://measuringu.com/ux-pets/">UX benchmark study conducted in 2019</a>, one of the participants’ tasks was to use the PetSmart website to start booking a grooming appointment in Glendale, CO for a one-year-old English Springer Spaniel, then stop after determining the cost of a bath and full haircut. In this section, we review the steps through the “happy path” and speculate about possible user behaviors that would be reasonable to track to provide background knowledge for understanding the problem lists presented later.</p>
<p>Appendix Figure 1 shows the home page. Before continuing, ask yourself, where would you start?</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48131" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png" alt="PetSmart home page for the pet grooming task." width="624" height="306" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF1.png 624w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF1-300x147.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF1-600x294.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 1: </strong>Home page for the pet grooming task.</p>
<p>For this task, the best first choice is to click “pet services” from the horizontal navigation menu close to the top of the page. From here, there are two paths to grooming, shown in Appendix Figure 2. Appendix Figure 2a shows the dropdown from which a user could drag the cursor down and release the button to select Grooming Salon. Appendix Figure 2b shows the pet services menu that appears after clicking “pet services” but releasing the mouse button without dragging, from which the user would click Grooming.</p>
<p><strong>Appendix Figure 2a: Click and drag path</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48132" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png" alt="Click and drag path to grooming" width="780" height="328" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-600x252.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><strong>Appendix Figure 2b: Click without dragging path</strong><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48132" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png" alt="Click and drag path to grooming" width="780" height="328" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-300x126.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-768x323.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2a-600x252.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48133" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png" alt="Click without dragging path." width="780" height="308" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-300x118.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-768x303.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF2b-600x237.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 2: </strong>Two paths to the grooming menu.</p>
<p>Appendix Figure 3 shows the grooming form. This is where users who are not in Glendale can change the location to Glendale, select dog, select the breed, select the age, then click the button to check prices.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48135" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png" alt="The grooming form." width="780" height="93" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-300x36.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-768x92.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF3-1-600x72.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 3: </strong>The grooming form.</p>
<p>Before checking prices, users needed to select a breed and age for their dog. As shown in Appendix Figure 4, clicking breed produced a searchable breed dropdown. Appendix Figure 4a shows the initial appearance of the dropdown; Appendix Figure 4b shows its appearance after typing “english&#8221; over the placeholder text &#8220;breed&#8221; in the combobox.</p>
<p><strong>Appendix Figure 4a: Initial appearance of the breed dropdown</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48136" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png" alt="Initial appearance of the breed dropdown." width="780" height="217" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-300x83.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-768x214.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4a-600x167.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p><strong>Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48137" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png" alt="Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”" width="780" height="177" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-300x68.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-768x174.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF4b-600x136.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 4: </strong>The searchable breed dropdown.</p>
<p>For the happy path, a user should type “english” into the combobox, so one of the potential problems we anticipated was users not realizing the dropdown list could be filtered. Compounding the complexity of this step in the process is that the location of the word “English” for various breeds was inconsistent. For example, after filtering, the list in Figure 4b included English Toy Spaniel, Old English Sheepdog, and Springer Spaniel &#8211; English. That’s less of a problem after filtering but could be more problematic if scrolling through the unfiltered list.</p>
<p>The age dropdown, shown in Appendix Figure 5, was relatively straightforward with only two choices.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48138" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png" alt="The age dropdown." width="780" height="128" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF5.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-300x49.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-768x126.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF5-600x98.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 5: </strong>The age dropdown.</p>
<p>With the grooming form completed, the next step is to click “check prices &amp; book now” to get the list of grooming options shown in Appendix Figure 6. Because the target option was the last one in the list and below the fold, we anticipated that some users might select an earlier option.</p>
<p><strong>Appendix Figure 6a: Completed grooming menu and first option in Grooming Salon Menu (above the fold)<a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48139 size-full" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png" alt="Completed grooming menu and first option in Grooming Salon Menu (above the fold)" width="780" height="359" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-300x138.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-768x353.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6a-600x276.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></strong></p>
<p><strong>Appendix Figure 6b: The other grooming options (below the fold; the last option is the target)</strong></p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48140" src="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png" alt="The other grooming options (below the fold, last option is the target)" width="780" height="304" srcset="https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b.png 780w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-300x117.png 300w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-768x299.png 768w, https://measuringu.com/wp-content/uploads/2026/08/081126-AF6b-600x234.png 600w" sizes="auto, (max-width: 780px) 100vw, 780px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Appendix Figure 6: </strong>Grooming options above and below the fold, showing the target Bath &amp; Full Haircut for $61.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Can We Trust AI to Moderate UX Interviews?</title>
		<link>https://measuringu.com/can-we-trust-ai-to-moderate-ux-interviews/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=can-we-trust-ai-to-moderate-ux-interviews</link>
		
		<dc:creator><![CDATA[Jeff Sauro, PhD •&nbsp;Lucas Plabst, PhD&nbsp;•&nbsp;Jim Lewis, PhD]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 23:01:15 +0000</pubDate>
				<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[mod]]></category>
		<category><![CDATA[moderated research]]></category>
		<category><![CDATA[Moderating]]></category>
		<category><![CDATA[Moderation]]></category>
		<category><![CDATA[moderator]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48092</guid>

					<description><![CDATA[Can AI Replace UX Researchers? Some criticized us for using that headline in 2023, calling it clickbait. Fair criticism, given the inflammatory nature of AI (for the record, we did have a subtitle). But here we are three years later, and we’re not talking about just the tedious task of having AI code comments—a task [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48106 size-medium" src="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-300x169.jpg" alt="Feature image showing an AI moderator" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/08/080426-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a><em>Can AI Replace UX Researchers?</em></p>
<p>Some criticized us for using that <a href="https://measuringu.com/classification-agreement-between-ux-researchers-and-chatgpt/">headline in 2023</a>, calling it clickbait. Fair criticism, given the inflammatory nature of AI (for the record, we did have a subtitle). But here we are three years later, and we’re not talking about just the tedious task of having AI code comments—a task that most researchers are happy to offload. We’re talking about AI being able to do some of the core jobs of the UX researcher, in particular, moderating participant sessions.</p>
<p>When we wrote our article three years ago about AI doing comment coding, the idea of replacing a UX moderator seemed about as probable as replacing a waiter in a restaurant. Both jobs seemed immune to AI. Yet here we are. AI isn’t just crunching numbers; it’s being tasked with doing qualitative work. <span data-preserver-spaces="true">It’s</span><span data-preserver-spaces="true"> a serious enough threat that </span><a class="editor-rtfLink" href="https://journals.sagepub.com/doi/full/10.1177/10778004251401851" target="_blank" rel="noopener"><span data-preserver-spaces="true">419 professionals </span></a><span data-preserver-spaces="true">have signed letters against the use of AI for qual research.</span></p>
<p>And just like we did with <a href="https://measuringu.com/review-of-experiments-with-synthetic-users/">synthetic users</a>, we need to separate the hype from the data. That means asking what the claim actually is, who it comes from, and how strong the evidence behind it is. Is it anecdotal? Is it peer-reviewed? Is it from a company selling AI moderators?</p>
<p>Before we dig into the claims about AI and moderation and start letter-writing campaigns, it’s important to understand what moderating is, what an AI moderator is, and what one can do.</p>
<h2><span lang="EN">What Is Moderating in UX Research?</span></h2>
<p>Moderating is a general term that’s not to be confused with content moderation on social networks. In UX research, <a href="https://measuringu.com/what-makes-a-good-ux-research-moderator/">moderating is a broad term</a> that spans several UX research methods: <a href="https://measuringu.com/three-goals/">usability testing</a>, interviews, <a href="https://measuringu.com/contextual-inquiry/">contextual inquiry</a>, and occasionally <a href="https://measuringu.com/f-word/">focus groups</a>. The moderator&#8217;s job changes depending on the method and goals.</p>
<p>An unstructured in-depth interview where you need to uncover problems and opportunities for a new product requires a different approach than a large-scale summative evaluation where you need minimal interaction to collect metrics (yes, you can moderate and collect metrics).</p>
<h2>What Makes a Moderator Good?</h2>
<p>Beyond knowing the study type, the practical skills that separate strong moderators from weak ones are things like building enough rapport that participants give honest reactions, probing appropriately (asking a follow-up once or twice to get past a surface answer, without dragging out a dead end), catching participants who are misrepresenting themselves or reading from a script, using silence to let someone think rather than rushing to fill it, and knowing when to go off-script versus when to stick to the guide.</p>
<p>None of this is really a checklist. Instead, it&#8217;s judgment applied in real time to what a specific person just said. Even with that list, there is no agreed-upon way to score whether a moderator did a good job. What makes a good moderator isn’t even clearly defined, let alone measured objectively.</p>
<h2><span lang="EN">What Is an AI Moderator?</span></h2>
<p>An AI moderator, as the name suggests, is software that uses AI to interactively interview participants using voice and even video to replicate the interactivity of a human. But it’s more than software reading a script and asking questions. An AI can be prompted to probe and follow up based on what the participants say. The technology has advanced considerably since Dragon NaturallySpeaking in the 1990s and Google’s cutting-edge transcription <a href="https://technologizer.com/2010/08/22/worst-google-voice-transcription-errors/index.html">failures of the 2010s</a>. AI conversations now flow with high accuracy and little latency.</p>
<h2><span lang="EN">AI Moderators (Not Quite HAL 9000)</span></h2>
<p>While the technology has come a long way, AI moderators are not like the diabolical AI in <em>2001: A Space Odyssey</em>. They are limited by the script you feed them and in their interactive ability.</p>
<p>An AI moderator is software (prompts) built on top of the same LLM models that make headlines (Claude, ChatGPT, Gemini). The moderators can be voice-only or have a range of appearances, from simple visualizations to more life-like avatars.</p>
<p><a href="https://chatgpt.com/features/voice/">ChatGPT</a>, <a href="https://gemini.google/overview/gemini-live/">Gemini</a>, and <a href="https://support.claude.com/en/articles/11101966-use-voice-mode">Claude</a> all ship voice modes that hold unstructured back-and-forth conversations, interruptions included, with no command list behind them. Figure 1 shows the appearance of ChatGPT&#8217;s voice interface compared with HAL 9000 from <em>2001: A Space Odyssey</em>.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1a.png"><img loading="lazy" decoding="async" class="alignnone wp-image-48093" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1a.png" alt="ChatGPT’s visual placeholder for its voice." width="198" height="216" /></a><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1b.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48094" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F1b.png" alt="HAL 9000" width="288" height="216" /> </a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>ChatGPT’s visual placeholder for its voice (left) and HAL 9000 (right).</p>
<p>An example of a simplified prompt for the LLM could be something like the following for doing research on a new sleep-aid product:</p>
<blockquote><p>You are a UX researcher running semi-structured interviews with participants in a research study. The goal of the study is to find out more about their sleeping behavior and to find potential market opportunities for wake-up devices. Follow the discussion guide you are given directly, covering all required questions, but follow up on interesting threads.</p></blockquote>
<p>AI moderators can also have avatars to go with the voice, so respondents aren’t talking to a circle or blank screen. Figure 2 shows an example of one used for job interviews.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48095" src="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png" alt="Example of a visual AI moderator used for job interviews from Humanly.io." width="624" height="301" srcset="https://measuringu.com/wp-content/uploads/2026/08/08042026-F2.png 624w, https://measuringu.com/wp-content/uploads/2026/08/08042026-F2-300x145.png 300w, https://measuringu.com/wp-content/uploads/2026/08/08042026-F2-600x289.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2: </strong>Example of a visual AI moderator used for job interviews from Humanly.io.</p>
<h2><span lang="EN">AI Moderation in Job Interviews</span></h2>
<p>We can talk about AI moderators as a concept, but are they really being used at scale? One clear application: interviewing for a job, where services like <a href="http://codesignal.com">CodeSignal</a>, <a href="http://humanly.io">Humanly</a>, or <a href="https://eightfold.ai/">Eightfold</a> offer AI interviewers.</p>
<p>One estimate suggests over <a href="https://www.greenhouse.com/newsroom/63-of-job-seekers-have-faced-an-ai-interview-most-havent-had-a-good-one-yet">60% of job seekers</a> have had exposure to an AI interview. Note that this estimate comes from a company selling AI moderating technology, so it’s not necessarily objective.</p>
<p>However, a large <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5395709">pre-registered experiment from 2025</a> with over 70,000 job applicants in the Philippines provides a more objective data point. Participants were randomly assigned to either a human interviewer (20% of sample), an AI job interviewer (60% of sample), or allowed to choose between the two (20% of sample) after a brief intro to the AI interviewer.</p>
<p>Candidates interviewed by the AI were 12% more likely to get an offer, 18% more likely to start, and 17% more likely to still be employed after a month. Of the subset that had a choice, a surprisingly high 78% picked the AI.</p>
<p>These numbers are hard to ignore. Consider a few caveats, though, before we fire all the job interviewers and UX moderators out there. First, this was for entry-level customer service jobs (call center) in the Philippines. Second, humans made the hiring decision, not AI. Probably most importantly, the value that was derived was in the perception and reality of a consistently delivered interviewer. The type of questions being asked were mostly closed, factual, verification-style questions, things like commute time, salary fit, availability, and contact info.</p>
<p>Job interviews are notorious for both explicit and implicit bias. There does seem to be a benefit if those biases can be significantly reduced (or even eliminated).</p>
<p>Interviewing a candidate for a job isn’t the same as interviewing a participant to understand usage patterns. While we are concerned about <a href="https://measuringu.com/ut-bias/">identifying and reducing bias</a> in UX research, UX moderation is often best deployed for unstructured and unknown problems.</p>
<p>Job interviews, especially screening interviews, are usually very structured with a clear set of questions. They are designed to be consistent. A structured UX interview is really more akin to a verbalized survey (something we’ll revisit). Where interviews really matter is when you don’t know what you don’t know.</p>
<h2><span lang="EN">AI in Qual Research</span></h2>
<p>Job interviewing is qualitatively different from the sort of inquiry in moderated research. In one of the best-known AI interviewer studies, in December 2025, Anthropic conducted a large study of 80k Claude users across 159 countries using their new Anthropic Interviewer tool. Their interview questions were a bit closer to the types of topics investigated in UX moderated interviews:</p>
<ol>
<li>What&#8217;s the last thing you used an AI chatbot for?</li>
<li>If you could wave a magic wand, what would AI do for you?</li>
<li>Has AI ever taken a step towards that vision for you?</li>
<li>Are there ways AI might be developed that would be contrary to your vision or what you value?</li>
</ol>
<p>Anthropic Interviewer followed up on each, probing for the underlying values and experiences behind people&#8217;s answers. Claude was then used to synthesize the transcripts for themes (much to the chagrin of the 419 professionals who rejected such usage).</p>
<p>Claude’s analysis determined that 88% to 98% of the open-ended comments were “substantive.” There aren’t many details to assess the quality relative to a human, but the authors reported that a subset of comments were validated as having at least 90% agreement with human coders on 25 labels. The authors were surprised by how candid some people were. Respondents shared things like grief, mental health crises, financial precarity, and relationship failures, all of which our human user researchers rarely encounter in traditional interviews. It’s unclear if those same comments would have been shared with a human, but it’s certainly a possible benefit worth investigating more when a topic may elicit more sensitive comments from participants.</p>
<p>What is clear is that this is a huge sample that would almost certainly never happen with human moderation. It’s less clear how well the insights from a smaller sample conducted by humans would have performed.</p>
<p>Anthropic isn’t the only one in the AI interviewer game, and others have noticed the self-disclosure. <a href="https://heymarvin.com/product/ai-moderated-interviewer">Marvin’s</a> AI Interviewer “leads human-like conversations, searches for the why behind responses, and gathers insights faster than ever before.” Or, from <a href="https://getperspective.ai/agents/interviewer">Perspective</a>, “customers share things in these conversations they&#8217;d never put in a form and would rarely say on a Zoom call with a stranger. Not a chatbot, not a survey, not a junior researcher reading from a script—the best interviewer you&#8217;ve ever seen, available at any hour, in any language, for every single customer.”</p>
<p>Bold vendor claims like these are worth investigating. <a href="https://www.nngroup.com/articles/ai-interviewers/">NN/Group’s</a> analysis of ten experienced researchers using two AI moderators (including <a href="https://heymarvin.com/product/ai-moderated-interviewer">Marvin</a> from above and <a href="https://userflix.de/">UserFlix</a>) found that AI interviewers can handle structured, scripted interviews (at scale). However, they didn’t think AI interviewers were adequate yet for semi-structured interviews.</p>
<h2>Are We Ready to Deploy the AI Moderators?</h2>
<p>So, we have some data suggesting that AI can be used at scale to interview job candidates. We have at least a proof of concept that AI can moderate a massive number of sessions for a handful of more open-ended research questions. And there’s some evidence that people may be more willing to disclose more to an AI moderator than to a human.</p>
<p>But if you recruit a dozen IT decision makers, or Chief Product Officers, and want to build them a better product by conducting a semi-structured interview that requires a moderator to go off script, would you use an AI moderator? NN/Group certainly suggests we aren’t there.</p>
<p>But this raises the question of how well an AI moderator would do compared to a human. That’s a question we’ll take up in our upcoming articles. First, we’ll review the published literature, and then we&#8217;ll report on the results of our own controlled experiment, where we compare an AI moderator with a human moderator in a UX research context.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Measuring the UX of AI</title>
		<link>https://measuringu.com/measuring-the-ux-of-ai/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=measuring-the-ux-of-ai</link>
		
		<dc:creator><![CDATA[Jeff Sauro, PhD&nbsp;•&nbsp;Jim Lewis, PhD]]></dc:creator>
		<pubDate>Tue, 28 Jul 2026 21:47:36 +0000</pubDate>
				<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48046</guid>

					<description><![CDATA[AI is everywhere and getting embedded in all of our products. If you ask the typical person in 2026 what AI is, they’ll probably say it’s a generative chat product like ChatGPT, Copilot, Claude, or Gemini. Of course, these are only the frontier Large Language Models (LLMs). That&#8217;s not all AI is. For example, the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-48076 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-300x169.jpg" alt="Feature image showing a researcher measuring the UX of AI" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/072826-FeatureImage-2.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>AI is everywhere and getting embedded in all of our products.</p>
<p>If you ask the <a href="https://hbr.org/2026/06/how-people-are-really-using-ai-in-2026">typical person in 2026 what AI is</a>, they’ll probably say it’s a generative chat product like ChatGPT, Copilot, Claude, or Gemini. Of course, these are only the <a href="https://www.iguazio.com/glossary/frontier-model/">frontier Large Language Models</a> (LLMs). That&#8217;s not all AI is.</p>
<p>For example, the algorithm recommending your next Netflix show, the AI drafting a formula in Excel, or the model flagging fraud on your credit card are all types of AI (none of which have chat interfaces).</p>
<p>But when most people say they&#8217;re “using AI,” they mean typing into a chat box, and that&#8217;s a good place to start when thinking about how to measure the UX of AI as it’s popularly understood.</p>
<p>How should we measure the quality of these experiences?</p>
<h2><span lang="EN-US">Measuring UX in General</span></h2>
<p>Measuring the user experience in general involves assessing <a href="https://measuringu.com/ux-measurement-purpose/">what people think and feel, and what people do</a>. That means using a mix of <a href="https://measuringu.com/get-comfortable-with-four-ux-metrics/">attitudinal and action measures</a>.</p>
<p>Action (behavioral) measures are more straightforward to interpret. A typical suite of action metrics includes a combination of effectiveness (<a href="https://measuringu.com/completion-rates/">completion rates</a>, <a href="https://measuringu.com/errors-ux/">errors</a>) and efficiency (<a href="https://measuringu.com/task-times/">time on task</a>). But they’re relatively hard to collect because you have to set up <a href="https://measuringu.com/task-based-metrics/">task scenarios</a> and record or observe behaviors.</p>
<p>Attitudinal measures are easier to collect, but you need to be sure you’re measuring the right thing.</p>
<p>For attitudes, we’ve found that standardized metrics like the <a href="https://measuringu.com/how-to-score-and-interpret-the-ux-lite/">UX-Lite</a><sup>®</sup> provide a good measure of overall attitudes toward a product’s usefulness (capabilities/features) and usability (ease of use). For example, see our <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">2026 retrospective benchmarks for ChatGPT, Claude, Gemini, and Grok</a>.</p>
<p>Having an assessment of usefulness and usability from the UX-Lite provides good high-level measures that can be compared to historical benchmarks. But even though it discriminates at a high level between usefulness and usability, it doesn’t provide very granular or diagnostic measures. Consequently, we’ll also want to explore more specific attitudes to see how they affect the way people think about their use of AI chatbots.</p>
<p>When measuring specific interfaces, it can be helpful to identify additional constructs, features, or interactions that participants can rate so you have a better idea of which aspects of the UX are perceived as good or bad. That gives you a more diagnostic set of items.</p>
<p>How do you do that for AI chat interfaces like ChatGPT, Claude, Gemini, and Grok? You follow the process for creating standardized measures (for example, see our <a href="https://measuringu.com/article/measuring-the-perceived-clutter-of-websites/">IJHCI paper on measuring the perceived clutter of websites</a>). You need items, data, and validation.</p>
<h2><span lang="EN-US">Picking The Items: What Matters When Interacting with Generative AI Chat Software?</span></h2>
<p>When building a standardized measure, the first step is to pick a set of items. We developed an initial set of 34 items based on input from the MeasuringU research team, drawing upon the existing literature and their experiences using these products to measure constructs like AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. Consistent with psychometric practice, we created at least three items per construct.</p>
<h3><span lang="EN-US">AI Productivity</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48078" src="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png" alt="A researcher using an AI for productivity" width="514" height="247" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-productivity-600x288.png 600w" sizes="auto, (max-width: 514px) 100vw, 514px" /></p>
<p>One of the most attractive capabilities of generative AI chatbots is the potential for enhanced productivity, making this an important construct to measure. A survey conducted by Microsoft and LinkedIn in 2024 found that <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part">90% of respondents said using AI helped them save time</a>.</p>
<p>In 2026, results from an Anthropic survey indicated <a href="https://www.anthropic.com/research/economic-index-june-2026-report">86% of respondents reported improvements in the speed</a> of their work. Table 1 shows the initial set of items we developed for this construct of increased productivity.</p>

<table id="tablepress-1055" class="tablepress tablepress-id-1055">
<thead>
<tr class="row-1">
	<th class="column-1">AI Productivity (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Using this AI chatbot greatly improves my productivity.</td>
</tr>
<tr class="row-3">
	<td class="column-1">This AI chatbot adds substantial value to my personal tasks.</td>
</tr>
<tr class="row-4">
	<td class="column-1">This AI chatbot adds substantial value to my professional tasks.</td>
</tr>
<tr class="row-5">
	<td class="column-1">Using this AI chatbot makes me feel more capable in my work or studies.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I feel comfortable being accountable for work that used this AI chatbot.</td>
</tr>
<tr class="row-7">
	<td class="column-1">Using this AI chatbot helps me achieve my goals.</td>
</tr>
<tr class="row-8">
	<td class="column-1">The amount of time it takes for this AI chatbot to respond is acceptable.</td>
</tr>
<tr class="row-9">
	<td class="column-1">This AI chatbot’s responses efficiently tell me the information I need.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1055 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>Initial item set for AI Productivity.</p>
<h3><span lang="EN-US">AI Trust</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48080" src="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png" alt="A researcher with computer screen showing &quot;AI trust&quot;" width="510" height="245" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-trust-600x288.png 600w" sizes="auto, (max-width: 510px) 100vw, 510px" /></p>
<p>The flip side of excitement about increased productivity is distrust in AI output and data security. Even recent models <a href="https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/">hallucinate</a> usability problems after reviewing videos of usability test sessions. A 2025 Melbourne-KPMG survey of 48,000 people across 47 countries found that <a href="https://fbe.unimelb.edu.au/newsroom/media-release-global-study-reveals-trust-of-ai-remains-a-critical-challenge-reflecting-tension-between-benefits-and-risks">less than half of the people regularly using AI were willing to trust it</a>. Table 2 shows our initial set of AI Trust items.</p>

<table id="tablepress-1056" class="tablepress tablepress-id-1056">
<thead>
<tr class="row-1">
	<th class="column-1">AI Trust (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I trust this AI chatbot to provide reliable information.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I feel confident relying on responses from this AI chatbot when making decisions.</td>
</tr>
<tr class="row-4">
	<td class="column-1">This AI chatbot always provides accurate responses.</td>
</tr>
<tr class="row-5">
	<td class="column-1">When this AI chatbot makes mistakes, they are usually easy to detect.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I don’t worry about how my data is used when interacting with this AI chatbot.</td>
</tr>
<tr class="row-7">
	<td class="column-1">It’s easy to understand what happens to the information I share with this AI chatbot.</td>
</tr>
<tr class="row-8">
	<td class="column-1">My professional value is not affected by products like this AI chatbot.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1056 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2: </strong>Initial item set for AI Trust.</p>
<h3><span lang="EN-US">AI Dependency</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48079" src="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png" alt="A relaxed researcher watching an AI assistant doing the work" width="518" height="249" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-dependency-600x288.png 600w" sizes="auto, (max-width: 518px) 100vw, 518px" /></p>
<p>As with Trust, there are particular concerns about AI users being overly dependent and uncritical of AI outputs. In the Melbourne-KPMG survey, 66% of respondents reported relying on AI output without evaluating its accuracy, and <a href="https://assets.kpmg.com/content/dam/kpmgsites/xx/pdf/2025/05/trust-attitudes-and-use-of-ai-global-report.pdf">56% reported making mistakes in their work due to uncritical acceptance of an AI output</a>. See Table 3 for the AI Dependency items.</p>

<table id="tablepress-1057" class="tablepress tablepress-id-1057">
<thead>
<tr class="row-1">
	<th class="column-1">AI Dependency (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I often rely on AI chatbots to perform tasks that I would otherwise do myself.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I tend to accept answers from AI chatbots without verifying their accuracy.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I rarely double-check information provided by AI chatbots.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1057 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 3: </strong>Initial item set for AI Dependency.</p>
<h3><span lang="EN-US">AI Anxiety</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48081" src="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png" alt="Stressed researcher looking at a line chart" width="512" height="246" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-anxiety-600x288.png 600w" sizes="auto, (max-width: 512px) 100vw, 512px" /></p>
<p>Conversations about AI often turn to anxiety about the potential negative effects of generative AI chatbots on society, the environment, and ethics. A 2025 University of Chicago AP-NORC survey found <a href="https://epic.uchicago.edu/wp-content/uploads/sites/5/2025/10/EPIC-AP-NORC-Poll_AI_2025_Fact-Sheet.pdf">44% believed AI would do more to hurt than help society</a>, compared with 22% who expected it to do more good, and 41% were extremely or very concerned about AI’s environmental impact. Table 4 shows our initial item set for AI Anxiety.</p>

<table id="tablepress-1058" class="tablepress tablepress-id-1058">
<thead>
<tr class="row-1">
	<th class="column-1">AI Anxiety (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">The increasing use of AI makes me uneasy.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am often concerned that AI could cause serious harm to society.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I often worry about the environmental impact of AI.</td>
</tr>
<tr class="row-5">
	<td class="column-1">AI development feels risky.</td>
</tr>
<tr class="row-6">
	<td class="column-1">AI development feels difficult to control.</td>
</tr>
<tr class="row-7">
	<td class="column-1">There should be more government regulation for AI development.</td>
</tr>
<tr class="row-8">
	<td class="column-1">Using AI chatbots for work or school feels unethical.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1058 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 4: </strong>Initial item set for AI Anxiety.</p>
<h3><span lang="EN-US">AI Personification</span></h3>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-48082" src="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png" alt="A researcher interacting with an AI in the form of an angelic woman" width="507" height="244" srcset="https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1024x492.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-300x144.png 300w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-768x369.png 768w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-1536x738.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-2048x984.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/ai-personification-600x288.png 600w" sizes="auto, (max-width: 507px) 100vw, 507px" /></p>
<p>Another aspect of interaction with generative AI chatbots that interests us is the extent to which users feel a personal relationship with the AI. This could range from feeling like you&#8217;re communicating with a human-like entity to feelings of friendship. <a href="https://www.nature.com/articles/s41598-025-19212-2">Not everyone has the same emotional reaction to generative AI chatbot products</a>, but we are interested in how products may differ in the extent to which they lead to social connection with their users. Our initial set of AI Personification items is listed in Table 5.</p>

<table id="tablepress-1059" class="tablepress tablepress-id-1059">
<thead>
<tr class="row-1">
	<th class="column-1">AI Personification (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">Interacting with this AI chatbot feels like communicating with a human.</td>
</tr>
<tr class="row-3">
	<td class="column-1">Sometimes I feel like this AI chatbot is more like a friend than a tool.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I feel like AI chatbots understand me well.</td>
</tr>
<tr class="row-5">
	<td class="column-1">I tend to feel a sense of connection when interacting with AI chatbots.</td>
</tr>
<tr class="row-6">
	<td class="column-1">I tend to feel like I’m socializing when I interact with AI chatbots.</td>
</tr>
<tr class="row-7">
	<td class="column-1">I’m more likely to share personal information with AI chatbots than with other people.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1059 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 5: </strong>Initial item set for AI Personification.</p>
<h3><span lang="EN-US">Early Adoption</span></h3>
<p>Although not related solely to AI, we included three items to assess respondents’ tendencies to be early adopters of new technologies (Table 6).</p>

<table id="tablepress-1060" class="tablepress tablepress-id-1060">
<thead>
<tr class="row-1">
	<th class="column-1">Early Adoption (Initial Set)</th>
</tr>
</thead>
<tbody class="row-striping row-hover">
<tr class="row-2">
	<td class="column-1">I like to experiment with new technologies before most people do.</td>
</tr>
<tr class="row-3">
	<td class="column-1">I am usually among the first to try new digital tools.</td>
</tr>
<tr class="row-4">
	<td class="column-1">I actively seek out new technologies to try.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1060 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 6: </strong>Initial item set for Early Adoption.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>In this article, we discussed how to measure the UX of AI in general and, specifically, generative AI chatbots like ChatGPT, Claude, Gemini, or Grok. The key points in the article were:</p>
<p><strong>Measuring AI UX Starts with Measuring UX. </strong>Measuring UX in general means using a mix of attitudinal and action metrics. Action metrics are more rigorous but harder to collect—you need task scenarios and observation. Attitudinal metrics are easier to collect and often take the form of standardized questionnaires. Standardized metrics like the UX-Lite can discriminate among products based on their perceived usefulness and usability and are a good place to start. But if you want more diagnostic insight—to know more than just that <em>something</em> is off—you need a deeper, more specialized set of measures.</p>
<p><strong>What Constructs Comprise the UX of AI? </strong>A deeper dive into the UX of generative AI chatbots requires investigation of specialized constructs. Based on our reading and experience with these types of products, we&#8217;ve proposed items for measuring AI Productivity, AI Trust, AI Dependency, AI Anxiety, and AI Personification.</p>
<p><strong>To Validate Items, You Need Data from Real People. </strong>Creating an initial set of items is an important first step to develop standardized metrics, but it is just a first step. Items that look sensible on paper don&#8217;t always hold up once real people respond to them. In future articles, we&#8217;ll report the results of psychometric evaluation to (1) determine if the initial items, as we expect, group into statistical factors, (2) examine item quality to determine which items to retain for a final streamlined instrument, and (3) explore the connection between the new constructs and higher-level constructs like brand attitude, intention to continue use, and intention to recommend. Stay tuned!<strong><br />
</strong></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>UX and NPS Benchmarks of Health Insurance Websites (2026)</title>
		<link>https://measuringu.com/ux-nps-benchmark-report-for-healthcare-websites-2026/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=ux-nps-benchmark-report-for-healthcare-websites-2026</link>
		
		<dc:creator><![CDATA[Eva Sundberg, MPH&nbsp;•&nbsp;Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 21 Jul 2026 20:52:06 +0000</pubDate>
				<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Healthcare]]></category>
		<category><![CDATA[NPS]]></category>
		<category><![CDATA[SUPR-Q]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48005</guid>

					<description><![CDATA[Tech has improved a lot over the last ten years. It’s hard to beat the promised convenience of managing health insurance online. Gone are the days of hunting through paper policy booklets, waiting on an employer’s benefits manager, or enduring the endless hold music of a telephone call center. Modern portals promise complete, on-demand control [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat.png"><img loading="lazy" decoding="async" class="alignleft wp-image-48043 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-300x169.png" alt="Feature image showing healthcare workers" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-600x338.png 600w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat.png 1280w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>Tech has improved a lot over the last ten years.</p>
<p>It’s hard to beat the <em>promised</em> convenience of managing health insurance online. Gone are the days of hunting through paper policy booklets, waiting on an employer’s benefits manager, or enduring the endless hold music of a telephone call center.</p>
<p>Modern portals promise complete, on-demand control over your healthcare: real-time network coverage and deductible tracking, a doctor search that works like any other search bar, and a full claims history available any time, day or night.</p>
<p>But this digital autonomy comes with high-stakes drawbacks. Because health insurance portals must organize fragmented clinical, network, and financial data, they frequently lack the intuitive ease that consumers expect from modern digital interfaces. It’s still hard to find out if a doctor is in your insurance network or if parts of your claim won&#8217;t be covered.</p>
<p>The circumstances behind the visit can also affect usability. Unlike low-stakes commercial retail websites, where users can casually learn navigational patterns over time, a health insurance portal is rarely visited for leisure. A user&#8217;s first interaction often occurs during a high-stress medical event where cognitive load is already strained, and a clunky, unfamiliar layout becomes a barrier to care.</p>
<p>Despite these issues, digital engagement is at an all-time high, with independent market data revealing that <a href="https://www.flexpa.com/blog/digital-by-default">85.5% of insured Americans have created an online account with their health insurance provider</a>. However, because of foundational usability flaws, the baseline experience fails consumers when they need help the most. Recent industry benchmarks show that <a href="https://www.jdpower.com/business/u-s-healthcare-digital-experience-study/">health plans fail to deliver an easy-to-use digital experience 39% of the time</a>, and their consumer satisfaction scores drastically trail other digital sectors like retail banking and property insurance. This makes the quantitative improvement of health insurance websites&#8217; user experience a critical priority for both insurers and patients.</p>
<p>This isn’t the first time we’ve checked in on this industry. When we <a href="https://measuringu.com/ux-health-insurance/">benchmarked six major health insurance websites in 2018</a>, the group’s average SUPR-Q score sat at the 67th percentile—solidly above average. Eight years, and presumably a lot of digital investment, later? The group founders at the 30th percentile.</p>
<p>To understand the quality of their online experiences today, we collected UX benchmark metrics on eight common healthcare websites and mobile applications.</p>
<ul>
<li>Aetna (CVS Health)</li>
<li>Blue Cross Blue Shield</li>
<li>Cigna</li>
<li>Elevance Health (Anthem)</li>
<li>Highmark</li>
<li>Humana</li>
<li>Kaiser Permanente</li>
<li>UnitedHealthcare</li>
</ul>
<p>We computed <a href="https://measuringu.com/product/suprq/">SUPR-Q</a><sup>®</sup> and <a href="https://measuringu.com/nps-reliability/">Net Promoter</a> scores, measured users’ attitudes regarding their experiences, conducted <a href="https://measuringu.com/key-drivers/">key driver</a> analyses, and analyzed reported usability problems. (Full details are in the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>.)</p>
<h2>Benchmark Study Details</h2>
<p>From February to March 2026, we asked 391 users of healthcare websites in the U.S. to recall their most recent experience and perceptions of one of these websites on their desktop and mobile app (if applicable) in the past year.</p>
<p>Respondents completed the eight-item <a href="https://measuringu.com/10-things-suprq/">SUPR-Q</a> (which includes the <a href="https://measuringu.com/nps-ux/">Net Promoter Score</a>), the two-item <a href="https://measuringu.com/evolution-of-the-ux-lite/">UX-Lite</a><sup>®</sup>, and the <a href="https://measuringu.com/article/streamlining-the-supr-qm-the-supr-qm-v2/">SUPR-Qm</a> standardized questionnaires, and answered questions about their brand attitudes, usage, and prior experiences.</p>
<h2>Quality of the Website User Experience: SUPR‑Q</h2>
<p>The SUPR-Q is a standardized questionnaire widely used for measuring attitudes toward the quality of a website user experience. Its norms are computed from a rolling database of around 200 websites across dozens of industries.</p>
<p>SUPR-Q scores are percentile ranks that tell you how a website’s experience ranks relative to other websites (50<sup>th</sup> percentile is average). The SUPR-Q provides an overall score as well as detailed scores for subdimensions of Usability, Trust, Appearance, and Loyalty.</p>
<p>The mean SUPR-Q across healthcare websites in this study was at the 30<sup>th</sup> percentile (substantially below average), ranging from the 12<sup>th</sup> percentile for Highmark to the 49<sup>th</sup> percentile for Cigna.</p>
<h3>Usability Scores</h3>
<p>Overall, usability scores were also well below average for the healthcare websites, averaging at the 21<sup>st</sup> percentile. Highmark had the lowest usability score at the 7<sup>th</sup> percentile, and Cigna had the highest (57<sup>th</sup> percentile).</p>
<p>Comments related to usability on Highmark included:</p>
<p style="padding-left: 25px;"><em>“Certain pages require multiple clicks to get to basic information, and the navigation isn’t always intuitive, which can make the experience a bit frustrating.”</em></p>
<p style="padding-left: 25px;"><em>“The layout is confusing. Too much time looking for what I need.”</em></p>
<h3>Loyalty/Net Promoter Scores</h3>
<p>By itself, the Net Promoter Scores isn’t necessarily a good metric for health insurance websites, since many users don&#8217;t get to choose their insurer (and the site they have to use). A low score could reflect more on the industry, a limited network, or a mandated plan than it does on the user experience. Despite this potential confusion, we still report it for two reasons. First, comparing scores across companies in the same industry (and over time periods) helps identify who&#8217;s delivering a relatively better or worse experience, as the underlying constraint is the same for everyone. Second, NPS remains one of the few metrics that needs no explanation in a boardroom: stakeholders already know what a score of 40% versus −20% means, which makes it a useful entry point when discussing more relevant UX metrics associated with the NPS.</p>
<p>All but two healthcare websites, Humana and Aetna (CVS Health), had negative Net Promoter Scores, led by Humana at 14% and followed by Aetna (CVS Health) at a neutral 0%. Highmark scored the worst behind the group at−47%. The average NPS for these websites was −14%, indicating a substantial surplus of brand detractors over promoters across the industry.</p>
<p>Comments related to NPS and loyalty included:</p>
<p style="padding-left: 25px;"><em>“For a website that has a variety of audiences visiting it who have very different needs, it does a good job and is more simple than having to navigate to a website specific to who you are (employer versus provider versus patient, etc.).” </em>— Blue Cross Blue Shield</p>
<p style="padding-left: 25px;"><em>“I just don&#8217;t think they are the best.  I am not sure the information is reliable especially when trying to find a provider or a specialist.” </em>— Highmark</p>
<h2>Websites and Mobile App Usage</h2>
<p>As a part of this benchmark, we asked respondents how they accessed the healthcare providers online. All respondents reported using their desktop/laptop computers (this was a requirement for participation in the survey), with 67% also using mobile apps and 69% using mobile websites. Most respondents reported visiting their healthcare websites on a desktop or a laptop computer a few times per year. Mobile app users showed more variation, with 40% of Aetna (CVS Health) and 39% of Kaiser Permanente users reporting using mobile apps a few times per month, and 33% of Cigna users reporting using mobile apps a few times per year. The majority of other healthcare website users mostly reported having never used the mobile app.</p>
<h2>Key Drivers of UX Quality</h2>
<p>To better understand what affects SUPR-Q scores and Likelihood-to-Recommend (LTR) ratings, we asked respondents to rate potentially important attributes of the healthcare websites on a five-point scale from 1 (Strongly disagree) to 5 (Strongly agree). We conducted key driver analyses (regression modeling) to quantify the extent to which ratings on these items drive (account for) variation in overall SUPR-Q scores and, separately, LTR (the rating from which the NPS is derived; full details are in the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>).</p>
<p>The top key driver of the overall SUPR-Q scores was the <strong>ease of finding providers</strong> (13%), followed by the abilities to find accurate claims information (9%), quickly find information on prior claims (9%), and easily find what they will pay for services (8%). Taken together, eight significant variables accounted for 63% of the variance in SUPR-Q scores.</p>
<p>For likelihood to recommend (LTR), the top key driver was the ease of finding providers (9%), followed by trust that personal and health information are secure (9%). Tying for third are plan coverage being easy to understand (7%) and the ease of finding how much services will cost (7%). Overall, six significant drivers accounted for 46% of the variance in LTR.</p>
<p>Figure 1 shows a scatterplot of the importance and opportunity for improvement for seven key drivers. The combination of importance and opportunity for improvement provides a basis for prioritizing which key drivers to improve. The importance score is the greater of the variance accounted for by the driver in the SUPR-Q and NPS analyses, where larger percentages indicate more importance. The opportunity score is the top-box percentage for the driver, so smaller percentages indicate greater opportunity for improvement (e.g., it would be harder to improve a driver with a top-box percentage of 100% than one with a top-box percentage of 10%).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-48035 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1024x345.jpg" alt="Scatterplot of importance and opportunity for improvement of key drivers." width="1024" height="345" srcset="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1024x345.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-300x101.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-768x259.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1536x518.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-2048x691.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-600x202.jpg 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Scatterplot of importance and opportunity for improvement of key drivers.</p>
<p>Five of these seven key drivers fell in the FIX quadrant (upper left) with relatively high importance and higher opportunity for improvement:</p>
<ul>
<li>Easy to find providers</li>
<li>Personal/health information is secure</li>
<li>Quick to find information on prior claims</li>
<li>Easy to find what I will pay for services</li>
<li>Accurate claims information</li>
</ul>
<p>Blue Cross Blue Shield achieved the highest top-box scores for quickly finding information about prior claims (46%), easily finding covered doctors or providers (32%), and consumer confidence that personal and health information is secure (30%). Humana led for trust in the accuracy of online claims information (24%), while Blue Cross Blue Shield and Cigna tied for the highest top-box score regarding the ease of finding out-of-pocket cost information (18%).</p>
<p>Conversely, the websites with the lowest top-box scores, suggesting the most room for digital improvement, were Humana for tracking prior claims (12%), Highmark for data security confidence (12%), and Humana again for finding covered providers (14%). Highmark scored the lowest for trust in claims information accuracy (9%), while Aetna (CVS Health) registered the worst top-box performance for clear out-of-pocket cost transparency (10%).</p>
<h2>UX Problems</h2>
<p>We examined the verbatim comments to better understand the problems users had.</p>
<h3>Difficulty Finding Information Was a Universal Grievance</h3>
<p>This issue affected all eight websites and was the single top complaint across every provider in the study. It was a massive pain point for Kaiser Permanente, Highmark, and UnitedHealthcare.</p>
<p style="padding-left: 25px;"><em>“Its navigation is non-intuitive.” </em>— Kaiser Permanente</p>
<p style="padding-left: 25px;"><em>“It is very difficult to find the information I need.” </em>— UnitedHealthcare</p>
<p style="padding-left: 25px;"><em>“I stumble through a lot of different pages before finally finding what I’m looking for.” </em>— Blue Cross Blue Shield</p>
<p>These comments speak to a navigation problem, not a content problem. The needed information is usually on the site <em>somewhere</em>. It’s just buried under menu structures built around the insurer’s org chart instead of the handful of things a member actually comes to do.</p>
<h3>Unreliable Provider Directories Degrade the Experience</h3>
<p>For seven of the eight platforms, outdated or inaccurate directory information emerged as a major barrier to care. Users routinely noted a stark disconnect between the doctors listed online and the reality of who actually accepts their insurance.</p>
<p style="padding-left: 25px;"><em>“They don’t update doctors who are in network and who isn’t.” </em>— Aetna</p>
<p style="padding-left: 25px;"><em>“It is not always accurate on whether a place/doctor takes your insurance or not.” </em>— Blue Cross Blue Shield</p>
<p style="padding-left: 25px;"><em>“Caregivers are listed as being in network or accepting new patients, but when I call to make an appointment, that&#8217;s not the case.” </em>— UnitedHealthcare</p>
<p>Figure 2 shows a critical directory system error on the Blue Cross Blue Shield search results page. Instead of presenting a diverse list of nearby options, the portal repeats the exact same facility multiple times in a row down the page. This type of technical redundancy unnecessarily inflates the layout, clutters the interface, and requires members to filter through duplicate data entries to locate distinct, alternative care providers.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48008" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png" alt="Duplicate provider search results on the Blue Cross Blue Shield platform." width="624" height="384" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F2-300x185.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F2-600x369.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Duplicate provider search results on the Blue Cross Blue Shield platform.</p>
<h3>Sluggish Performance and Site Errors Delay Essential Tasks</h3>
<p>Users reported slow loading times, timeouts, or complete site outages across all eight websites. Performance issues were especially disruptive for Highmark, Elevance Health (Anthem), and Kaiser Permanente.</p>
<p style="padding-left: 25px;"><em>“Sometimes when I log in, the site is down.” </em>— Aetna</p>
<p style="padding-left: 25px;"><em>“It&#8217;s clunky, it freezes, it&#8217;s under construction, constant error messages.” </em>— Humana</p>
<p style="padding-left: 25px;"><em>“Sometimes there are error bugs on certain pages although rare. When these different bugs pop up certain information or pages are not accessible.” — Elevance Health (Anthem)</em></p>
<p style="padding-left: 25px;"><em>“Occasional slow page loads or timeouts especially during peak hours.”</em> — UnitedHealthcare</p>
<p style="padding-left: 25px;"><em>“The website is very slow at times and pages fail to load altogether. Getting a generic error message when trying to access my medical records is extremely frustrating.” </em>— Kaiser Permanente</p>
<p>Figure 3 illustrates the exact technical dead end described by these participants. Instead of loading the requested medical records or dashboard utility, the Kaiser Permanente platform encounters a critical error loop, leaving the page blank save for a red exclamation warning badge and a disruptive &#8220;Back to Sign In&#8221; command. This system failure completely breaks the user session, forcing members to entirely restart the multi-step login process to attempt their task again.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48009" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png" alt="A systemic backend routing error on the Kaiser Permanente portal that blocks member access to internal records and triggers session termination." width="624" height="343" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F3-300x165.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F3-600x330.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> A systemic backend routing error on the Kaiser Permanente portal that blocks member access to internal records and triggers session termination.</p>
<p>These aren’t cosmetic bugs. A frozen page or a timeout in the middle of a claims dispute can mean a missed deadline, not just an annoyance. Every one of the eight sites drew reliability complaints, suggesting that basic uptime and performance are still struggles for an industry pushing members toward digital-only service.</p>
<h3>Cluttered Layouts and Irrelevant Popups Prevent Smooth Navigation</h3>
<p>Navigational bottlenecks heavily impacted users’ baseline experience. Six of the eight sites suffered from significant login friction, while others overwhelmed users with busy, crowded interfaces.</p>
<p style="padding-left: 25px;"><em>“Overwhelming amount of information … the website sometimes feels cluttered with information, making it difficult for me to filter out what doesn&#8217;t apply to me.”</em> — Humana</p>
<p>Figure 4 shows why: the interface presents an overly dense, text-heavy dashboard layout for accessing basic member documents. Rather than organizing resources into clear, scannable visual segments, the platform confronts users with exhaustive walls of prose, complex text blocks outlining compliance requirements, and generic sidebar links. This format spikes cognitive load, forcing members to meticulously read through paragraphs of operational data just to find a specific form link.<a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48010" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png" alt="Dense, text-heavy documentation layouts on the Humana platform that contribute to visual clutter and search fatigue." width="936" height="610" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png 936w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-300x196.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-768x501.png 768w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-600x391.png 600w" sizes="auto, (max-width: 936px) 100vw, 936px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> Dense, text-heavy documentation layouts on the Humana platform that contribute to visual clutter and search fatigue.</p>
<p>This layout fatigue is far from an isolated issue; members across other major insurance platforms frequently reported navigating highly saturated interfaces where non-essential material routinely buries functional tools:</p>
<p style="padding-left: 25px;"><em>“It can be confusing at times due to the website being crowded with information.” </em>— Kaiser Permanente</p>
<p style="padding-left: 25px;"><em>“There is a lot of fluff and links to articles that I don&#8217;t care about or have no bearing on the services provided.” </em>— Elevance Health (Anthem)</p>
<p>Figure 5 shows another source of that clutter: a promotional pop-up covering the entire &#8220;Find Care&#8221; page on UnitedHealthcare, which members have to dismiss before they can do anything else. A marketing message is blocking the one task this page exists for.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48011" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png" alt="An intrusive modal pop-up on the UnitedHealthcare website." width="624" height="379" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F5-300x182.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F5-600x364.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 5:</strong> An intrusive modal pop-up on the UnitedHealthcare website.</p>
<h2>Summary and Takeaways</h2>
<p>Healthcare payers are massive enterprises, serving millions of Americans who increasingly rely on digital self-service tools. An analysis of the user experience of eight major health insurance platforms using benchmark data found:</p>
<ol>
<li><strong>Cigna and Humana lead; Highmark lags.</strong> Cigna achieved the highest overall SUPR-Q score, falling at the 49<sup>th</sup> percentile, while Highmark had the lowest score, falling in the 12<sup>th</sup> percentile. Humana achieved the highest Net Promoter Score at 14%, while Highmark fell the furthest behind the group at −47%.</li>
</ol>
<ol start="2">
<li><strong>Provider verification and claims transparency drive UX scores.</strong> The top key driver of overall SUPR-Q scores was the ease of finding doctors or providers covered by the user’s plan (13%), followed by the accuracy of claims information (9%), quickly finding info on prior claims (9%), and and easily finding what they will pay for services (8%). Taken together, eight significant variables accounted for 63% of the variance in the SUPR-Q scores. Other critical drivers include the ability to get information without calling customer service (6%) and confidence that personal/health information is secure (5%).</li>
</ol>
<ol start="3">
<li><strong>The top opportunities for improvement are helping members easily find covered providers and allowing them to quickly track prior claims. </strong>One way to prioritize attention to key drivers is to consider both their importance (variability in regression) and how well the websites achieve the stated goal (top-box scores). The two key drivers with the most potential for improvement (high impact percentages and low top-box scores) were the ability to easily find doctors or providers covered by a plan (average top-box score: 21%) and the ability to quickly find information about prior claims (average top-box score: 22%). For finding covered providers and for tracking prior claims, the leader was Blue Cross Blue Shield (Humana lagging).</li>
</ol>
<ol start="4">
<li><strong>Top frustrations are hidden information, website performance issues, and unreliable provider searches.</strong> The most frequently reported issue, affecting all eight websites, was a general difficulty in locating information. Website performance and slow loading times were also universal issues logged across all eight platforms. For seven of the eight systems, users specifically cited that the provider search directory was unreliable or difficult to parse. General login and authentication friction was also highly prevalent, acting as a major pain point on six of the platforms.</li>
</ol>
<ol start="5">
<li><strong>Health insurance members have trouble finding vital information.</strong> Quantitative and qualitative signals from our findings converge on a clear trend: users of these digital health platforms have trouble efficiently uncovering vital policy information. Quantitative signals emphasize broken pathways toward verifying provider networks, reviewing claims, and estimating out-of-pocket pricing. The qualitative &#8220;why&#8221; behind the numbers highlights extensive complaints surrounding buried information, clunky login processes, and sluggish site performance. Quantitative metrics and user commentary both suggest that the platforms that would especially benefit from a strong focus on baseline usability are Highmark, UnitedHealthcare, and Elevance Health.</li>
</ol>
<p>For more details, see the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Understanding Alpha Inflation</title>
		<link>https://measuringu.com/understanding-alpha-inflation/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=understanding-alpha-inflation</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 14 Jul 2026 22:23:34 +0000</pubDate>
				<category><![CDATA[Statistics]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[alpha inflation]]></category>
		<category><![CDATA[false positive]]></category>
		<category><![CDATA[Null Hypothesis]]></category>
		<category><![CDATA[Type 1 error]]></category>
		<category><![CDATA[Type I error]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47974</guid>

					<description><![CDATA[The large sample size. The right statistical tests. A compelling and statistically significant finding! All the ingredients of a successful quantitative project that gets stakeholder buy-in. But inflation isn’t just a worry for the Federal Reserve on interest rate policy. It affects research decisions, too. When you conduct something like a usage and attitude survey [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47987 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-300x169.jpg" alt="Feature image showing a researcher inflating a balloon with alpha written on it" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>The large sample size. The right statistical tests. A compelling and statistically significant finding! All the ingredients of a successful quantitative project that gets stakeholder buy-in. But inflation isn’t just a worry for the Federal Reserve on interest rate policy. It affects research decisions, too.</p>
<p>When you conduct something like a usage and attitude survey with hundreds or thousands of participants (like what we do with our <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">industry reports</a>), you have the benefit of statistical power (the ability to detect differences). This allows you to look not only for differences between main effects (like between software products), but also for interactions (differences by type of user by product).</p>
<p>With enough power, you can dig into these more nuanced differences. For example, you may learn that younger adults have different usage patterns and attitudes than older cohorts for certain products. It would be a compelling story, and it&#8217;s something that stakeholders can make sense of and act on. We’ve seen it and reported it. But how do you know these differences aren’t a fluke?</p>
<p>You may think limiting yourself to reporting only statistically significant findings will protect you from these false positives, and that approach certainly helps. But like printing money to boost the economy, the more statistical tests you run, the more you inflate your chances of finding things that aren’t actually there.</p>
<p>This isn’t monetary policy; it’s alpha inflation. It’s a subtle concept that you should understand when making multiple statistical comparisons.</p>
<h2><span lang="EN-US">Alpha (False Alarm Rate) in Hypothesis Testing</span></h2>
<p>In <a href="https://measuringu.com/how-does-hypothesis-testing-work/">null hypothesis significance testing</a> (NHST), the alpha criterion is the value selected for decisions of statistical significance. This is the process through which you decide if a difference is statistically significant. As shown in Figure 1, when the <em>p</em>-value you get from a statistical test is less than the alpha value you set, you reject the hypothesis of no difference (the null hypothesis, H<sub>0</sub>). It’s statistical significance! Otherwise, you fail to reject H<sub>0</sub> (you don’t accept it; you just don’t have enough evidence to reject it).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-47983 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-1024x352.jpg" alt="Figure 1: High-level flowchart of statistical hypothesis testing." width="1024" height="352" srcset="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-1024x352.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-300x103.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-768x264.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-600x206.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1.jpg 1461w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> High-level flowchart of statistical hypothesis testing.</p>
<p>The whole point of this process is to control the percentage of false alarms (rejecting H<sub>0</sub> when there really is no difference) <strong>in the long run</strong>. We’ve previously discussed why <a href="https://measuringu.com/setting-alpha/">the alpha criterion doesn’t have to be the standard <em>p</em> &lt; .05</a>, but the rationale for using .05 was published by R. A. Fisher in 1929: “It is a common practice to judge a result significant, if it is of such a magnitude that it would have been produced by chance not more frequently than once in twenty trials [.05]. This is an arbitrary, but convenient, level of significance for the practical investigator.”</p>
<p>So, when you run one test of significance, if there is really no difference, you have a 5% chance of mistakenly concluding that there is a difference. That’s fine if you’ve collected data from two groups (A and B), but what if you’ve collected data from three groups (A, B, and C) and want to compare A with B, B with C, and A with C? Or, as in the case of the usage and attitude survey, you want to make dozens of comparisons of subgroups (such as age or usage levels across product versions)?</p>
<p>For all these scenarios, you’ll have to deal with the consequences of alpha inflation.</p>
<h2><span lang="EN-US">What Is Alpha Inflation?</span></h2>
<p>When we declare a difference as statistically significant when <em>p</em> &lt; .05, we reduce our chance of being fooled by random noise in our sample. But these false alarms, called <a href="https://measuringu.com/hypothesis-testing-what-can-go-wrong/">Type I</a> errors in statistical jargon, will happen and are expected.</p>
<p>At the <em>p</em> &lt; .05 level of significance, 5% of our conclusions will be false alarms over the long run. That’s 5% for running just one statistical test (such as comparing the <a href="https://measuringu.com/sample-sizes-for-comparison-of-ux-lite-scores/">UX-Lite</a><sup>®</sup> scores of two products). But what if you run more than one statistical test? Software makes it very easy to run all sorts of comparisons such as the differences between age cohorts, experience levels, and products. What if you run 5, 10, 20, or 100 tests? You’re printing money and, like the inflation rate, your false alarm rate goes up too.</p>
<p>To understand how running more tests increases your false alarm (Type I) error rate, Table 1 shows the probability of zero false alarms, one false alarm, two false alarms, and so forth.</p>

<table id="tablepress-1054" class="tablepress tablepress-id-1054">
<thead>
<tr class="row-1">
	<th class="column-1">x (Number of false alarms)</th><th class="column-2"><i>p</i>(x) when <i>α<i> = .05</th><th class="column-3"><i>p</i>(at least x) when <i>α<i> = .05</th>
</tr>
</thead>
<tbody class="row-striping">
<tr class="row-2">
	<td class="column-1">0</td><td class="column-2">0.36</td><td class="column-3">1.00</td>
</tr>
<tr class="row-3">
	<td class="column-1">1</td><td class="column-2">0.38</td><td class="column-3">0.64</td>
</tr>
<tr class="row-4">
	<td class="column-1">2</td><td class="column-2">0.19</td><td class="column-3">0.26</td>
</tr>
<tr class="row-5">
	<td class="column-1">3</td><td class="column-2">0.06</td><td class="column-3">0.08</td>
</tr>
<tr class="row-6">
	<td class="column-1">4</td><td class="column-2">0.01</td><td class="column-3">0.02</td>
</tr>
<tr class="row-7">
	<td class="column-1">5</td><td class="column-2">0.002</td><td class="column-3">0.003</td>
</tr>
<tr class="row-8">
	<td class="column-1">6</td><td class="column-2">0.0003</td><td class="column-3">0.0003</td>
</tr>
<tr class="row-9">
	<td class="column-1">7</td><td class="column-2">0.00003</td><td class="column-3">0.00003</td>
</tr>
<tr class="row-10">
	<td class="column-1">8–20</td><td class="column-2">0.00000</td><td class="column-3">0.00000</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1054 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Alpha inflation for 20 tests conducted with <em>α</em> = .05.</p>
<p>As shown in Table 1, the most likely number of Type I errors in a set of 20 independent tests with <em>α</em> = 0.05 is one, with a point probability of 0.38. That is, when you run 20 comparisons, there’s a <strong>38% chance of getting one false alarm</strong>. The next highest point probability is 0.36 for 0 false alarms, meaning you have about the same chance of having no false alarms in a set of 20 comparisons as you do one false alarm.</p>
<p>Unfortunately, you can encounter more than one false alarm in a set of 20 tests. The likelihood of at least one Type I error, however, is higher (specifically, 1 − <em>p</em>(0) = 1 − 0.36 = 0.64). That’s a 64% chance of 1, 2, 3, or more. (Although more than two or three false alarms in 20 comparisons is rare.) So, rather than having a 5% chance of encountering a Type I error when there is no real difference, <em>α</em> has inflated to 64%.</p>
<h2><span lang="EN-US">What Can You Do About Alpha Inflation?</span></h2>
<p>You can’t raise interest rates to tame alpha inflation. But since the middle of the 20th century, many strategies and techniques have been published to guide the analysis of multiple comparisons, such as omnibus tests (e.g., <a href="https://en.wikipedia.org/wiki/Analysis_of_variance">ANOVA</a> and <a href="https://en.wikipedia.org/wiki/Multivariate_analysis_of_variance">MANOVA</a>) and procedures for the comparison of pairs of means (e.g., Tukey’s <a href="https://en.wikipedia.org/wiki/Tukey's_B_method">WSD</a> and <a href="https://en.wikipedia.org/wiki/Tukey's_range_test">HSD</a> procedures, the <a href="https://en.wikipedia.org/wiki/Newman%E2%80%93Keuls_method">Student–Newman–Keuls test</a>, <a href="https://en.wikipedia.org/wiki/Dunnett%27s_test">Dunnett’s test</a>, the <a href="https://en.wikipedia.org/wiki/Duncan%27s_new_multiple_range_test">Duncan procedure</a>, the <a href="https://en.wikipedia.org/wiki/Scheff%C3%A9%27s_method">Scheffé procedure</a>, the <a href="https://en.wikipedia.org/wiki/Bonferroni_correction">Bonferroni adjustment</a>, and the <a href="https://en.wikipedia.org/wiki/False_discovery_rate#Benjamini%E2%80%93Hochberg_procedure">Benjamini–Hochberg adjustment</a>).</p>
<p>With all these methods available to handle alpha inflation, the problem is solved—right?</p>
<h2><span lang="EN-US">Controlling Alpha Inflation Has Its Consequences</span></h2>
<p>When the null hypothesis is not true, applying techniques to control alpha inflation (decreasing the number of Type I errors) necessarily increases the number of Type II errors—the failure to detect differences that are real. An overemphasis on the prevention of Type I errors leads to the proliferation of Type II errors.</p>
<p>Unless, for your situation, the cost of a Type I error is much greater than the cost of a Type II error, you should avoid applying any of the techniques designed to suppress alpha inflation. As Perneger (<a href="https://jcsites.juniata.edu/faculty/merovich/QuantEcol_files/ESS335Perneger1998_Bonferroni_adjustments.pdf">1998</a>, p. 1236) wrote, “Simply describing what tests of significance have been performed, and why, is generally the best way of dealing with multiple comparisons.” In UX and many applied research settings, failing to detect a real difference can be just as harmful as falsely declaring a difference. Context matters, so don’t think Type I errors should take all the focus.</p>
<p>Or, as Abelson (<a href="https://www.google.com/books/edition/Statistics_as_Principled_Argument/BHiNEQAAQBAJ?hl=en&amp;gbpv=1&amp;pg=PT8&amp;printsec=frontcover">1995</a>, p. 70) put it, “Random patterns will seem to contain something systematic when scrutinized in many particular ways. If you look at enough boulders, there is bound to be one that looks like a sculpted human face. Knowing this, if you apply extremely strict criteria for what is to be recognized as an intentionally carved face, you might miss the whole show on Easter Island.”</p>
<p>We pay more attention to alpha inflation when we’re making many unplanned comparisons. But even then, we have a balanced approach to managing Type I and Type II errors, which we’ll cover in an upcoming article.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Alpha inflation is what happens when running multiple statistical tests quietly erodes the reliability of your <em>p</em> &lt; .05 threshold, the same way printing money erodes the value of a dollar. The main points from this article are:</p>
<p><strong>Alpha inflation is real. </strong>Run 20 independent comparisons at <em>α</em> = .05, and you don&#8217;t have a 5% chance of a false alarm anymore; you have a 64% chance of <em>at least</em> one. The reality of alpha inflation can easily be demonstrated using binomial probabilities (as in Table 1).</p>
<p><strong>Many methods have been developed to control alpha inflation. </strong>The methods, primarily developed in the 20th century, vary considerably in their relative conservatism (judging fewer contrasts to be significant) and liberalism (judging more contrasts to be significant).</p>
<p><strong>But you can&#8217;t just raise interest rates to fix alpha inflation. </strong>Controlling alpha inflation has hidden costs. Controlling only Type I errors (false alarms) leads to a proliferation of Type II errors (misses). The decision to use methods to control alpha inflation depends strongly on the relative costs of Type I and Type II errors in a specific research context. We&#8217;ll cover how to strike that balance in an upcoming article.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>UX Benchmarks for AI-Based Chat Software (2026)</title>
		<link>https://measuringu.com/ai-based-chat-software-ux-2026/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=ai-based-chat-software-ux-2026</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 07 Jul 2026 22:28:13 +0000</pubDate>
				<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47842</guid>

					<description><![CDATA[AI may be rapidly increasing the speed at which we can generate software and find answers, but where there&#8217;s an interface for a human, interactions aren&#8217;t always smooth. As AI providers compete for users, they add features. Although features add capabilities, they can also add complexity. Simple chat interfaces are giving way to tabbed experiences [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47943 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-300x169.jpg" alt="Feature image showing a person holding a smartphone and using AI chat software" width="300" height="169" data-wp-editing="1" srcset="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>AI may be rapidly increasing the speed at which we can generate software and find answers, but where there&#8217;s an interface for a human, interactions aren&#8217;t always smooth. As AI providers compete for users, they add features. Although features add capabilities, <a href="https://www.digitalisleofman.com/ai-news/how-generative-ai-usage-has-changed-in-a-year/">they can also add complexity</a>. Simple chat interfaces are giving way to tabbed experiences with more complicated terminology and navigation structures. It&#8217;s time to see how that may impact the user experience.</p>
<p>To follow-up our <a href="https://measuringu.com/ai-based-chat-software-ux-2025/">2025 benchmark of AI-based chat software</a>, we conducted another retrospective study of four AI-based chat software products: ChatGPT, Claude, Gemini, and Grok. In this article, we present key UX findings from our 2026 investigation with some comparisons to those 2025 findings. For more details, see <a href="https://measuringu.com/product/ux-benchmark-report-for-ai-based-chat-software-2026">the full report</a>.</p>
<h2><span lang="EN-US">AI-Based Chat Software Benchmark Study</span></h2>
<p>In May 2026, we conducted a retrospective study of four AI-based chat software products with 420 U.S.-based panel participants. This study included the metrics we typically collect in our <a href="https://measuringu.com/consumer-software-ux-2025/">standard UX and NPS study of consumer software</a>.</p>
<p>There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several questionnaires, including the <a href="https://measuringu.com/nps-ux/">NPS</a>, <a href="https://measuringu.com/10-things-sus/">SUS</a>, <a href="https://measuringu.com/from-umux-lite-to-ux-lite/">UX-Lite</a><sup>®</sup>, and <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" />. The AI-based chat products and sample sizes were:</p>
<ul>
<li>ChatGPT: 113</li>
<li>Claude: 103</li>
<li>Gemini: 101</li>
<li>Grok: 103</li>
</ul>
<p>The sample sizes are modest but <a href="https://measuringu.com/might-not-be-a-magic-number-but-there-are-magic-ranges/">adequate to establish baselines and identify medium-sized differences</a> relative to each other and other software products we measure (e.g., ±10% for binary metrics; ±5 for 0–100-point rating scales).</p>
<h2><span lang="EN-US">UX Metrics Results</span></h2>
<p>Ease of use and usefulness affect whether people use and recommend products. In this section, we review the UX-Lite, SUS, and NPS results.</p>
<h3><span lang="EN-US">Perceived Usefulness and Ease (UX-Lite)</span></h3>
<p>The <a href="https://measuringu.com/how-to-score-and-interpret-the-ux-lite/">UX-Lite</a> has emerged as an industry standard for assessing the two constructs that matter most in technology adoption: ease of use and usefulness. These aren&#8217;t arbitrary choices. Research on the <a href="https://en.wikipedia.org/wiki/Technology_acceptance_model">Technology Acceptance Model</a> (TAM) has repeatedly shown, as far back as the mid-1980s, that ease and usefulness are key drivers of intention to use a product, which is in turn a <a href="https://measuringu.com/article/effect-of-perceived-ease-of-use-and-usefulness-on-ux-and-behavioral-outcomes/">significant driver of actual use</a>.</p>
<p>The UX-Lite captures both constructs with just two items: one rating perceived ease of use (&#8220;This product is easy to use&#8221;) and one rating usefulness (&#8220;This product&#8217;s features meet my needs&#8221;). Together, they give UX researchers a compact but validated measure of acceptance—or more broadly, satisfaction and product quality.</p>
<p>Figure 1 shows the ease and usefulness scores for 2025 (blue) and 2026 (green). The dashed red lines indicate the overall means.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure1-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47950 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure1-1024x331.png" alt="Scatterplot of the two UX-Lite subscales for the AI-based chat products in 2025 and 2026 (means across products indicated by red dashed lines; Grok collected in 2026 only)." width="1024" height="331" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure1-1024x331.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-300x97.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-768x248.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-1536x496.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-2048x661.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-600x194.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Scatterplot of the two UX-Lite subscales for the AI-based chat products in 2025 and 2026 (means across products indicated by red dashed lines; Grok collected in 2026 only).</p>
<p>The mean UX-Lite scores for 2026 ranged by just 5.5 points, from 77.8 for Grok to 83.3 for ChatGPT. There was no significant difference for the main effect of Product (<em>F</em>(3, 416) = 1.7, <em>p</em> = .16).</p>
<p>The most notable mover was Gemini, which dropped from its 2025 position in the upper right quadrant to the lower left in 2026. This was the largest year-over-year shift of any product. ChatGPT remained the most stable, with nearly identical scores in both years. Claude improved meaningfully in perceived usefulness but still lags in ease, keeping it left of center. Grok, measured for the first time, has room to grow as it sits slightly below the group average on both dimensions.</p>
<h3><span lang="EN-US">Perceived Usability (SUS)</span></h3>
<p>While the UX-Lite is a compact way to measure ease, many organizations still use the System Usability Scale (SUS) for historical comparability. SUS is a ten-item questionnaire with possible scores ranging from 0 to 100. The average <a href="https://measuringu.com/product/suspack/">SUS score from over 500 products</a> (including websites and business software) is 68 (a grade of C on the <a href="https://measuringu.com/interpret-sus-score/">Sauro-Lewis curved grading scale</a>).</p>
<p>Our 2026 finding was that ChatGPT led with the highest SUS score, but the range for the products was just 3.1 points (78.4 to 81.5, no statistically significant difference, <em>F</em>(3,416) = .97, <em>p</em> = .41). All products had above-average perceived usability (at least a grade of B+).</p>
<p>As shown in Figure 2, the SUS scores were reasonably stable from 2025 to 2026. For the products measured in both years, the main effects were not statistically significant (Product: <em>F</em>(2, 464) = .63, <em>p</em> = .63; Year: <em>F</em>(2, 464) = .58, <em>p</em> = .45), but there was some indication of interaction likely due to the 5.3-point increase for ChatGPT (<em>F</em>(2, 464) = 2.3, <em>p</em> = .10).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure2-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47946 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure2-1024x347.png" alt="SUS with 95% confidence intervals from data collected in 2025 and 2026." width="1024" height="347" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure2-1024x347.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-300x102.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-768x260.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-1536x520.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-2048x693.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-600x203.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> SUS with 95% confidence intervals from data collected in 2025 and 2026.</p>
<h3><span lang="EN-US">Recommendation Intention (NPS)</span></h3>
<p>AI software usage has grown dramatically from word of mouth as friends and colleagues describe the latest thing they did with AI. The Net Promoter Score provides a good gauge of word-of-mouth recommendations that can portend rapid growth of products. NPS is calculated using an 11-point (0 to 10) likelihood-to-recommend (LTR) question, computed by subtracting the percentage of detractors (0–6) from promoters (9–10).</p>
<p>Figure 3 shows the NPS we obtained for these AI-based chat products compared to our 2025 findings. General <a href="https://www.qualtrics.com/experience-management/customer/good-net-promoter-score/">guidelines for the interpretation of the NPS</a> are that anything above 0 is good (more promoters than detractors), above 20 is favorable, and above 50 is excellent.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure3-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47947 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure3-1024x339.png" alt="NPS with 95% confidence intervals from data collected in 2025 and 2026." width="1024" height="339" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure3-1024x339.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-300x99.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-768x254.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-1536x509.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-2048x678.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-600x199.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> NPS with 95% confidence intervals from data collected in 2025 and 2026.</p>
<p>NPS declined for all the products measured in 2025. Although the number of ChatGPT users hasn’t changed, its <a href="https://fortune.com/2026/02/05/chatgpt-openai-market-share-app-slip-google-rivals-close-the-gap/">percentage of market share declined in 2026</a> as competitors’ share has increased, likely contributing to the observed drop in likelihood to recommend. Most notably, Claude now has the relatively highest NPS, reflecting its <a href="https://www.kucoin.com/news/flash/claude-paid-user-growth-outpaces-chatgpt-in-2026">recent dominant growth</a>. Gemini’s user base has increased, though much of this growth is not from user choice but <a href="https://www.tomsguide.com/ai/chatgpts-traffic-share-hits-lowest-point-since-2023-as-gemini-surges-new-report-exposes-pressure-on-openai">due to integration in the Google ecosystem</a>, possibly diluting recommendation intensity. The confidence intervals show that an NPS of 0% for 2026 is plausible for ChatGPT and Gemini, but not for Claude or Grok.</p>
<h2><span lang="EN-US">Analysis of Verbatim Comments</span></h2>
<p>To dig into the “why” behind the current numbers, we asked participants to name one thing they disliked about the product they rated. Table 1 shows the top three issues for each product (with sample participant quotes). Some clear themes emerged.</p>
<p>Accuracy and reliability were the most common complaints across all four products, with ChatGPT, Gemini, and Grok all drawing criticism for <strong>inaccurate or inconsistent responses</strong> (Grok in particular for hallucination loops). Claude stood apart, with users citing <strong>usage limits</strong> and prompt misinterpretation rather than accuracy, while slow performance was a recurring frustration for Gemini and Grok users.</p>

<table id="tablepress-1053" class="tablepress tablepress-id-1053 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<th class="column-1">Product</th><th class="column-2">Top Three Issues</th><th class="column-3">Sample User Quote</th>
</tr>
</thead>
<tbody class="row-striping">
<tr class="row-2">
	<td rowspan="3" class="column-1"><strong>ChatGPT</strong></td><td class="column-2">Accuracy/Reliability Issues</td><td class="column-3">"It's very confident in its errors, so I need to pay careful attention to make sure I'm getting the right information."</td>
</tr>
<tr class="row-3">
	<td class="column-2">Slow to Respond</td><td class="column-3">"It gets laggy when the conversation gets long."</td>
</tr>
<tr class="row-4">
	<td class="column-2">Subscription/Access Limitations</td><td class="column-3">"Limited Features in the Free Version."</td>
</tr>
<tr class="row-5">
	<td rowspan="3" class="column-1"><strong>Claude</strong></td><td class="column-2">Limited Capabilities/Restrictions</td><td class="column-3">"Usage limits from both the free and paid tiers can be frustrating sometimes."</td>
</tr>
<tr class="row-6">
	<td class="column-2">Performance Issues</td><td class="column-3">“Sometimes the responses are slow, and the website can feel limited when handling long conversations or complex tasks."</td>
</tr>
<tr class="row-7">
	<td class="column-2">Task Limitations</td><td class="column-3">“It does not translate data as well as I’d like. It tends to complicate tasks and I have to check its work.“</td>
</tr>
<tr class="row-8">
	<td rowspan="3" class="column-1"><strong>Gemini</strong></td><td class="column-2">Inaccuracy/Inconsistent Responses</td><td class="column-3">“Sometimes, I have to do some fact checking about information Gemini found."</td>
</tr>
<tr class="row-9">
	<td class="column-2">Not Responding as Prompted</td><td class="column-3">“Occasionally it misinterprets what I'm asking so I figure out a different way to prompt what I need."</td>
</tr>
<tr class="row-10">
	<td class="column-2">Slow Performance/Processing Issues</td><td class="column-3">“I think that I've had a lot of lagging or that it’s slow.“</td>
</tr>
<tr class="row-11">
	<td rowspan="3" class="column-1"><strong>Grok</strong></td><td class="column-2">Inaccuracy/Inconsistent Responses</td><td class="column-3">“Sometimes the bot can get stuck in a loop if it experiences a hallucination."</td>
</tr>
<tr class="row-12">
	<td class="column-2">Slow Performance/Processing Issues</td><td class="column-3">“It takes too long to get answers."</td>
</tr>
<tr class="row-13">
	<td class="column-2">Image/Video Creation/Editing Errors</td><td class="column-3">“Inconsistent results when generating, annoying navigation, user privacy doesn't seem great.“</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1053 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Top issues for AI-based chat software products.</p>
<h2>Do Users of These Products Differ in Tech Savviness?</h2>
<p>We don’t want differences in metrics to just be a result of differences in participant tech-savviness. It could be that less tech-savvy users gravitate toward ChatGPT, and more tech-savvy users prefer Claude or Grok. To differentiate between differences in participant ability and differences in the user experience, we collected tech-savviness scores using our <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10 measure</a> for each of the three products.</p>
<p>The <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a> (Technical Activity Checklist with ten items) is a reliable (consistent) and valid (predictive) measure of tech savviness. The TAC-10 score for a person is the number of items selected from its checklist.</p>
<p>In 2026, the TAC-10 scores for the four AI products ranged from 6.3 to 7.0 (see Figure 4), a statistically significant main effect (<em>F</em>(3, 416) = 2.8, <em>p</em> = .04), due to the difference between ChatGPT and Claude (a pattern similar to what we saw in 2025 when the difference was 2.3 points, but less pronounced in 2026). All mean scores were in the lower part of the high range for two-group classification and medium range for three-group classification except for ChatGPT in 2025, which just missed the cut-off of 5 to be in those groups.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure4-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47948 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure4-1024x300.png" alt="TAC-10 scores by product and year (with 95% confidence intervals)." width="1024" height="300" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure4-1024x300.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-300x88.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-768x225.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-1536x450.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-2048x600.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-600x176.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> TAC-10 scores by product and year (with 95% confidence intervals).</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Results of our AI-based chat benchmarks based on 420 participants revealed:</p>
<p><strong>ChatGPT was the leader in perceived usability, but just barely.</strong> ChatGPT led in SUS scores in 2026, enjoying a 5.3-point increase from 2025. The range of SUS scores was, however, just 3.1 points (not statistically significant). On the UX-Lite, ChatGPT also led in perceived ease of use and usefulness, though differences among products were not statistically significant.</p>
<p><strong>For Net Promoter Scores, Claude was the leader, and ChatGPT was the laggard.</strong> In 2025, the NPS for Claude was only 4 points higher than ChatGPT, but in 2026 the difference was 21 points in favor of Claude (28% vs. 7% for ChatGPT), with Gemini (12%) and Grok (17%) in between. This reflects Claude&#8217;s recent surge in growth.</p>
<p><strong>Claude showed the biggest gain in perceived usefulness.</strong> Of the products measured in both years, Claude had the largest year-over-year improvement on the UX-Lite usefulness dimension, though it still lags behind ChatGPT on perceived ease of use.</p>
<p><strong>Frequently reported issues included inaccurate responses and limited capabilities.</strong> Respondents reported accuracy issues with ChatGPT, Gemini, and Grok, and criticized Claude for limited capabilities and frustrating usage limits. Other reported problems were slow response time and not responding as prompted.</p>
<p><strong>ChatGPT users had slightly lower tech savviness than Claude users.</strong> The mean TAC-10 scores were all in the medium range of tech savviness. Within that range, there was a statistically significant difference (6.3 for ChatGPT vs. 7.0 for Claude), though the 0.7-point gap on a 10-point scale suggests it may not be practically significant (compared to the 2.3-point difference in 2025).</p>
<p>For more details on these products, see <a href="https://measuringu.com/product/ux-benchmark-report-for-ai-based-chat-software-2026">the full report</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>What Are the Different Types of Synthetic Users?</title>
		<link>https://measuringu.com/what-are-the-different-types-of-synthetic-users/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=what-are-the-different-types-of-synthetic-users</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 23 Jun 2026 21:54:24 +0000</pubDate>
				<category><![CDATA[Survey]]></category>
		<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Synthetic user]]></category>
		<category><![CDATA[Synthetic users]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47761</guid>

					<description><![CDATA[Recruiting participants for research is expensive. It’s also rife with problems: Are these people really who they say they are? Are they actually paying attention? Or is the data from some survey farm where people click through and make money? AI is disrupting UX research. But the disruption is leading to more software, not less. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47827 size-medium" src="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-300x169.jpg" alt="Feature image showing 5 different AI bots representing 5 synthetic user types" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>Recruiting participants for research is expensive. It’s also rife with problems: Are these people really who they say they are? Are they actually paying attention? Or is the data from some <a href="https://www.intellisurvey.com/blog/fighting-survey-farms">survey farm</a> where people click through and make money?</p>
<p>AI is disrupting UX research. But the disruption is leading to more software, not less. The need for insights into how people will use that software isn’t going away.</p>
<p>But can AI help? Can we use AI to synthesize people’s attitudes, beliefs, and behaviors? Instead of trying to find the right people to take surveys, could these <strong>synthetic users</strong> generate insights faster and at almost no cost? News about synthetic users would certainly make headlines. <a href="https://measuringu.com/review-of-experiments-with-synthetic-users/">And they do</a>.</p>
<p>But what exactly <em>is</em> a synthetic user? Is that the same as a digital twin? Or a synthetic persona?</p>
<p>To properly assess the effectiveness of AI tools, we think it’s important to have a good understanding of the terms and how they fit together.</p>
<p>In this article, we propose a preliminary taxonomy of five distinct types of synthetic users, organized by how grounded they are in real human data. Before we get to the taxonomy, though, it helps to ask a question that sounds simpler than it is.</p>
<h2>What Birds Can Teach Us about Synthetic Users</h2>
<p>How do we know that a bird is a bird?</p>
<p>Is it because a bird can fly? Well, bats are mammals that can fly, while penguins are birds that can’t fly.</p>
<p>Is it because they lay eggs? Platypuses are mammals that lay eggs (as do most reptiles, amphibians, fish, and arthropods).</p>
<p>Maybe it’s because birds have feathers rather than scales or fur? That might be true in the present, but in the past, many dinosaurs not in the lineage leading to birds are <a href="https://www.smithsonianmag.com/science-nature/dinosaurs-evolved-feathers-for-far-more-than-flight-180985012/">now known to have had feathers</a>.</p>
<p>The answer, as <a href="https://www.linnean.org/learning/who-was-linnaeus/career-and-legacy">Linnaeus worked out in the 1700s</a>, is that no single feature is sufficient. A bird is defined by a <em>constellation</em> of characteristics (feathers, beak, two wings, two feet, warm blood, hard-shelled eggs) organized within a hierarchy of kingdom, class, genus, and species. Figure 1 shows how that plays out from the animal kingdom down to a single species. And this is our guide for classifying and understanding synthetic users.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png" rel="attachment wp-att-47768"><img loading="lazy" decoding="async" class="alignnone wp-image-47831 size-full" src="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png" alt="Figure 2: Classification of different types (species) of synthetic users. " width="2560" height="1349" srcset="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-300x158.png 300w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-1024x540.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-768x405.png 768w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-1536x809.png 1536w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-2048x1079.png 2048w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-600x316.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Example of classification from the animal kingdom to species of stork.</p>
<h2>Synthetic Users: More of a Genus than a Species</h2>
<p>The topic of classifying different types of synthetic users is in a state of flux (lots of labels, overlapping meanings, vendor-specific definitions). Despite this, in Figure 2, we attempt a preliminary classification scheme similar to Figure 1 for five types (species) under the genus of Synthetic User.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png" rel="attachment wp-att-47769"><img loading="lazy" decoding="async" class="alignnone wp-image-47832 size-full" src="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png" alt="Classification of different types (species) of synthetic users." width="1185" height="868" srcset="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png 1185w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-300x220.png 300w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-1024x750.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-768x563.png 768w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-600x439.png 600w" sizes="auto, (max-width: 1185px) 100vw, 1185px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Classification of different types (species) of synthetic users.</p>
<p>Table 1 lists the identifying characteristics for each of these types of synthetic users, primarily focusing on the type of data used to create the synthetic user and how grounded the synthetic user is in actual user data.</p>

<table id="tablepress-1052" class="tablepress tablepress-id-1052">
<thead>
<tr class="row-1">
	<th class="column-1">Synthetic User Type</th><th class="column-2">Identifying Characteristics/Descriptions</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">AI Proto Persona</td><td class="column-2">This is the weakest (least grounded) type of synthetic user generated with simple role-playing prompts (e.g., “You are a world-class Python programmer”). This method produces preliminary user profiles based on broad assumptions rather than research.</td>
</tr>
<tr class="row-3">
	<td class="column-1">Demographic Based</td><td class="column-2">Prompts specify age, gender, occupation, region, etc. to approximate group-level tendencies. This method is somewhat more grounded than a proto persona but is still limited in the quality of its output, especially when demographics have only weak relationships with research topics (e.g., much UX research).</td>
</tr>
<tr class="row-4">
	<td class="column-1">Persona Based</td><td class="column-2">Prompts focus on richer persona paragraphs (e.g., “Bill G. is a 27-year old male graphic designer who always has his sketchbook at hand, has a track record of being creative and innovative, and is up-to-date on current design trends. How would he complete the following questionnaire?”). Because these synthetic users are still weakly grounded, they are limited to approximate group-level tendencies.</td>
</tr>
<tr class="row-5">
	<td class="column-1">Research Grounded</td><td class="column-2">Prompts refer to actual research artifacts with traceable sources but do not attempt to model individual human responses. These are based on actual interviews, survey results, analytics, customer-support logs, or other user data that are typically not available to publicly generated LLMs.</td>
</tr>
<tr class="row-6">
	<td class="column-1">Digital Twins</td><td class="column-2">Prompts refer to rich individual-level data for the purpose of modeling each individual in a dataset. This approach has the strongest grounding in actual user data but its accuracy in real-world deployments is still an open research question.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1052 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1</strong>: Brief descriptions of types of synthetic users.</p>
<h2>Discussion</h2>
<p>In this preliminary taxonomy, we’ve defined five types of synthetic users: AI proto persona, demographic based, persona based, research grounded, and digital twins, based on differences in the types of data (e.g., demographic, persona) and the strength of the relationship between the synthetic user and human user data.</p>
<p>Preliminary taxonomies change over time. In this article, we’ve used the levels originally defined by Linnaeus because they were adequate for our purpose. Modern biological taxonomies have eight levels (Domain, Kingdom, Phylum, Class, Order, Family, Genus, Species), and the number of kingdoms has increased to six (Bacteria, Archaea, Protista, Fungi, Plants, Animals).</p>
<p>We fully expect changes to our classification scheme over time, but it’s a start.</p>
<p>For example, we have not included generative agents in this taxonomy because they are qualitatively different from synthetic users that simulate responses and are more like simulated actors, trying to model what people might do over time. This may eventually become its own branch from the genus of synthetic users, separate from the synthetic respondents. Time will tell.</p>
<p>Just like how there are hybrids in the animal kingdom (e.g., mules, ligers), in practice, there may be hybrids of different types of synthetic users. For example, in Bisbee et al.’s “Synthetic Replacements for Human Survey Data? The Perils of Large Language Models” (<a href="https://www.cambridge.org/core/journals/political-analysis/article/synthetic-replacements-for-human-survey-data-the-perils-of-large-language-models/B92267DC26195C7F36E63EA04A47D2FE">2024</a>), the researchers used the following prompt to elicit 30 synthetic responses for each of the 7,350 human respondents for each of the 7,350 human respondents in the 2016–2020 ANES survey to get a final dataset with 3,614,400 responses:</p>
<p style="padding-left: 25px;"><em>It is [YEAR]. You are a [AGE] year-old, [MARST], [RACETH] [GENDER] with [EDUCATION] making [INCOME] per year, living in the United States. You are [IDEO], [REGIS] [PID] who [INTEREST] pays attention to what’s going on in government and politics. Provide responses from this person’s perspective. Use only knowledge about politics that they would have.</em></p>
<p>Each bracketed item is a variable with values corresponding to a real respondent in the two waves of the ANES. For example, [YEAR] was 2016 or 2020, [AGE] matched the selected respondent’s age, [MARST] was marital status (e.g., married, divorced, single), [IDEO] was political ideology from extremely liberal to extremely conservative, [REGIS] was voter registration status, and [PID] was party membership (Democrat, Independent, Republican).</p>
<p>Thus, this is a hybrid between demographic- and persona-based types with a light sprinkle of digital twinning. It’s less than a fully research-grounded respondent or digital twin because the prompt doesn’t include access to the respondent’s prior answers, interview transcript, open-ended comments, voting history, occupation, or religion. It uses selected ANES variables as conditioning attributes and asks the LLM to answer from that perspective (multiple times for each human respondent).</p>
<h2>Summary</h2>
<p>Our key conclusions from this exercise are:</p>
<p><strong>“Synthetic users” is more of an umbrella term (like a genus) than a type (like a species). </strong></p>
<p>All five types we’ve described can fall under the umbrella of synthetic users. In practice, that means when we talk about synthetic users, it’s like talking about storks. There are a variety of storks, so knowing which bird we’re talking about helps move the conversation forward.</p>
<p><strong>Key criteria for discriminating among types of synthetic users include data type and grounding. </strong></p>
<p>The types we’ve defined differ in the kind of data used to model responses (e.g., demographic, persona) and the extent to which they are grounded in real user data.</p>
<p><strong>Taxonomies change over time. </strong></p>
<p>We consider this article a necessary exercise in a preliminary taxonomy of synthetic users, but fully expect it to evolve over time, maybe quickly due to rapid changes in these technologies.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!-- plugin=object-cache-pro client=phpredis metric#hits=8978 metric#misses=132 metric#hit-ratio=98.6 metric#bytes=4229006 metric#prefetches=243 metric#store-reads=134 metric#store-writes=159 metric#store-hits=385 metric#store-misses=114 metric#sql-queries=52 metric#ms-total=879.77 metric#ms-cache=42.60 metric#ms-cache-avg=0.1459 metric#ms-cache-ratio=4.8 -->
