<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>MeasuringU</title>
	<atom:link href="https://measuringu.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://measuringu.com</link>
	<description>UX Research and Software</description>
	<lastBuildDate>Thu, 23 Jul 2026 17:00:52 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.9.5</generator>

<image>
	<url>https://measuringu.com/wp-content/uploads/2020/11/site-icon.png</url>
	<title>MeasuringU</title>
	<link>https://measuringu.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>UX and NPS Benchmarks of Health Insurance Websites (2026)</title>
		<link>https://measuringu.com/ux-nps-benchmark-report-for-healthcare-websites-2026/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=ux-nps-benchmark-report-for-healthcare-websites-2026</link>
		
		<dc:creator><![CDATA[Eva Sundberg, MPH&nbsp;•&nbsp;Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 21 Jul 2026 20:52:06 +0000</pubDate>
				<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Healthcare]]></category>
		<category><![CDATA[NPS]]></category>
		<category><![CDATA[SUPR-Q]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=48005</guid>

					<description><![CDATA[Tech has improved a lot over the last ten years. It’s hard to beat the promised convenience of managing health insurance online. Gone are the days of hunting through paper policy booklets, waiting on an employer’s benefits manager, or enduring the endless hold music of a telephone call center. Modern portals promise complete, on-demand control [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat.png"><img fetchpriority="high" decoding="async" class="alignleft wp-image-48043 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-300x169.png" alt="Feature image showing healthcare workers" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat-600x338.png 600w, https://measuringu.com/wp-content/uploads/2026/07/072126_Feat.png 1280w" sizes="(max-width: 300px) 100vw, 300px" /></a>Tech has improved a lot over the last ten years.</p>
<p>It’s hard to beat the <em>promised</em> convenience of managing health insurance online. Gone are the days of hunting through paper policy booklets, waiting on an employer’s benefits manager, or enduring the endless hold music of a telephone call center.</p>
<p>Modern portals promise complete, on-demand control over your healthcare: real-time network coverage and deductible tracking, a doctor search that works like any other search bar, and a full claims history available any time, day or night.</p>
<p>But this digital autonomy comes with high-stakes drawbacks. Because health insurance portals must organize fragmented clinical, network, and financial data, they frequently lack the intuitive ease that consumers expect from modern digital interfaces. It’s still hard to find out if a doctor is in your insurance network or if parts of your claim won&#8217;t be covered.</p>
<p>The circumstances behind the visit can also affect usability. Unlike low-stakes commercial retail websites, where users can casually learn navigational patterns over time, a health insurance portal is rarely visited for leisure. A user&#8217;s first interaction often occurs during a high-stress medical event where cognitive load is already strained, and a clunky, unfamiliar layout becomes a barrier to care.</p>
<p>Despite these issues, digital engagement is at an all-time high, with independent market data revealing that <a href="https://www.flexpa.com/blog/digital-by-default">85.5% of insured Americans have created an online account with their health insurance provider</a>. However, because of foundational usability flaws, the baseline experience fails consumers when they need help the most. Recent industry benchmarks show that <a href="https://www.jdpower.com/business/u-s-healthcare-digital-experience-study/">health plans fail to deliver an easy-to-use digital experience 39% of the time</a>, and their consumer satisfaction scores drastically trail other digital sectors like retail banking and property insurance. This makes the quantitative improvement of health insurance websites&#8217; user experience a critical priority for both insurers and patients.</p>
<p>This isn’t the first time we’ve checked in on this industry. When we <a href="https://measuringu.com/ux-health-insurance/">benchmarked six major health insurance websites in 2018</a>, the group’s average SUPR-Q score sat at the 67th percentile—solidly above average. Eight years, and presumably a lot of digital investment, later? The group founders at the 30th percentile.</p>
<p>To understand the quality of their online experiences today, we collected UX benchmark metrics on eight common healthcare websites and mobile applications.</p>
<ul>
<li>Aetna (CVS Health)</li>
<li>Blue Cross Blue Shield</li>
<li>Cigna</li>
<li>Elevance Health (Anthem)</li>
<li>Highmark</li>
<li>Humana</li>
<li>Kaiser Permanente</li>
<li>UnitedHealthcare</li>
</ul>
<p>We computed <a href="https://measuringu.com/product/suprq/">SUPR-Q</a><sup>®</sup> and <a href="https://measuringu.com/nps-reliability/">Net Promoter</a> scores, measured users’ attitudes regarding their experiences, conducted <a href="https://measuringu.com/key-drivers/">key driver</a> analyses, and analyzed reported usability problems. (Full details are in the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>.)</p>
<h2>Benchmark Study Details</h2>
<p>From February to March 2026, we asked 391 users of healthcare websites in the U.S. to recall their most recent experience and perceptions of one of these websites on their desktop and mobile app (if applicable) in the past year.</p>
<p>Respondents completed the eight-item <a href="https://measuringu.com/10-things-suprq/">SUPR-Q</a> (which includes the <a href="https://measuringu.com/nps-ux/">Net Promoter Score</a>), the two-item <a href="https://measuringu.com/evolution-of-the-ux-lite/">UX-Lite</a><sup>®</sup>, and the <a href="https://measuringu.com/article/streamlining-the-supr-qm-the-supr-qm-v2/">SUPR-Qm</a> standardized questionnaires, and answered questions about their brand attitudes, usage, and prior experiences.</p>
<h2>Quality of the Website User Experience: SUPR‑Q</h2>
<p>The SUPR-Q is a standardized questionnaire widely used for measuring attitudes toward the quality of a website user experience. Its norms are computed from a rolling database of around 200 websites across dozens of industries.</p>
<p>SUPR-Q scores are percentile ranks that tell you how a website’s experience ranks relative to other websites (50<sup>th</sup> percentile is average). The SUPR-Q provides an overall score as well as detailed scores for subdimensions of Usability, Trust, Appearance, and Loyalty.</p>
<p>The mean SUPR-Q across healthcare websites in this study was at the 30<sup>th</sup> percentile (substantially below average), ranging from the 12<sup>th</sup> percentile for Highmark to the 49<sup>th</sup> percentile for Cigna.</p>
<h3>Usability Scores</h3>
<p>Overall, usability scores were also well below average for the healthcare websites, averaging at the 21<sup>st</sup> percentile. Highmark had the lowest usability score at the 7<sup>th</sup> percentile, and Cigna had the highest (57<sup>th</sup> percentile).</p>
<p>Comments related to usability on Highmark included:</p>
<p style="padding-left: 25px;"><em>“Certain pages require multiple clicks to get to basic information, and the navigation isn’t always intuitive, which can make the experience a bit frustrating.”</em></p>
<p style="padding-left: 25px;"><em>“The layout is confusing. Too much time looking for what I need.”</em></p>
<h3>Loyalty/Net Promoter Scores</h3>
<p>By itself, the Net Promoter Scores isn’t necessarily a good metric for health insurance websites, since many users don&#8217;t get to choose their insurer (and the site they have to use). A low score could reflect more on the industry, a limited network, or a mandated plan than it does on the user experience. Despite this potential confusion, we still report it for two reasons. First, comparing scores across companies in the same industry (and over time periods) helps identify who&#8217;s delivering a relatively better or worse experience, as the underlying constraint is the same for everyone. Second, NPS remains one of the few metrics that needs no explanation in a boardroom: stakeholders already know what a score of 40% versus −20% means, which makes it a useful entry point when discussing more relevant UX metrics associated with the NPS.</p>
<p>All but two healthcare websites, Humana and Aetna (CVS Health), had negative Net Promoter Scores, led by Humana at 14% and followed by Aetna (CVS Health) at a neutral 0%. Highmark scored the worst behind the group at−47%. The average NPS for these websites was −14%, indicating a substantial surplus of brand detractors over promoters across the industry.</p>
<p>Comments related to NPS and loyalty included:</p>
<p style="padding-left: 25px;"><em>“For a website that has a variety of audiences visiting it who have very different needs, it does a good job and is more simple than having to navigate to a website specific to who you are (employer versus provider versus patient, etc.).” </em>— Blue Cross Blue Shield</p>
<p style="padding-left: 25px;"><em>“I just don&#8217;t think they are the best.  I am not sure the information is reliable especially when trying to find a provider or a specialist.” </em>— Highmark</p>
<h2>Websites and Mobile App Usage</h2>
<p>As a part of this benchmark, we asked respondents how they accessed the healthcare providers online. All respondents reported using their desktop/laptop computers (this was a requirement for participation in the survey), with 67% also using mobile apps and 69% using mobile websites. Most respondents reported visiting their healthcare websites on a desktop or a laptop computer a few times per year. Mobile app users showed more variation, with 40% of Aetna (CVS Health) and 39% of Kaiser Permanente users reporting using mobile apps a few times per month, and 33% of Cigna users reporting using mobile apps a few times per year. The majority of other healthcare website users mostly reported having never used the mobile app.</p>
<h2>Key Drivers of UX Quality</h2>
<p>To better understand what affects SUPR-Q scores and Likelihood-to-Recommend (LTR) ratings, we asked respondents to rate potentially important attributes of the healthcare websites on a five-point scale from 1 (Strongly disagree) to 5 (Strongly agree). We conducted key driver analyses (regression modeling) to quantify the extent to which ratings on these items drive (account for) variation in overall SUPR-Q scores and, separately, LTR (the rating from which the NPS is derived; full details are in the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>).</p>
<p>The top key driver of the overall SUPR-Q scores was the <strong>ease of finding providers</strong> (13%), followed by the abilities to find accurate claims information (9%), quickly find information on prior claims (9%), and easily find what they will pay for services (8%). Taken together, eight significant variables accounted for 63% of the variance in SUPR-Q scores.</p>
<p>For likelihood to recommend (LTR), the top key driver was the ease of finding providers (9%), followed by trust that personal and health information are secure (9%). Tying for third are plan coverage being easy to understand (7%) and the ease of finding how much services will cost (7%). Overall, six significant drivers accounted for 46% of the variance in LTR.</p>
<p>Figure 1 shows a scatterplot of the importance and opportunity for improvement for seven key drivers. The combination of importance and opportunity for improvement provides a basis for prioritizing which key drivers to improve. The importance score is the greater of the variance accounted for by the driver in the SUPR-Q and NPS analyses, where larger percentages indicate more importance. The opportunity score is the top-box percentage for the driver, so smaller percentages indicate greater opportunity for improvement (e.g., it would be harder to improve a driver with a top-box percentage of 100% than one with a top-box percentage of 10%).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-scaled.jpg"><img decoding="async" class="alignnone wp-image-48035 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1024x345.jpg" alt="Scatterplot of importance and opportunity for improvement of key drivers." width="1024" height="345" srcset="https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1024x345.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-300x101.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-768x259.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-1536x518.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-2048x691.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/07/072126-Figure-1-600x202.jpg 600w" sizes="(max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Scatterplot of importance and opportunity for improvement of key drivers.</p>
<p>Five of these seven key drivers fell in the FIX quadrant (upper left) with relatively high importance and higher opportunity for improvement:</p>
<ul>
<li>Easy to find providers</li>
<li>Personal/health information is secure</li>
<li>Quick to find information on prior claims</li>
<li>Easy to find what I will pay for services</li>
<li>Accurate claims information</li>
</ul>
<p>Blue Cross Blue Shield achieved the highest top-box scores for quickly finding information about prior claims (46%), easily finding covered doctors or providers (32%), and consumer confidence that personal and health information is secure (30%). Humana led for trust in the accuracy of online claims information (24%), while Blue Cross Blue Shield and Cigna tied for the highest top-box score regarding the ease of finding out-of-pocket cost information (18%).</p>
<p>Conversely, the websites with the lowest top-box scores, suggesting the most room for digital improvement, were Humana for tracking prior claims (12%), Highmark for data security confidence (12%), and Humana again for finding covered providers (14%). Highmark scored the lowest for trust in claims information accuracy (9%), while Aetna (CVS Health) registered the worst top-box performance for clear out-of-pocket cost transparency (10%).</p>
<h2>UX Problems</h2>
<p>We examined the verbatim comments to better understand the problems users had.</p>
<h3>Difficulty Finding Information Was a Universal Grievance</h3>
<p>This issue affected all eight websites and was the single top complaint across every provider in the study. It was a massive pain point for Kaiser Permanente, Highmark, and UnitedHealthcare.</p>
<p style="padding-left: 25px;"><em>“Its navigation is non-intuitive.” </em>— Kaiser Permanente</p>
<p style="padding-left: 25px;"><em>“It is very difficult to find the information I need.” </em>— UnitedHealthcare</p>
<p style="padding-left: 25px;"><em>“I stumble through a lot of different pages before finally finding what I’m looking for.” </em>— Blue Cross Blue Shield</p>
<p>These comments speak to a navigation problem, not a content problem. The needed information is usually on the site <em>somewhere</em>. It’s just buried under menu structures built around the insurer’s org chart instead of the handful of things a member actually comes to do.</p>
<h3>Unreliable Provider Directories Degrade the Experience</h3>
<p>For seven of the eight platforms, outdated or inaccurate directory information emerged as a major barrier to care. Users routinely noted a stark disconnect between the doctors listed online and the reality of who actually accepts their insurance.</p>
<p style="padding-left: 25px;"><em>“They don’t update doctors who are in network and who isn’t.” </em>— Aetna</p>
<p style="padding-left: 25px;"><em>“It is not always accurate on whether a place/doctor takes your insurance or not.” </em>— Blue Cross Blue Shield</p>
<p style="padding-left: 25px;"><em>“Caregivers are listed as being in network or accepting new patients, but when I call to make an appointment, that&#8217;s not the case.” </em>— UnitedHealthcare</p>
<p>Figure 2 shows a critical directory system error on the Blue Cross Blue Shield search results page. Instead of presenting a diverse list of nearby options, the portal repeats the exact same facility multiple times in a row down the page. This type of technical redundancy unnecessarily inflates the layout, clutters the interface, and requires members to filter through duplicate data entries to locate distinct, alternative care providers.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png"><img decoding="async" class="alignnone size-full wp-image-48008" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png" alt="Duplicate provider search results on the Blue Cross Blue Shield platform." width="624" height="384" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F2.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F2-300x185.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F2-600x369.png 600w" sizes="(max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Duplicate provider search results on the Blue Cross Blue Shield platform.</p>
<h3>Sluggish Performance and Site Errors Delay Essential Tasks</h3>
<p>Users reported slow loading times, timeouts, or complete site outages across all eight websites. Performance issues were especially disruptive for Highmark, Elevance Health (Anthem), and Kaiser Permanente.</p>
<p style="padding-left: 25px;"><em>“Sometimes when I log in, the site is down.” </em>— Aetna</p>
<p style="padding-left: 25px;"><em>“It&#8217;s clunky, it freezes, it&#8217;s under construction, constant error messages.” </em>— Humana</p>
<p style="padding-left: 25px;"><em>“Sometimes there are error bugs on certain pages although rare. When these different bugs pop up certain information or pages are not accessible.” — Elevance Health (Anthem)</em></p>
<p style="padding-left: 25px;"><em>“Occasional slow page loads or timeouts especially during peak hours.”</em> — UnitedHealthcare</p>
<p style="padding-left: 25px;"><em>“The website is very slow at times and pages fail to load altogether. Getting a generic error message when trying to access my medical records is extremely frustrating.” </em>— Kaiser Permanente</p>
<p>Figure 3 illustrates the exact technical dead end described by these participants. Instead of loading the requested medical records or dashboard utility, the Kaiser Permanente platform encounters a critical error loop, leaving the page blank save for a red exclamation warning badge and a disruptive &#8220;Back to Sign In&#8221; command. This system failure completely breaks the user session, forcing members to entirely restart the multi-step login process to attempt their task again.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48009" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png" alt="A systemic backend routing error on the Kaiser Permanente portal that blocks member access to internal records and triggers session termination." width="624" height="343" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F3.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F3-300x165.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F3-600x330.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> A systemic backend routing error on the Kaiser Permanente portal that blocks member access to internal records and triggers session termination.</p>
<p>These aren’t cosmetic bugs. A frozen page or a timeout in the middle of a claims dispute can mean a missed deadline, not just an annoyance. Every one of the eight sites drew reliability complaints, suggesting that basic uptime and performance are still struggles for an industry pushing members toward digital-only service.</p>
<h3>Cluttered Layouts and Irrelevant Popups Prevent Smooth Navigation</h3>
<p>Navigational bottlenecks heavily impacted users’ baseline experience. Six of the eight sites suffered from significant login friction, while others overwhelmed users with busy, crowded interfaces.</p>
<p style="padding-left: 25px;"><em>“Overwhelming amount of information … the website sometimes feels cluttered with information, making it difficult for me to filter out what doesn&#8217;t apply to me.”</em> — Humana</p>
<p>Figure 4 shows why: the interface presents an overly dense, text-heavy dashboard layout for accessing basic member documents. Rather than organizing resources into clear, scannable visual segments, the platform confronts users with exhaustive walls of prose, complex text blocks outlining compliance requirements, and generic sidebar links. This format spikes cognitive load, forcing members to meticulously read through paragraphs of operational data just to find a specific form link.<a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48010" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png" alt="Dense, text-heavy documentation layouts on the Humana platform that contribute to visual clutter and search fatigue." width="936" height="610" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F4.png 936w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-300x196.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-768x501.png 768w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F4-600x391.png 600w" sizes="auto, (max-width: 936px) 100vw, 936px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> Dense, text-heavy documentation layouts on the Humana platform that contribute to visual clutter and search fatigue.</p>
<p>This layout fatigue is far from an isolated issue; members across other major insurance platforms frequently reported navigating highly saturated interfaces where non-essential material routinely buries functional tools:</p>
<p style="padding-left: 25px;"><em>“It can be confusing at times due to the website being crowded with information.” </em>— Kaiser Permanente</p>
<p style="padding-left: 25px;"><em>“There is a lot of fluff and links to articles that I don&#8217;t care about or have no bearing on the services provided.” </em>— Elevance Health (Anthem)</p>
<p>Figure 5 shows another source of that clutter: a promotional pop-up covering the entire &#8220;Find Care&#8221; page on UnitedHealthcare, which members have to dismiss before they can do anything else. A marketing message is blocking the one task this page exists for.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-48011" src="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png" alt="An intrusive modal pop-up on the UnitedHealthcare website." width="624" height="379" srcset="https://measuringu.com/wp-content/uploads/2026/07/07212026-F5.png 624w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F5-300x182.png 300w, https://measuringu.com/wp-content/uploads/2026/07/07212026-F5-600x364.png 600w" sizes="auto, (max-width: 624px) 100vw, 624px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 5:</strong> An intrusive modal pop-up on the UnitedHealthcare website.</p>
<h2>Summary and Takeaways</h2>
<p>Healthcare payers are massive enterprises, serving millions of Americans who increasingly rely on digital self-service tools. An analysis of the user experience of eight major health insurance platforms using benchmark data found:</p>
<ol>
<li><strong>Cigna and Humana lead; Highmark lags.</strong> Cigna achieved the highest overall SUPR-Q score, falling at the 49<sup>th</sup> percentile, while Highmark had the lowest score, falling in the 12<sup>th</sup> percentile. Humana achieved the highest Net Promoter Score at 14%, while Highmark fell the furthest behind the group at −47%.</li>
</ol>
<ol start="2">
<li><strong>Provider verification and claims transparency drive UX scores.</strong> The top key driver of overall SUPR-Q scores was the ease of finding doctors or providers covered by the user’s plan (13%), followed by the accuracy of claims information (9%), quickly finding info on prior claims (9%), and and easily finding what they will pay for services (8%). Taken together, eight significant variables accounted for 63% of the variance in the SUPR-Q scores. Other critical drivers include the ability to get information without calling customer service (6%) and confidence that personal/health information is secure (5%).</li>
</ol>
<ol start="3">
<li><strong>The top opportunities for improvement are helping members easily find covered providers and allowing them to quickly track prior claims. </strong>One way to prioritize attention to key drivers is to consider both their importance (variability in regression) and how well the websites achieve the stated goal (top-box scores). The two key drivers with the most potential for improvement (high impact percentages and low top-box scores) were the ability to easily find doctors or providers covered by a plan (average top-box score: 21%) and the ability to quickly find information about prior claims (average top-box score: 22%). For finding covered providers and for tracking prior claims, the leader was Blue Cross Blue Shield (Humana lagging).</li>
</ol>
<ol start="4">
<li><strong>Top frustrations are hidden information, website performance issues, and unreliable provider searches.</strong> The most frequently reported issue, affecting all eight websites, was a general difficulty in locating information. Website performance and slow loading times were also universal issues logged across all eight platforms. For seven of the eight systems, users specifically cited that the provider search directory was unreliable or difficult to parse. General login and authentication friction was also highly prevalent, acting as a major pain point on six of the platforms.</li>
</ol>
<ol start="5">
<li><strong>Health insurance members have trouble finding vital information.</strong> Quantitative and qualitative signals from our findings converge on a clear trend: users of these digital health platforms have trouble efficiently uncovering vital policy information. Quantitative signals emphasize broken pathways toward verifying provider networks, reviewing claims, and estimating out-of-pocket pricing. The qualitative &#8220;why&#8221; behind the numbers highlights extensive complaints surrounding buried information, clunky login processes, and sluggish site performance. Quantitative metrics and user commentary both suggest that the platforms that would especially benefit from a strong focus on baseline usability are Highmark, UnitedHealthcare, and Elevance Health.</li>
</ol>
<p>For more details, see the <a href="https://measuringu.com/product/ux-nps-benchmark-report-for-healthcare-websites-2026">downloadable report</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Understanding Alpha Inflation</title>
		<link>https://measuringu.com/understanding-alpha-inflation/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=understanding-alpha-inflation</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 14 Jul 2026 22:23:34 +0000</pubDate>
				<category><![CDATA[Statistics]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[alpha inflation]]></category>
		<category><![CDATA[false positive]]></category>
		<category><![CDATA[Null Hypothesis]]></category>
		<category><![CDATA[Type 1 error]]></category>
		<category><![CDATA[Type I error]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47974</guid>

					<description><![CDATA[The large sample size. The right statistical tests. A compelling and statistically significant finding! All the ingredients of a successful quantitative project that gets stakeholder buy-in. But inflation isn’t just a worry for the Federal Reserve on interest rate policy. It affects research decisions, too. When you conduct something like a usage and attitude survey [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47987 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-300x169.jpg" alt="Feature image showing a researcher inflating a balloon with alpha written on it" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1-600x338.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/071426-FeatureImage-1.jpg 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>The large sample size. The right statistical tests. A compelling and statistically significant finding! All the ingredients of a successful quantitative project that gets stakeholder buy-in. But inflation isn’t just a worry for the Federal Reserve on interest rate policy. It affects research decisions, too.</p>
<p>When you conduct something like a usage and attitude survey with hundreds or thousands of participants (like what we do with our <a href="https://measuringu.com/ai-based-chat-software-ux-2026/">industry reports</a>), you have the benefit of statistical power (the ability to detect differences). This allows you to look not only for differences between main effects (like between software products), but also for interactions (differences by type of user by product).</p>
<p>With enough power, you can dig into these more nuanced differences. For example, you may learn that younger adults have different usage patterns and attitudes than older cohorts for certain products. It would be a compelling story, and it&#8217;s something that stakeholders can make sense of and act on. We’ve seen it and reported it. But how do you know these differences aren’t a fluke?</p>
<p>You may think limiting yourself to reporting only statistically significant findings will protect you from these false positives, and that approach certainly helps. But like printing money to boost the economy, the more statistical tests you run, the more you inflate your chances of finding things that aren’t actually there.</p>
<p>This isn’t monetary policy; it’s alpha inflation. It’s a subtle concept that you should understand when making multiple statistical comparisons.</p>
<h2><span lang="EN-US">Alpha (False Alarm Rate) in Hypothesis Testing</span></h2>
<p>In <a href="https://measuringu.com/how-does-hypothesis-testing-work/">null hypothesis significance testing</a> (NHST), the alpha criterion is the value selected for decisions of statistical significance. This is the process through which you decide if a difference is statistically significant. As shown in Figure 1, when the <em>p</em>-value you get from a statistical test is less than the alpha value you set, you reject the hypothesis of no difference (the null hypothesis, H<sub>0</sub>). It’s statistical significance! Otherwise, you fail to reject H<sub>0</sub> (you don’t accept it; you just don’t have enough evidence to reject it).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1.jpg"><img loading="lazy" decoding="async" class="alignnone wp-image-47983 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-1024x352.jpg" alt="Figure 1: High-level flowchart of statistical hypothesis testing." width="1024" height="352" srcset="https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-1024x352.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-300x103.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-768x264.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1-600x206.jpg 600w, https://measuringu.com/wp-content/uploads/2026/07/071426-Figure1.jpg 1461w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> High-level flowchart of statistical hypothesis testing.</p>
<p>The whole point of this process is to control the percentage of false alarms (rejecting H<sub>0</sub> when there really is no difference) <strong>in the long run</strong>. We’ve previously discussed why <a href="https://measuringu.com/setting-alpha/">the alpha criterion doesn’t have to be the standard <em>p</em> &lt; .05</a>, but the rationale for using .05 was published by R. A. Fisher in 1929: “It is a common practice to judge a result significant, if it is of such a magnitude that it would have been produced by chance not more frequently than once in twenty trials [.05]. This is an arbitrary, but convenient, level of significance for the practical investigator.”</p>
<p>So, when you run one test of significance, if there is really no difference, you have a 5% chance of mistakenly concluding that there is a difference. That’s fine if you’ve collected data from two groups (A and B), but what if you’ve collected data from three groups (A, B, and C) and want to compare A with B, B with C, and A with C? Or, as in the case of the usage and attitude survey, you want to make dozens of comparisons of subgroups (such as age or usage levels across product versions)?</p>
<p>For all these scenarios, you’ll have to deal with the consequences of alpha inflation.</p>
<h2><span lang="EN-US">What Is Alpha Inflation?</span></h2>
<p>When we declare a difference as statistically significant when <em>p</em> &lt; .05, we reduce our chance of being fooled by random noise in our sample. But these false alarms, called <a href="https://measuringu.com/hypothesis-testing-what-can-go-wrong/">Type I</a> errors in statistical jargon, will happen and are expected.</p>
<p>At the <em>p</em> &lt; .05 level of significance, 5% of our conclusions will be false alarms over the long run. That’s 5% for running just one statistical test (such as comparing the <a href="https://measuringu.com/sample-sizes-for-comparison-of-ux-lite-scores/">UX-Lite</a><sup>®</sup> scores of two products). But what if you run more than one statistical test? Software makes it very easy to run all sorts of comparisons such as the differences between age cohorts, experience levels, and products. What if you run 5, 10, 20, or 100 tests? You’re printing money and, like the inflation rate, your false alarm rate goes up too.</p>
<p>To understand how running more tests increases your false alarm (Type I) error rate, Table 1 shows the probability of zero false alarms, one false alarm, two false alarms, and so forth.</p>

<table id="tablepress-1054" class="tablepress tablepress-id-1054">
<thead>
<tr class="row-1">
	<th class="column-1">x (Number of false alarms)</th><th class="column-2"><i>p</i>(x) when <i>α<i> = .05</th><th class="column-3"><i>p</i>(at least x) when <i>α<i> = .05</th>
</tr>
</thead>
<tbody class="row-striping">
<tr class="row-2">
	<td class="column-1">0</td><td class="column-2">0.36</td><td class="column-3">1.00</td>
</tr>
<tr class="row-3">
	<td class="column-1">1</td><td class="column-2">0.38</td><td class="column-3">0.64</td>
</tr>
<tr class="row-4">
	<td class="column-1">2</td><td class="column-2">0.19</td><td class="column-3">0.26</td>
</tr>
<tr class="row-5">
	<td class="column-1">3</td><td class="column-2">0.06</td><td class="column-3">0.08</td>
</tr>
<tr class="row-6">
	<td class="column-1">4</td><td class="column-2">0.01</td><td class="column-3">0.02</td>
</tr>
<tr class="row-7">
	<td class="column-1">5</td><td class="column-2">0.002</td><td class="column-3">0.003</td>
</tr>
<tr class="row-8">
	<td class="column-1">6</td><td class="column-2">0.0003</td><td class="column-3">0.0003</td>
</tr>
<tr class="row-9">
	<td class="column-1">7</td><td class="column-2">0.00003</td><td class="column-3">0.00003</td>
</tr>
<tr class="row-10">
	<td class="column-1">8–20</td><td class="column-2">0.00000</td><td class="column-3">0.00000</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1054 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Alpha inflation for 20 tests conducted with <em>α</em> = .05.</p>
<p>As shown in Table 1, the most likely number of Type I errors in a set of 20 independent tests with <em>α</em> = 0.05 is one, with a point probability of 0.38. That is, when you run 20 comparisons, there’s a <strong>38% chance of getting one false alarm</strong>. The next highest point probability is 0.36 for 0 false alarms, meaning you have about the same chance of having no false alarms in a set of 20 comparisons as you do one false alarm.</p>
<p>Unfortunately, you can encounter more than one false alarm in a set of 20 tests. The likelihood of at least one Type I error, however, is higher (specifically, 1 − <em>p</em>(0) = 1 − 0.36 = 0.64). That’s a 64% chance of 1, 2, 3, or more. (Although more than two or three false alarms in 20 comparisons is rare.) So, rather than having a 5% chance of encountering a Type I error when there is no real difference, <em>α</em> has inflated to 64%.</p>
<h2><span lang="EN-US">What Can You Do About Alpha Inflation?</span></h2>
<p>You can’t raise interest rates to tame alpha inflation. But since the middle of the 20th century, many strategies and techniques have been published to guide the analysis of multiple comparisons, such as omnibus tests (e.g., <a href="https://en.wikipedia.org/wiki/Analysis_of_variance">ANOVA</a> and <a href="https://en.wikipedia.org/wiki/Multivariate_analysis_of_variance">MANOVA</a>) and procedures for the comparison of pairs of means (e.g., Tukey’s <a href="https://en.wikipedia.org/wiki/Tukey's_B_method">WSD</a> and <a href="https://en.wikipedia.org/wiki/Tukey's_range_test">HSD</a> procedures, the <a href="https://en.wikipedia.org/wiki/Newman%E2%80%93Keuls_method">Student–Newman–Keuls test</a>, <a href="https://en.wikipedia.org/wiki/Dunnett%27s_test">Dunnett’s test</a>, the <a href="https://en.wikipedia.org/wiki/Duncan%27s_new_multiple_range_test">Duncan procedure</a>, the <a href="https://en.wikipedia.org/wiki/Scheff%C3%A9%27s_method">Scheffé procedure</a>, the <a href="https://en.wikipedia.org/wiki/Bonferroni_correction">Bonferroni adjustment</a>, and the <a href="https://en.wikipedia.org/wiki/False_discovery_rate#Benjamini%E2%80%93Hochberg_procedure">Benjamini–Hochberg adjustment</a>).</p>
<p>With all these methods available to handle alpha inflation, the problem is solved—right?</p>
<h2><span lang="EN-US">Controlling Alpha Inflation Has Its Consequences</span></h2>
<p>When the null hypothesis is not true, applying techniques to control alpha inflation (decreasing the number of Type I errors) necessarily increases the number of Type II errors—the failure to detect differences that are real. An overemphasis on the prevention of Type I errors leads to the proliferation of Type II errors.</p>
<p>Unless, for your situation, the cost of a Type I error is much greater than the cost of a Type II error, you should avoid applying any of the techniques designed to suppress alpha inflation. As Perneger (<a href="https://jcsites.juniata.edu/faculty/merovich/QuantEcol_files/ESS335Perneger1998_Bonferroni_adjustments.pdf">1998</a>, p. 1236) wrote, “Simply describing what tests of significance have been performed, and why, is generally the best way of dealing with multiple comparisons.” In UX and many applied research settings, failing to detect a real difference can be just as harmful as falsely declaring a difference. Context matters, so don’t think Type I errors should take all the focus.</p>
<p>Or, as Abelson (<a href="https://www.google.com/books/edition/Statistics_as_Principled_Argument/BHiNEQAAQBAJ?hl=en&amp;gbpv=1&amp;pg=PT8&amp;printsec=frontcover">1995</a>, p. 70) put it, “Random patterns will seem to contain something systematic when scrutinized in many particular ways. If you look at enough boulders, there is bound to be one that looks like a sculpted human face. Knowing this, if you apply extremely strict criteria for what is to be recognized as an intentionally carved face, you might miss the whole show on Easter Island.”</p>
<p>We pay more attention to alpha inflation when we’re making many unplanned comparisons. But even then, we have a balanced approach to managing Type I and Type II errors, which we’ll cover in an upcoming article.</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Alpha inflation is what happens when running multiple statistical tests quietly erodes the reliability of your <em>p</em> &lt; .05 threshold, the same way printing money erodes the value of a dollar. The main points from this article are:</p>
<p><strong>Alpha inflation is real. </strong>Run 20 independent comparisons at <em>α</em> = .05, and you don&#8217;t have a 5% chance of a false alarm anymore; you have a 64% chance of <em>at least</em> one. The reality of alpha inflation can easily be demonstrated using binomial probabilities (as in Table 1).</p>
<p><strong>Many methods have been developed to control alpha inflation. </strong>The methods, primarily developed in the 20th century, vary considerably in their relative conservatism (judging fewer contrasts to be significant) and liberalism (judging more contrasts to be significant).</p>
<p><strong>But you can&#8217;t just raise interest rates to fix alpha inflation. </strong>Controlling alpha inflation has hidden costs. Controlling only Type I errors (false alarms) leads to a proliferation of Type II errors (misses). The decision to use methods to control alpha inflation depends strongly on the relative costs of Type I and Type II errors in a specific research context. We&#8217;ll cover how to strike that balance in an upcoming article.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>UX Benchmarks for AI-Based Chat Software (2026)</title>
		<link>https://measuringu.com/ai-based-chat-software-ux-2026/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=ai-based-chat-software-ux-2026</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 07 Jul 2026 22:28:13 +0000</pubDate>
				<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Grok]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47842</guid>

					<description><![CDATA[AI may be rapidly increasing the speed at which we can generate software and find answers, but where there&#8217;s an interface for a human, interactions aren&#8217;t always smooth. As AI providers compete for users, they add features. Although features add capabilities, they can also add complexity. Simple chat interfaces are giving way to tabbed experiences [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47943 size-medium" src="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-300x169.jpg" alt="Feature image showing a person holding a smartphone and using AI chat software" width="300" height="169" data-wp-editing="1" srcset="https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/07/FeatureImage-070726-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>AI may be rapidly increasing the speed at which we can generate software and find answers, but where there&#8217;s an interface for a human, interactions aren&#8217;t always smooth. As AI providers compete for users, they add features. Although features add capabilities, <a href="https://www.digitalisleofman.com/ai-news/how-generative-ai-usage-has-changed-in-a-year/">they can also add complexity</a>. Simple chat interfaces are giving way to tabbed experiences with more complicated terminology and navigation structures. It&#8217;s time to see how that may impact the user experience.</p>
<p>To follow-up our <a href="https://measuringu.com/ai-based-chat-software-ux-2025/">2025 benchmark of AI-based chat software</a>, we conducted another retrospective study of four AI-based chat software products: ChatGPT, Claude, Gemini, and Grok. In this article, we present key UX findings from our 2026 investigation with some comparisons to those 2025 findings. For more details, see <a href="https://measuringu.com/product/ux-benchmark-report-for-ai-based-chat-software-2026">the full report</a>.</p>
<h2><span lang="EN-US">AI-Based Chat Software Benchmark Study</span></h2>
<p>In May 2026, we conducted a retrospective study of four AI-based chat software products with 420 U.S.-based panel participants. This study included the metrics we typically collect in our <a href="https://measuringu.com/consumer-software-ux-2025/">standard UX and NPS study of consumer software</a>.</p>
<p>There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several questionnaires, including the <a href="https://measuringu.com/nps-ux/">NPS</a>, <a href="https://measuringu.com/10-things-sus/">SUS</a>, <a href="https://measuringu.com/from-umux-lite-to-ux-lite/">UX-Lite</a><sup>®</sup>, and <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" />. The AI-based chat products and sample sizes were:</p>
<ul>
<li>ChatGPT: 113</li>
<li>Claude: 103</li>
<li>Gemini: 101</li>
<li>Grok: 103</li>
</ul>
<p>The sample sizes are modest but <a href="https://measuringu.com/might-not-be-a-magic-number-but-there-are-magic-ranges/">adequate to establish baselines and identify medium-sized differences</a> relative to each other and other software products we measure (e.g., ±10% for binary metrics; ±5 for 0–100-point rating scales).</p>
<h2><span lang="EN-US">UX Metrics Results</span></h2>
<p>Ease of use and usefulness affect whether people use and recommend products. In this section, we review the UX-Lite, SUS, and NPS results.</p>
<h3><span lang="EN-US">Perceived Usefulness and Ease (UX-Lite)</span></h3>
<p>The <a href="https://measuringu.com/how-to-score-and-interpret-the-ux-lite/">UX-Lite</a> has emerged as an industry standard for assessing the two constructs that matter most in technology adoption: ease of use and usefulness. These aren&#8217;t arbitrary choices. Research on the <a href="https://en.wikipedia.org/wiki/Technology_acceptance_model">Technology Acceptance Model</a> (TAM) has repeatedly shown, as far back as the mid-1980s, that ease and usefulness are key drivers of intention to use a product, which is in turn a <a href="https://measuringu.com/article/effect-of-perceived-ease-of-use-and-usefulness-on-ux-and-behavioral-outcomes/">significant driver of actual use</a>.</p>
<p>The UX-Lite captures both constructs with just two items: one rating perceived ease of use (&#8220;This product is easy to use&#8221;) and one rating usefulness (&#8220;This product&#8217;s features meet my needs&#8221;). Together, they give UX researchers a compact but validated measure of acceptance—or more broadly, satisfaction and product quality.</p>
<p>Figure 1 shows the ease and usefulness scores for 2025 (blue) and 2026 (green). The dashed red lines indicate the overall means.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure1-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47950 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure1-1024x331.png" alt="Scatterplot of the two UX-Lite subscales for the AI-based chat products in 2025 and 2026 (means across products indicated by red dashed lines; Grok collected in 2026 only)." width="1024" height="331" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure1-1024x331.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-300x97.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-768x248.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-1536x496.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-2048x661.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure1-600x194.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Scatterplot of the two UX-Lite subscales for the AI-based chat products in 2025 and 2026 (means across products indicated by red dashed lines; Grok collected in 2026 only).</p>
<p>The mean UX-Lite scores for 2026 ranged by just 5.5 points, from 77.8 for Grok to 83.3 for ChatGPT. There was no significant difference for the main effect of Product (<em>F</em>(3, 416) = 1.7, <em>p</em> = .16).</p>
<p>The most notable mover was Gemini, which dropped from its 2025 position in the upper right quadrant to the lower left in 2026. This was the largest year-over-year shift of any product. ChatGPT remained the most stable, with nearly identical scores in both years. Claude improved meaningfully in perceived usefulness but still lags in ease, keeping it left of center. Grok, measured for the first time, has room to grow as it sits slightly below the group average on both dimensions.</p>
<h3><span lang="EN-US">Perceived Usability (SUS)</span></h3>
<p>While the UX-Lite is a compact way to measure ease, many organizations still use the System Usability Scale (SUS) for historical comparability. SUS is a ten-item questionnaire with possible scores ranging from 0 to 100. The average <a href="https://measuringu.com/product/suspack/">SUS score from over 500 products</a> (including websites and business software) is 68 (a grade of C on the <a href="https://measuringu.com/interpret-sus-score/">Sauro-Lewis curved grading scale</a>).</p>
<p>Our 2026 finding was that ChatGPT led with the highest SUS score, but the range for the products was just 3.1 points (78.4 to 81.5, no statistically significant difference, <em>F</em>(3,416) = .97, <em>p</em> = .41). All products had above-average perceived usability (at least a grade of B+).</p>
<p>As shown in Figure 2, the SUS scores were reasonably stable from 2025 to 2026. For the products measured in both years, the main effects were not statistically significant (Product: <em>F</em>(2, 464) = .63, <em>p</em> = .63; Year: <em>F</em>(2, 464) = .58, <em>p</em> = .45), but there was some indication of interaction likely due to the 5.3-point increase for ChatGPT (<em>F</em>(2, 464) = 2.3, <em>p</em> = .10).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure2-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47946 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure2-1024x347.png" alt="SUS with 95% confidence intervals from data collected in 2025 and 2026." width="1024" height="347" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure2-1024x347.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-300x102.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-768x260.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-1536x520.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-2048x693.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure2-600x203.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> SUS with 95% confidence intervals from data collected in 2025 and 2026.</p>
<h3><span lang="EN-US">Recommendation Intention (NPS)</span></h3>
<p>AI software usage has grown dramatically from word of mouth as friends and colleagues describe the latest thing they did with AI. The Net Promoter Score provides a good gauge of word-of-mouth recommendations that can portend rapid growth of products. NPS is calculated using an 11-point (0 to 10) likelihood-to-recommend (LTR) question, computed by subtracting the percentage of detractors (0–6) from promoters (9–10).</p>
<p>Figure 3 shows the NPS we obtained for these AI-based chat products compared to our 2025 findings. General <a href="https://www.qualtrics.com/experience-management/customer/good-net-promoter-score/">guidelines for the interpretation of the NPS</a> are that anything above 0 is good (more promoters than detractors), above 20 is favorable, and above 50 is excellent.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure3-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47947 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure3-1024x339.png" alt="NPS with 95% confidence intervals from data collected in 2025 and 2026." width="1024" height="339" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure3-1024x339.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-300x99.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-768x254.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-1536x509.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-2048x678.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure3-600x199.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> NPS with 95% confidence intervals from data collected in 2025 and 2026.</p>
<p>NPS declined for all the products measured in 2025. Although the number of ChatGPT users hasn’t changed, its <a href="https://fortune.com/2026/02/05/chatgpt-openai-market-share-app-slip-google-rivals-close-the-gap/">percentage of market share declined in 2026</a> as competitors’ share has increased, likely contributing to the observed drop in likelihood to recommend. Most notably, Claude now has the relatively highest NPS, reflecting its <a href="https://www.kucoin.com/news/flash/claude-paid-user-growth-outpaces-chatgpt-in-2026">recent dominant growth</a>. Gemini’s user base has increased, though much of this growth is not from user choice but <a href="https://www.tomsguide.com/ai/chatgpts-traffic-share-hits-lowest-point-since-2023-as-gemini-surges-new-report-exposes-pressure-on-openai">due to integration in the Google ecosystem</a>, possibly diluting recommendation intensity. The confidence intervals show that an NPS of 0% for 2026 is plausible for ChatGPT and Gemini, but not for Claude or Grok.</p>
<h2><span lang="EN-US">Analysis of Verbatim Comments</span></h2>
<p>To dig into the “why” behind the current numbers, we asked participants to name one thing they disliked about the product they rated. Table 1 shows the top three issues for each product (with sample participant quotes). Some clear themes emerged.</p>
<p>Accuracy and reliability were the most common complaints across all four products, with ChatGPT, Gemini, and Grok all drawing criticism for <strong>inaccurate or inconsistent responses</strong> (Grok in particular for hallucination loops). Claude stood apart, with users citing <strong>usage limits</strong> and prompt misinterpretation rather than accuracy, while slow performance was a recurring frustration for Gemini and Grok users.</p>

<table id="tablepress-1053" class="tablepress tablepress-id-1053 tbody-has-connected-cells">
<thead>
<tr class="row-1">
	<th class="column-1">Product</th><th class="column-2">Top Three Issues</th><th class="column-3">Sample User Quote</th>
</tr>
</thead>
<tbody class="row-striping">
<tr class="row-2">
	<td rowspan="3" class="column-1"><strong>ChatGPT</strong></td><td class="column-2">Accuracy/Reliability Issues</td><td class="column-3">"It's very confident in its errors, so I need to pay careful attention to make sure I'm getting the right information."</td>
</tr>
<tr class="row-3">
	<td class="column-2">Slow to Respond</td><td class="column-3">"It gets laggy when the conversation gets long."</td>
</tr>
<tr class="row-4">
	<td class="column-2">Subscription/Access Limitations</td><td class="column-3">"Limited Features in the Free Version."</td>
</tr>
<tr class="row-5">
	<td rowspan="3" class="column-1"><strong>Claude</strong></td><td class="column-2">Limited Capabilities/Restrictions</td><td class="column-3">"Usage limits from both the free and paid tiers can be frustrating sometimes."</td>
</tr>
<tr class="row-6">
	<td class="column-2">Performance Issues</td><td class="column-3">“Sometimes the responses are slow, and the website can feel limited when handling long conversations or complex tasks."</td>
</tr>
<tr class="row-7">
	<td class="column-2">Task Limitations</td><td class="column-3">“It does not translate data as well as I’d like. It tends to complicate tasks and I have to check its work.“</td>
</tr>
<tr class="row-8">
	<td rowspan="3" class="column-1"><strong>Gemini</strong></td><td class="column-2">Inaccuracy/Inconsistent Responses</td><td class="column-3">“Sometimes, I have to do some fact checking about information Gemini found."</td>
</tr>
<tr class="row-9">
	<td class="column-2">Not Responding as Prompted</td><td class="column-3">“Occasionally it misinterprets what I'm asking so I figure out a different way to prompt what I need."</td>
</tr>
<tr class="row-10">
	<td class="column-2">Slow Performance/Processing Issues</td><td class="column-3">“I think that I've had a lot of lagging or that it’s slow.“</td>
</tr>
<tr class="row-11">
	<td rowspan="3" class="column-1"><strong>Grok</strong></td><td class="column-2">Inaccuracy/Inconsistent Responses</td><td class="column-3">“Sometimes the bot can get stuck in a loop if it experiences a hallucination."</td>
</tr>
<tr class="row-12">
	<td class="column-2">Slow Performance/Processing Issues</td><td class="column-3">“It takes too long to get answers."</td>
</tr>
<tr class="row-13">
	<td class="column-2">Image/Video Creation/Editing Errors</td><td class="column-3">“Inconsistent results when generating, annoying navigation, user privacy doesn't seem great.“</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1053 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Top issues for AI-based chat software products.</p>
<h2>Do Users of These Products Differ in Tech Savviness?</h2>
<p>We don’t want differences in metrics to just be a result of differences in participant tech-savviness. It could be that less tech-savvy users gravitate toward ChatGPT, and more tech-savvy users prefer Claude or Grok. To differentiate between differences in participant ability and differences in the user experience, we collected tech-savviness scores using our <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10 measure</a> for each of the three products.</p>
<p>The <a href="https://measuringu.com/how-to-use-the-tac/">TAC-10</a> (Technical Activity Checklist with ten items) is a reliable (consistent) and valid (predictive) measure of tech savviness. The TAC-10 score for a person is the number of items selected from its checklist.</p>
<p>In 2026, the TAC-10 scores for the four AI products ranged from 6.3 to 7.0 (see Figure 4), a statistically significant main effect (<em>F</em>(3, 416) = 2.8, <em>p</em> = .04), due to the difference between ChatGPT and Claude (a pattern similar to what we saw in 2025 when the difference was 2.3 points, but less pronounced in 2026). All mean scores were in the lower part of the high range for two-group classification and medium range for three-group classification except for ChatGPT in 2025, which just missed the cut-off of 5 to be in those groups.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/07/Figure4-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47948 size-large" src="https://measuringu.com/wp-content/uploads/2026/07/Figure4-1024x300.png" alt="TAC-10 scores by product and year (with 95% confidence intervals)." width="1024" height="300" srcset="https://measuringu.com/wp-content/uploads/2026/07/Figure4-1024x300.png 1024w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-300x88.png 300w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-768x225.png 768w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-1536x450.png 1536w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-2048x600.png 2048w, https://measuringu.com/wp-content/uploads/2026/07/Figure4-600x176.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> TAC-10 scores by product and year (with 95% confidence intervals).</p>
<h2><span lang="EN-US">Summary and Discussion</span></h2>
<p>Results of our AI-based chat benchmarks based on 420 participants revealed:</p>
<p><strong>ChatGPT was the leader in perceived usability, but just barely.</strong> ChatGPT led in SUS scores in 2026, enjoying a 5.3-point increase from 2025. The range of SUS scores was, however, just 3.1 points (not statistically significant). On the UX-Lite, ChatGPT also led in perceived ease of use and usefulness, though differences among products were not statistically significant.</p>
<p><strong>For Net Promoter Scores, Claude was the leader, and ChatGPT was the laggard.</strong> In 2025, the NPS for Claude was only 4 points higher than ChatGPT, but in 2026 the difference was 21 points in favor of Claude (28% vs. 7% for ChatGPT), with Gemini (12%) and Grok (17%) in between. This reflects Claude&#8217;s recent surge in growth.</p>
<p><strong>Claude showed the biggest gain in perceived usefulness.</strong> Of the products measured in both years, Claude had the largest year-over-year improvement on the UX-Lite usefulness dimension, though it still lags behind ChatGPT on perceived ease of use.</p>
<p><strong>Frequently reported issues included inaccurate responses and limited capabilities.</strong> Respondents reported accuracy issues with ChatGPT, Gemini, and Grok, and criticized Claude for limited capabilities and frustrating usage limits. Other reported problems were slow response time and not responding as prompted.</p>
<p><strong>ChatGPT users had slightly lower tech savviness than Claude users.</strong> The mean TAC-10 scores were all in the medium range of tech savviness. Within that range, there was a statistically significant difference (6.3 for ChatGPT vs. 7.0 for Claude), though the 0.7-point gap on a 10-point scale suggests it may not be practically significant (compared to the 2.3-point difference in 2025).</p>
<p>For more details on these products, see <a href="https://measuringu.com/product/ux-benchmark-report-for-ai-based-chat-software-2026">the full report</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>What Are the Different Types of Synthetic Users?</title>
		<link>https://measuringu.com/what-are-the-different-types-of-synthetic-users/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=what-are-the-different-types-of-synthetic-users</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 23 Jun 2026 21:54:24 +0000</pubDate>
				<category><![CDATA[Survey]]></category>
		<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Synthetic user]]></category>
		<category><![CDATA[Synthetic users]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47761</guid>

					<description><![CDATA[Recruiting participants for research is expensive. It’s also rife with problems: Are these people really who they say they are? Are they actually paying attention? Or is the data from some survey farm where people click through and make money? AI is disrupting UX research. But the disruption is leading to more software, not less. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-scaled.jpg"><img loading="lazy" decoding="async" class="alignleft wp-image-47827 size-medium" src="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-300x169.jpg" alt="Feature image showing 5 different AI bots representing 5 synthetic user types" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-300x169.jpg 300w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-1024x576.jpg 1024w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-768x432.jpg 768w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-1536x864.jpg 1536w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-2048x1152.jpg 2048w, https://measuringu.com/wp-content/uploads/2026/06/062323-FeatureImage-1-600x338.jpg 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>Recruiting participants for research is expensive. It’s also rife with problems: Are these people really who they say they are? Are they actually paying attention? Or is the data from some <a href="https://www.intellisurvey.com/blog/fighting-survey-farms">survey farm</a> where people click through and make money?</p>
<p>AI is disrupting UX research. But the disruption is leading to more software, not less. The need for insights into how people will use that software isn’t going away.</p>
<p>But can AI help? Can we use AI to synthesize people’s attitudes, beliefs, and behaviors? Instead of trying to find the right people to take surveys, could these <strong>synthetic users</strong> generate insights faster and at almost no cost? News about synthetic users would certainly make headlines. <a href="https://measuringu.com/review-of-experiments-with-synthetic-users/">And they do</a>.</p>
<p>But what exactly <em>is</em> a synthetic user? Is that the same as a digital twin? Or a synthetic persona?</p>
<p>To properly assess the effectiveness of AI tools, we think it’s important to have a good understanding of the terms and how they fit together.</p>
<p>In this article, we propose a preliminary taxonomy of five distinct types of synthetic users, organized by how grounded they are in real human data. Before we get to the taxonomy, though, it helps to ask a question that sounds simpler than it is.</p>
<h2>What Birds Can Teach Us about Synthetic Users</h2>
<p>How do we know that a bird is a bird?</p>
<p>Is it because a bird can fly? Well, bats are mammals that can fly, while penguins are birds that can’t fly.</p>
<p>Is it because they lay eggs? Platypuses are mammals that lay eggs (as do most reptiles, amphibians, fish, and arthropods).</p>
<p>Maybe it’s because birds have feathers rather than scales or fur? That might be true in the present, but in the past, many dinosaurs not in the lineage leading to birds are <a href="https://www.smithsonianmag.com/science-nature/dinosaurs-evolved-feathers-for-far-more-than-flight-180985012/">now known to have had feathers</a>.</p>
<p>The answer, as <a href="https://www.linnean.org/learning/who-was-linnaeus/career-and-legacy">Linnaeus worked out in the 1700s</a>, is that no single feature is sufficient. A bird is defined by a <em>constellation</em> of characteristics (feathers, beak, two wings, two feet, warm blood, hard-shelled eggs) organized within a hierarchy of kingdom, class, genus, and species. Figure 1 shows how that plays out from the animal kingdom down to a single species. And this is our guide for classifying and understanding synthetic users.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png" rel="attachment wp-att-47768"><img loading="lazy" decoding="async" class="alignnone wp-image-47831 size-full" src="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png" alt="Figure 2: Classification of different types (species) of synthetic users. " width="2560" height="1349" srcset="https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-scaled.png 2560w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-300x158.png 300w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-1024x540.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-768x405.png 768w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-1536x809.png 1536w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-2048x1079.png 2048w, https://measuringu.com/wp-content/uploads/2026/06/Figure-1-Example-of-classification-600x316.png 600w" sizes="auto, (max-width: 2560px) 100vw, 2560px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Example of classification from the animal kingdom to species of stork.</p>
<h2>Synthetic Users: More of a Genus than a Species</h2>
<p>The topic of classifying different types of synthetic users is in a state of flux (lots of labels, overlapping meanings, vendor-specific definitions). Despite this, in Figure 2, we attempt a preliminary classification scheme similar to Figure 1 for five types (species) under the genus of Synthetic User.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png" rel="attachment wp-att-47769"><img loading="lazy" decoding="async" class="alignnone wp-image-47832 size-full" src="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png" alt="Classification of different types (species) of synthetic users." width="1185" height="868" srcset="https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification.png 1185w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-300x220.png 300w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-1024x750.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-768x563.png 768w, https://measuringu.com/wp-content/uploads/2026/06/Figure-2-Example-of-classification-600x439.png 600w" sizes="auto, (max-width: 1185px) 100vw, 1185px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Classification of different types (species) of synthetic users.</p>
<p>Table 1 lists the identifying characteristics for each of these types of synthetic users, primarily focusing on the type of data used to create the synthetic user and how grounded the synthetic user is in actual user data.</p>

<table id="tablepress-1052" class="tablepress tablepress-id-1052">
<thead>
<tr class="row-1">
	<th class="column-1">Synthetic User Type</th><th class="column-2">Identifying Characteristics/Descriptions</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">AI Proto Persona</td><td class="column-2">This is the weakest (least grounded) type of synthetic user generated with simple role-playing prompts (e.g., “You are a world-class Python programmer”). This method produces preliminary user profiles based on broad assumptions rather than research.</td>
</tr>
<tr class="row-3">
	<td class="column-1">Demographic Based</td><td class="column-2">Prompts specify age, gender, occupation, region, etc. to approximate group-level tendencies. This method is somewhat more grounded than a proto persona but is still limited in the quality of its output, especially when demographics have only weak relationships with research topics (e.g., much UX research).</td>
</tr>
<tr class="row-4">
	<td class="column-1">Persona Based</td><td class="column-2">Prompts focus on richer persona paragraphs (e.g., “Bill G. is a 27-year old male graphic designer who always has his sketchbook at hand, has a track record of being creative and innovative, and is up-to-date on current design trends. How would he complete the following questionnaire?”). Because these synthetic users are still weakly grounded, they are limited to approximate group-level tendencies.</td>
</tr>
<tr class="row-5">
	<td class="column-1">Research Grounded</td><td class="column-2">Prompts refer to actual research artifacts with traceable sources but do not attempt to model individual human responses. These are based on actual interviews, survey results, analytics, customer-support logs, or other user data that are typically not available to publicly generated LLMs.</td>
</tr>
<tr class="row-6">
	<td class="column-1">Digital Twins</td><td class="column-2">Prompts refer to rich individual-level data for the purpose of modeling each individual in a dataset. This approach has the strongest grounding in actual user data but its accuracy in real-world deployments is still an open research question.</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1052 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1</strong>: Brief descriptions of types of synthetic users.</p>
<h2>Discussion</h2>
<p>In this preliminary taxonomy, we’ve defined five types of synthetic users: AI proto persona, demographic based, persona based, research grounded, and digital twins, based on differences in the types of data (e.g., demographic, persona) and the strength of the relationship between the synthetic user and human user data.</p>
<p>Preliminary taxonomies change over time. In this article, we’ve used the levels originally defined by Linnaeus because they were adequate for our purpose. Modern biological taxonomies have eight levels (Domain, Kingdom, Phylum, Class, Order, Family, Genus, Species), and the number of kingdoms has increased to six (Bacteria, Archaea, Protista, Fungi, Plants, Animals).</p>
<p>We fully expect changes to our classification scheme over time, but it’s a start.</p>
<p>For example, we have not included generative agents in this taxonomy because they are qualitatively different from synthetic users that simulate responses and are more like simulated actors, trying to model what people might do over time. This may eventually become its own branch from the genus of synthetic users, separate from the synthetic respondents. Time will tell.</p>
<p>Just like how there are hybrids in the animal kingdom (e.g., mules, ligers), in practice, there may be hybrids of different types of synthetic users. For example, in Bisbee et al.’s “Synthetic Replacements for Human Survey Data? The Perils of Large Language Models” (<a href="https://www.cambridge.org/core/journals/political-analysis/article/synthetic-replacements-for-human-survey-data-the-perils-of-large-language-models/B92267DC26195C7F36E63EA04A47D2FE">2024</a>), the researchers used the following prompt to elicit 30 synthetic responses for each of the 7,350 human respondents for each of the 7,350 human respondents in the 2016–2020 ANES survey to get a final dataset with 3,614,400 responses:</p>
<p style="padding-left: 25px;"><em>It is [YEAR]. You are a [AGE] year-old, [MARST], [RACETH] [GENDER] with [EDUCATION] making [INCOME] per year, living in the United States. You are [IDEO], [REGIS] [PID] who [INTEREST] pays attention to what’s going on in government and politics. Provide responses from this person’s perspective. Use only knowledge about politics that they would have.</em></p>
<p>Each bracketed item is a variable with values corresponding to a real respondent in the two waves of the ANES. For example, [YEAR] was 2016 or 2020, [AGE] matched the selected respondent’s age, [MARST] was marital status (e.g., married, divorced, single), [IDEO] was political ideology from extremely liberal to extremely conservative, [REGIS] was voter registration status, and [PID] was party membership (Democrat, Independent, Republican).</p>
<p>Thus, this is a hybrid between demographic- and persona-based types with a light sprinkle of digital twinning. It’s less than a fully research-grounded respondent or digital twin because the prompt doesn’t include access to the respondent’s prior answers, interview transcript, open-ended comments, voting history, occupation, or religion. It uses selected ANES variables as conditioning attributes and asks the LLM to answer from that perspective (multiple times for each human respondent).</p>
<h2>Summary</h2>
<p>Our key conclusions from this exercise are:</p>
<p><strong>“Synthetic users” is more of an umbrella term (like a genus) than a type (like a species). </strong></p>
<p>All five types we’ve described can fall under the umbrella of synthetic users. In practice, that means when we talk about synthetic users, it’s like talking about storks. There are a variety of storks, so knowing which bird we’re talking about helps move the conversation forward.</p>
<p><strong>Key criteria for discriminating among types of synthetic users include data type and grounding. </strong></p>
<p>The types we’ve defined differ in the kind of data used to model responses (e.g., demographic, persona) and the extent to which they are grounded in real user data.</p>
<p><strong>Taxonomies change over time. </strong></p>
<p>We consider this article a necessary exercise in a preliminary taxonomy of synthetic users, but fully expect it to evolve over time, maybe quickly due to rapid changes in these technologies.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Should a PhD Count as Years of Experience?</title>
		<link>https://measuringu.com/should-a-phd-count-as-years-of-experience/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=should-a-phd-count-as-years-of-experience</link>
		
		<dc:creator><![CDATA[Jeff Sauro&nbsp;•&nbsp;Jim Lewis, PhD]]></dc:creator>
		<pubDate>Wed, 17 Jun 2026 02:29:37 +0000</pubDate>
				<category><![CDATA[UX]]></category>
		<category><![CDATA[PhD]]></category>
		<category><![CDATA[UX Research]]></category>
		<category><![CDATA[work]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47740</guid>

					<description><![CDATA[You put in the grueling hours. You know what it takes to learn, to persevere, to deliver. There’s nothing quite like the experience of a PhD program. But perhaps a career in academia isn’t what you’re looking for. You decide to leave academia and apply your skills in the UX industry. You look for jobs [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-scaled.png"><img loading="lazy" decoding="async" class="alignleft wp-image-47811 size-medium" src="https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-300x169.png" alt="Feature image showing a graduate and an infographic" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-2048x1152.png 2048w, https://measuringu.com/wp-content/uploads/2026/06/061626-FeatureImage1-600x338.png 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>You put in the grueling hours. You know what it takes to learn, to persevere, to deliver.</p>
<p>There’s nothing quite like the experience of a PhD program. But perhaps a career in academia isn’t what you’re looking for.</p>
<p>You decide to leave academia and apply your skills in the UX industry. You look for jobs and see that most require a minimum of three to five years of experience.</p>
<p>The PhD program took five years. Don’t those years in the program count as years of job experience? Should they count?</p>
<p>Probably not.</p>
<h2>The Value of a PhD</h2>
<p>First, we value what a PhD offers. We know. We both went through the experience—while working and raising kids! There’s nothing quite like it. The classwork, teaching, papers, dissertation, defense, and the pile of academic rules.</p>
<p>Averaged over the past decade, about 10% of respondents to the UXPA salary survey <a href="https://uxpa.org/salary-surveys/">had a PhD</a>. That’s seven times higher than the general U.S. population, so it’s a bit of an exclusive club. Although intellectually rewarding, it’s not clear that a PhD pays off financially if entering industry after being a full-time student.</p>
<p>While PhDs do get paid more than their peers with master’s and bachelor’s degrees, the delay before getting into the workforce <a href="https://measuringu.com/does-an-advanced-degree-pay-off/">offsets much of the gain</a>. It takes about 24 years of employment to break even (Figure 1). The value of the PhD to us is not in the clear financial payoff.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47809" src="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1-1024x815.png" alt="Figure 1: “Lifetime” earnings of PhD and master’s degree recipients vs. not having an advanced degree. " width="800" height="637" srcset="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1-1024x815.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1-300x239.png 300w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1-768x611.png 768w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1-600x478.png 600w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure1.png 1226w" sizes="auto, (max-width: 800px) 100vw, 800px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> “Lifetime” earnings of PhD and master’s degree recipients vs. not having an advanced degree.</p>
<h2>PhD Skills Are Transferable</h2>
<p>At MeasuringU, we hire PhDs and probably have one of the highest percentages of PhDs among UX research companies and related fields. (Figure 2).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47810" src="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2-300x234.png" alt="Figure 2: Doctorates at MeasuringU – we like to hire them!" width="800" height="625" srcset="https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2-300x234.png 300w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2-1024x800.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2-768x600.png 768w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2-600x469.png 600w, https://measuringu.com/wp-content/uploads/2026/06/061626-Figure2.png 1184w" sizes="auto, (max-width: 800px) 100vw, 800px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Doctorates at MeasuringU—we like to hire them!</p>
<p>We do this because many of the skills learned in a PhD program are transferable. Table 1 shows examples of the demands of PhD programs and industrial UX research that lead to similar skills.</p>

<table id="tablepress-1050" class="tablepress tablepress-id-1050">
<thead>
<tr class="row-1">
	<th class="column-1">Skill / Demand</th><th class="column-2">PhD Program</th><th class="column-3">Industry UX Research</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">Designing and running studies</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Extensive</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Extensive</td>
</tr>
<tr class="row-3">
	<td class="column-1">Statistical analysis and interpretation</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Rigorous</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Applied</td>
</tr>
<tr class="row-4">
	<td class="column-1">Communicating findings to a skeptical audience</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Committees, advisors</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Stakeholders, leadership</td>
</tr>
<tr class="row-5">
	<td class="column-1">Working under deadlines</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Milestones, defenses</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Sprint cycles, launches</td>
</tr>
<tr class="row-6">
	<td class="column-1">Managing ambiguity in data</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Common</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Constant</td>
</tr>
<tr class="row-7">
	<td class="column-1">Writing clearly and persuasively</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Dissertations, papers</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Reports, readouts</td>
</tr>
<tr class="row-8">
	<td class="column-1">Defending methodology under scrutiny</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Peer review, committee</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Cross-functional critique</td>
</tr>
<tr class="row-9">
	<td class="column-1">Literature synthesis and benchmarking</td><td class="column-2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Deep</td><td class="column-3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2705.png" alt="✅" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Applied</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1050 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Comparisons of skills and demands in PhD programs and industrial UX research.</p>
<h2>Where the Comparison Breaks Down</h2>
<p>Beyond the overlap shown in Table 1, industry experience adds something different: critical understanding of company politics, constraints, and compromises of conducting research inside a business.</p>
<p>PhD programs rarely teach these skills because they’re not primarily what graduate school is for. Years of professional UX experience expose practitioners to these important skills for navigating the workforce:</p>
<p><strong>Communicating with upper management.</strong> Academic communication is oriented toward depth, precision, and scholarly audiences. It’s an entirely different skill to effectively present findings to a vice president who has eight minutes until the next meeting and who has little interest in methods, much less your confidence intervals.</p>
<p><strong>Cost-justifying research.</strong> In industry, research doesn&#8217;t happen because it&#8217;s intellectually interesting. You have to make the business case for why this study is worth three weeks of execution, headcount, and recruiting budget.</p>
<p><strong>Making recommendations. </strong>This is probably one of the harder challenges for fresh PhDs. In academic work, published research is the product where the interpretative challenge is to show how the data support or fail to support competing theories. The publication process has its own timeline, which can span over years. In industrial research, the timeline is much faster, and the interpretive challenge is to make connections between results and decisions. You’re expected to be the expert and to use the data as a guardrail when making specific recommendations.</p>
<p><strong>Making decisions with incomplete data.</strong> Making recommendations is hard, but making them with incomplete, partial data and small sample sizes is even harder, and it&#8217;s pretty much guaranteed to be part of the job. Academic research is built around the idea that you don&#8217;t publish until the evidence is sufficient. Industry doesn&#8217;t wait. Researchers routinely must make actionable calls with <em>n</em> = 6, two weeks of data, and a product manager asking, &#8220;So what do we do?&#8221; Being comfortable with that is a skill that only comes from real-world practice.</p>
<p><strong>Navigating organizational dynamics.</strong> Whose priorities take precedence? What happens when engineering ignores your findings? How do you influence a roadmap when you don&#8217;t own it? These questions don&#8217;t appear on qualifying exams.</p>
<p><strong>Learning to say “no.”</strong> As a student, you’re expected to do everything you’re asked to do without any significant pushback. When you join an industrial team as a UX researcher, you can count on being asked to do more than you can. Learning how to say “no” in socially and politically savvy ways becomes an important skill.</p>
<p><strong>Learning to deal with the “PowerPoint problem.”</strong> Translating nuanced research into a compelling, executive-ready slide deck (one that lands the insight without losing the truth) is a learned, practiced skill. Many new PhDs have never had to do it, and it shows. This is not writing an academic paper where findings and conclusions come after the methodological details. Stakeholders assume you know how to do the work. You need to learn to start with the findings and recommendations and then provide the supporting details.</p>
<p><strong>Being okay with &#8220;good enough&#8221; research.</strong> Graduate training optimizes for rigor. Industry optimizes for decision-making. Sometimes the right study is a fast-and-dirty five-session usability test, not a mixed-methods longitudinal study. Knowing when to go &#8220;good enough&#8221; (and not feeling bad about it) takes time and context that academia doesn&#8217;t provide.</p>
<h2>Do Years of Experience Count Toward a PhD?</h2>
<p>While we’re sympathetic to PhD applicants who want their years of program experience to count as industry experience, should the inverse also hold?</p>
<p>That is, if a PhD counts as five years of industry experience, should five years of UX research experience earn an honorary doctorate?</p>
<p>Most PhDs would immediately bristle at this suggestion, and they&#8217;d be right to do so. Years of industry work, however excellent, don&#8217;t replicate what a doctoral program demands: the years of independent scholarship, the depth of the literature review, the grueling process of designing and defending original research, the personal reckoning of being solely responsible for a body of work over years. Nobody gets a PhD by accident or by accumulating time. It&#8217;s a specific, demanding thing.</p>
<p>But the fact that this inverse is obviously false signals that the original equivalence is weaker than it sounds. Despite some overlapping skills, the demands of the two research contexts are very different.</p>
<h2>The Bottom Line</h2>
<p>A PhD and years of industry experience are not the same thing. They produce overlapping but distinct skill sets, so treating them as interchangeable flattens what is genuinely valuable about each.</p>
<p>But that&#8217;s not the same as saying a PhD doesn&#8217;t matter in industry. For certain roles—particularly in applied research firms where methodological depth, statistical rigor, and the ability to defend findings under scrutiny are core to the work (like MeasuringU!), a PhD is a genuine competitive advantage. It&#8217;s not a substitute for industry experience, but it&#8217;s a strong foundation to build on, and the gap closes quickly for PhDs who are self-aware about what they still need to learn.</p>
<p>If you have a PhD, be proud of it. Use it. The methodological foundation you built is real and hard to replace.</p>
<p>The PhD is a head start on the craft. Experience is a head start on the context. The best industrial researchers, eventually, have both.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Do Statistics Really Require 30 Participants?</title>
		<link>https://measuringu.com/do-statistics-really-require-30-participants/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=do-statistics-really-require-30-participants</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 09 Jun 2026 21:37:45 +0000</pubDate>
				<category><![CDATA[Research]]></category>
		<category><![CDATA[Sample Size]]></category>
		<category><![CDATA[Usability Testing]]></category>
		<category><![CDATA[User Experience]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[confidence interval]]></category>
		<category><![CDATA[Confidence Intervals]]></category>
		<category><![CDATA[t-distribution]]></category>
		<category><![CDATA[t-test]]></category>
		<category><![CDATA[z-distribution]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47696</guid>

					<description><![CDATA[Should the sample size n be greater than 30? If you’ve taken any introductory statistics course or an AP statistics class (or helped your child with it), you’ve encountered the n ≥ 30 rule. The “magic number 5” rule we’ve written extensively about applies (with its important caveats) to problem discovery for usability testing. But the [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-scaled.png"><img loading="lazy" decoding="async" class="alignleft wp-image-47752 size-medium" src="https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-300x169.png" alt="Feature image showing an icon representing participants and &quot;≥&quot;" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-2048x1152.png 2048w, https://measuringu.com/wp-content/uploads/2026/06/060926-FeatureImage1-600x338.png 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>Should the sample size <em>n</em> be greater than 30?</p>
<p>If you’ve taken any introductory statistics course or an AP statistics class (or helped your child with it), you’ve encountered the <em>n</em> ≥ 30 rule.</p>
<p>The “<a href="https://measuringu.com/specific-sample-sizes-in-discovery-studies/">magic number 5</a>” rule we’ve written extensively about applies (with its important caveats) to problem discovery for usability testing.</p>
<p>But the <em>n</em> ≥ 30 rule goes beyond usability testing, coming up across disciplines and even in classrooms. It will often be mentioned by skeptical stakeholders and during the peer review process (probably from <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8505560/">Reviewer Number 2</a>). Violating it can feel like a methodological sin.</p>
<p>But where does it actually come from? And does it hold up in general and in UX research in particular?</p>
<p>The short answer is that the rule has real statistical roots, but they’re often misunderstood and misapplied.</p>
<h2>On One Hand: Arguments for Why Researchers Need <em>n </em>≥ 30</h2>
<p>The <em>n</em> ≥ 30 rule is grounded in two related concerns: whether your statistical analyses will perform accurately with smaller sample sizes (1) when your raw data is normally distributed and (2) when raw data is not normally distributed.</p>
<h3>The <em>t</em>-Distribution Converges to <em>Z</em> at about <em>n</em> = 30 for Continuous Data</h3>
<p>When learning statistics, you’ll often start with the normal <em>z</em>-distribution for statistical tests. You can use tables or simple formulas to look up <em>z</em> values when making computations. But using the normal <em>z</em>-distribution and tables means you need to know the population standard deviation.</p>
<p>Unfortunately, in applied settings, we rarely know the population standard deviation! Fortunately, the alternative is to use <em>t</em>-distribution tables and computations. They have one additional input compared to <em>z</em> computations, which is the sample size (strictly speaking, <em>n</em> − 1, the <a href="https://sites.utexas.edu/sos/degreesfreedom/">degrees of freedom</a> for the <em>t</em>-distributions). However, when the sample size gets to about 30, <em>z</em> and <em>t</em> values converge, so they are roughly the same (see Figure 1).</p>
<p>Consequently, when <em>n</em> is at least 30, you don’t have to deal with somewhat more complicated small-sample statistics. Over time, this statistical footnote calcified into a general-purpose sample size rule.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/06/Figure1-scaled.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47754" src="https://measuringu.com/wp-content/uploads/2026/06/Figure1-300x225.png" alt="Figure 1: Approach of t to z as a function of degrees of freedom (n - 1)." width="800" height="599" srcset="https://measuringu.com/wp-content/uploads/2026/06/Figure1-300x225.png 300w, https://measuringu.com/wp-content/uploads/2026/06/Figure1-1024x767.png 1024w, https://measuringu.com/wp-content/uploads/2026/06/Figure1-768x575.png 768w, https://measuringu.com/wp-content/uploads/2026/06/Figure1-1536x1150.png 1536w, https://measuringu.com/wp-content/uploads/2026/06/Figure1-2048x1534.png 2048w, https://measuringu.com/wp-content/uploads/2026/06/Figure1-600x449.png 600w" sizes="auto, (max-width: 800px) 100vw, 800px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Approach of <em>t</em> to <em>z</em> as a function of degrees of freedom (<em>n</em> − 1).</p>
<h3>Traditional “Wald” Confidence Intervals Are Less Accurate at Small Sample Sizes</h3>
<p>The <em>n</em> ≥ 30 rule isn’t just for continuous data. It’s also been applied to binary (0/1, yes/no) data. Like the convergence of <em>t</em> to <em>z</em> shown in Figure 1, there is a similar convergence of the binomial distribution to the <em>z</em>-distribution, which becomes approximately normal when <em>n</em> = 30 and the expected proportion is not very close to 0 or 1. The most widely taught method for calculating binomial confidence intervals (the Wald method) grossly understates the width of the true interval when sample sizes are small because it’s based on the <em>z</em>-distribution. We demonstrated the inaccuracy of the Wald method in <a href="https://measuringu.com/article/estimating-completion-rates-from-small-samples-using-binomial-confidence-intervals-comparisons-and-recommendations/">our 2005 paper</a> using real-world completion rate data. For example, a 95% confidence interval around a completion rate with <em>n</em> = 15 constructed with the Wald method is more like a 72% confidence interval (wildly inaccurate—see below for other findings from that paper).</p>
<h3>UX Data Generally Isn’t Normally Distributed</h3>
<p>If you’ve ever stared at a histogram of task completion rates, time-on-task, or even SUS scores, you know that raw UX data rarely forms the classic bell curve. Completion rates are binary (0 or 1), usually with more successes than failures. Task times are often-right skewed due to a long tail of slow participants. Likert-scale items such as the Single Ease Question (<a href="https://measuringu.com/evolution-of-seq/">SEQ</a><sup>®</sup>) tend to cluster toward the top (more scores above the median than below it). As we showed in our article <a href="https://measuringu.com/is-ux-data-normal/">Is UX Data Normally Distributed?</a>, none of these distributions look remotely normal (see Figure 2).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions.png" rel="attachment wp-att-47698"><img loading="lazy" decoding="async" class="alignnone wp-image-47698 size-full" src="https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions.png" alt="Figure 2: Distributions of four UX metrics showing their non-normal raw distributions. " width="1041" height="1100" srcset="https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions.png 1041w, https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions-284x300.png 284w, https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions-969x1024.png 969w, https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions-768x812.png 768w, https://measuringu.com/wp-content/uploads/2026/05/Figure-2-Distributions-of-four-UX-metrics-showing-their-non-normal-raw-distributions-600x634.png 600w" sizes="auto, (max-width: 1041px) 100vw, 1041px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Distributions of four UX metrics showing their non-normal raw distributions.</p>
<p>That might sound alarming. Many statistical tests, such as confidence intervals, <em>t</em>-tests, and ANOVA, assume normality. If the raw data isn’t normal, are those analyses invalid?</p>
<h2>On the Other Hand: Arguments Why Sometimes <em>n</em> &lt; 30 Is OK</h2>
<p>Fortunately, neither the convergence of <em>t</em> to <em>z</em> nor the non-normality of raw data means confidence intervals or statistical comparison tests are invalid below <em>n</em> = 30. For continuous data, the <em>t</em>-distribution works correctly at any sample size. For binary data, the standard-Wald method can be replaced with better-performing alternatives. And for the normality concern, the Central Limit Theorem (CLT) means the distribution of your raw data matters far less than most researchers assume. This is where the distinction between the distribution of your <em>raw data</em> and the distribution of your <em>sample means</em> becomes critical. We’ll start with the normality issue.</p>
<h3>The Central Limit Theorem Solves Most Normality Issues</h3>
<p>One of the most important concepts in all of statistics is the CLT. According to the CLT, as the sample size increases, the distribution of the mean becomes more and more normal, <a href="https://measuringu.com/is-ux-data-normal/">regardless of the normality of the underlying distribution</a>.</p>
<p>How quickly does the CLT kick in? Considering a wide variety of distributions, most achieve a normal or near-normal distribution of the means when <em>n</em> is 30, making that a reasonably safe bet when you don’t know anything about the distributions.</p>
<p>For many common UX metrics, however, the distributions of the means approach normality sooner than you’d expect. Using bootstrap simulations on real UX datasets (repeatedly drawing sub-samples and computing means), the sampling distributions of SEQ ratings and SUS scores approach normality by <em>n</em> = 10. Even binomial completion rates (which are maximally non-normal) and completion times approach a normal sampling distribution by <em>n</em> = 20–30 (see Figure 3). But it turns out there’s a fix for even smaller binomial sample sizes (see below).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes.png" rel="attachment wp-att-47699"><img loading="lazy" decoding="async" class="alignnone wp-image-47699 size-full" src="https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes.png" alt="Figure 3: Distributions of the means for four UX metrics with varying sample sizes." width="1185" height="1445" srcset="https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes.png 1185w, https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes-246x300.png 246w, https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes-840x1024.png 840w, https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes-768x937.png 768w, https://measuringu.com/wp-content/uploads/2026/05/Figure-3-Distributions-of-the-means-for-four-UX-metrics-with-varying-sample-sizes-600x732.png 600w" sizes="auto, (max-width: 1185px) 100vw, 1185px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 3:</strong> Distributions of the means for four UX metrics with varying sample sizes.</p>
<p>This is why means, <em>t</em>-tests, and confidence intervals work reasonably well even when your raw responses are skewed (e.g., completion times) or bounded (e.g., rating scales).</p>
<h3>The <em>t</em>-Distribution Was Built for Small Samples</h3>
<p>The deep irony of being advised that “you need 30 to use a <em>t</em>-test” is that the <em>t</em>-distribution was invented specifically for small samples.</p>
<p>In 1899, William S. Gossett (Figure 4), a recent graduate of New College, Oxford with degrees in chemistry and mathematics, became one of the first scientists to join the Guinness brewery.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett.jpg" rel="attachment wp-att-47700"><img loading="lazy" decoding="async" class="alignnone wp-image-47700 size-full" src="https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett.jpg" alt="Figure 4: An anachronistic interpretation of William S. Gossett (Student), adapted from his Wikipedia page photo (public domain) with AI assistance." width="837" height="1075" srcset="https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett.jpg 837w, https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett-234x300.jpg 234w, https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett-797x1024.jpg 797w, https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett-768x986.jpg 768w, https://measuringu.com/wp-content/uploads/2026/05/Figure-4-An-anachronistic-interpretation-of-William-S.-Gossett-600x771.jpg 600w" sizes="auto, (max-width: 837px) 100vw, 837px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 4:</strong> An anachronistic interpretation of William S. Gossett (Student), adapted from his Wikipedia page photo (public domain) with AI assistance.</p>
<p>As Michael Cowles wrote in his book <em>Statistics in Psychology: An Historical Perspective</em>, “Compared with the giants of his day, he published very little, but his contribution is of critical importance. … The nature of the process of brewing, with its variability in temperature and ingredients, means that it is not possible to take large samples over a long run” (pp. 108–109).</p>
<p>Gossett couldn’t use <em>z</em>-scores because they don’t perform well with small samples. After analyzing the deficiencies of the <em>z</em>-distribution for small-sample statistical tests, he worked out the necessary adjustments as a function of degrees of freedom to produce the <em>t</em>-distribution—published in 1908 under the pseudonym “Student,” because Guinness prohibited employees from publishing. In the work that led to his tables, Gossett performed an early version of <a href="https://en.wikipedia.org/wiki/Monte_Carlo_method">Monte Carlo simulations</a>. He prepared 3,000 cards labeled with physical measurements taken on criminals, shuffled them, then dealt them into 750 groups of size 4 (<em>n</em> much smaller than 30).</p>
<p>The point is that the <em>t</em>-distribution was designed precisely to handle small samples correctly. The idea that you need <em>n</em> ≥ 30 even when using the <em>t</em>-distribution contradicts both the history and purpose of the statistic.</p>
<p>Historians of statistics widely regard Gossett&#8217;s publication of Student&#8217;s <em>t</em>-test as a landmark event. In a letter to <a href="https://en.wikipedia.org/wiki/Ronald_Fisher">Ronald A. Fisher</a> containing an early copy of the <em>t</em>-tables, Gossett wrote, &#8220;<a href="https://www.physoc.org/magazine-articles/the-strange-origins-of-the-students-t-test/">You are the only man that&#8217;s ever likely to use them</a>.&#8221;</p>
<p>Gossett got a lot of things right. He certainly got that wrong.</p>
<h3>The Wald Method Can Be Fixed with a Simple Adjustment</h3>
<p>While binary data generates inaccurate confidence intervals with the standard-Wald method when <em>n</em> &lt; 20–30, there&#8217;s a fix. It turns out that a slight adjustment makes even small binomial samples generate accurate confidence intervals. In the same <a href="https://math.unm.edu/~james/Agresti1998.pdf">2005 paper</a> where we demonstrated the problem with the Wald method, we also showed that a simple adjustment brings 95% confidence intervals back to accurate coverage. The adjustment is very close to just adding two successes and two failures to your observed data, then computing a standard-Wald interval around that adjusted proportion. This is now often called the adjusted-Wald method, formalized by work by Agresti and Coull. A small tweak to the math turns an unreliable method into a reliable one, even with small samples.</p>
<h2>Our Recommendation: Stats Work Below <em>n</em> = 30 with the Right Approach</h2>
<p>When the cost of sampling is high, as it often is, always insisting on at least 30 users regardless of study goals is wasteful at best and infeasible at worst. A more appropriate approach is to use sample size formulas derived from the specific statistical analysis (e.g., confidence interval estimation, significance test), accounting for the data type, expected variability, desired confidence, and target effect size. We’ve published articles on this for several common UX scenarios, for example:</p>
<ul>
<li><a href="https://measuringu.com/ux-lite-sample-sizes-for-confidence-intervals/">UX-Lite Sample Sizes for Confidence Intervals</a></li>
<li><a href="https://measuringu.com/ux-lite-sample-sizes-for-comparison-to-a-benchmark/">UX-Lite Sample Sizes for Comparison to a Benchmark</a></li>
<li><a href="https://measuringu.com/sample-sizes-for-comparison-of-ux-lite-scores/">Sample Sizes for Comparing UX-Lite Scores</a></li>
</ul>
<p>Knowing that small samples <em>can</em> work statistically doesn&#8217;t tell you how to handle them. The right approach depends on the type of data you&#8217;re analyzing:</p>
<ul>
<li><strong>Rating scales</strong> (SUS, SEQ, SUPR-Q): Use the <em>t</em>-distribution with the correct degrees of freedom for confidence intervals and tests of significance. It was designed for exactly this situation.</li>
<li><strong>Binary data</strong> (completion rates, yes/no): Use adjusted methods (e.g., <a href="https://measuringu.com/article/estimating-completion-rates-from-small-samples-using-binomial-confidence-intervals-comparisons-and-recommendations/">adjusted-Wald</a> for confidence intervals, <a href="https://measuringu.com/what-is-the-n-1-two-proportion-test/"><em>N</em>−1 two-proportion</a> method for significance tests), which perform accurately at small samples where standard methods based on the <em>z</em>-distribution break down.</li>
<li><strong>Time data</strong>: For confidence intervals, log-transform the raw data to correct for right-skew, then transform back to the original scale (usually no need to transform for tests of significance, but it is always an option).</li>
</ul>
<p>When samples are small, the concern shifts from normality and sampling distributions to sensitivity and power—the accuracy of an estimate (confidence intervals) or your ability to detect a true difference when one exists (hypothesis testing).</p>
<p>The procedures work correctly with small samples, but you&#8217;re limited to relatively imprecise estimates (confidence intervals) or reliably detecting only large differences (hypothesis tests). Subtle or moderate effects will likely go undetected, not because the statistics are broken but because small samples carry more uncertainty. We cover how to plan for adequate sensitivity and power in all our articles on sample size estimation (e.g., <a href="https://measuringu.com/sample-sizes-for-rating-scale-comparisons/">Sample Sizes for Comparing Rating Scale Means</a>). Sometimes a small sample is all you need to achieve your research goal.</p>
<p>This controversy is similar to the &#8220;magic number 5&#8221; controversy but applied to <a href="https://measuringu.com/three-goals/">summative rather than formative</a> research. The &#8220;magic number 30&#8221; has real empirical rationale, as it&#8217;s roughly where the CLT kicks in for a wide variety of distributions and where <em>t</em> converges on <em>z</em>. In practice, however, it&#8217;s applied far too rigidly. The appropriate sample size depends on the  distribution, the expected variability, the desired confidence and power, and the minimum <a href="https://measuringu.com/an-introduction-to-effect-sizes/">effect size</a> you need to detect. A sample of 30 is almost never exactly right for any specific situation.</p>
<p>It isn&#8217;t much more complicated to use the <em>t</em>-distribution than the <em>z</em>-distribution, or the adjusted-Wald instead of the Wald method (you just need to account for the sample size). The entire reason the <em>t</em>-distribution was developed was to enable the analysis of small samples. This is just one of the less obvious ways usability practitioners benefit from the science and practice of beer brewing.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Using the TAC-10 for Screening and Data Cleaning</title>
		<link>https://measuringu.com/using-the-tac10-for-screening-and-data-cleaning/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=using-the-tac10-for-screening-and-data-cleaning</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 02 Jun 2026 23:09:08 +0000</pubDate>
				<category><![CDATA[Methods]]></category>
		<category><![CDATA[User Research]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Online Panels]]></category>
		<category><![CDATA[TAC-10]]></category>
		<category><![CDATA[Tech savviness]]></category>
		<category><![CDATA[UX Research]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47676</guid>

					<description><![CDATA[It’s hard to collect data for UX research, and once you have it, you have to clean it. In a simpler world, all respondents would be honest and focused on providing high-quality information rather than maximizing income, but that’s not the world we live in. From past research, we estimate the prevalence of cheating on [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1.png" rel="attachment wp-att-24969"><img loading="lazy" decoding="async" class="alignleft wp-image-47712 size-medium" src="https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-300x169.png" alt="Feature image showing TAC-10 being used for screening and data cleaning" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1-600x338.png 600w, https://measuringu.com/wp-content/uploads/2026/05/060226-FeatureImage-1.png 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>It’s hard to collect data for UX research, and once you have it, you have to clean it.</p>
<p>In a simpler world, all respondents would be honest and focused on providing high-quality information rather than maximizing income, but that’s not the world we live in. From <a href="https://measuringu.com/cheat-survey/">past research</a>, we estimate the prevalence of cheating on paid panels to be about 10% of respondents (ranging from 3–20%).</p>
<p>UX researchers can use <a href="https://measuringu.com/cleaning-data/">numerous strategies</a> for screening (stopping bad actors before they get to the actual study) and cleaning (finding and removing poor quality respondents after study completion). These include:</p>
<ul>
<li>Identification of speeders</li>
<li>Disqualifying questions</li>
<li>Attention checks</li>
<li>Review of open-ended responses</li>
<li>Internal consistency</li>
<li>Straightlining</li>
<li>Review of session recordings (when available)</li>
<li>Duplicate and bot detection</li>
</ul>
<p>AI complicates all these approaches. Modern AI can mimic attentive respondent behavior well enough to slip past most of these detection methods. We are encouraged, however, that many panel operators have taken active steps to restrict AI fraud at the source.</p>
<p>When those safeguards are in place, or when participants come from a verified human population such as a customer list, we propose another dual-purpose and quick approach.</p>
<p>In this article, we demonstrate how to use TAC-10<img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2122.png" alt="™" class="wp-smiley" style="height: 1em; max-height: 1em;" /> response patterns not only for its primary purpose as a measure of tech savviness but also as a type of internal consistency check for screening and detecting inattentive or misrepresenting human respondents.</p>
<h2>TAC-10 Basics</h2>
<p>In a <a href="https://measuringu.com/how-to-use-the-tac/">series of articles</a>, we reviewed the findings of eight years of research into measuring tech savviness. In that research program, we explored <a href="https://measuringu.com/in-search-of-tech-savvy-measures/">several methods</a> for measuring tech savviness, including quizzes (what people know), self-assessment questionnaires (what people feel), and technical activity checklists (what people are confident doing).</p>
<p>After analyzing thousands of participants’ data to see how measures of tech savviness predict task performance, we determined that technical activity checklists had better measurement properties than quizzes or questionnaires. Of the various versions of checklists that we studied, we determined that the one with ten activities and a none-of-the-above option (the TAC-10 shown in Figure 1) has the best balance between conciseness and completeness.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/1.png" rel="attachment wp-att-47677"><img loading="lazy" decoding="async" class="alignnone wp-image-47677 size-full" src="https://measuringu.com/wp-content/uploads/2026/05/1.png" alt="Figure 1: The current version of the TAC-10 (image from the MUiQ® platform). " width="683" height="625" srcset="https://measuringu.com/wp-content/uploads/2026/05/1.png 683w, https://measuringu.com/wp-content/uploads/2026/05/1-300x275.png 300w, https://measuringu.com/wp-content/uploads/2026/05/1-600x549.png 600w" sizes="auto, (max-width: 683px) 100vw, 683px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> The current version of the TAC-10 (image from the <a href="https://measuringu.com/muiq/">MUiQ<sup>®</sup> platform</a>).</p>
<p>The TAC-10 score for a person is the number of selected items. It’s a reliable (consistent) and valid (predictive) measure of tech savviness that, for its primary purpose, can be used (1) to classify participants into groups with low, medium, or high levels of tech savviness and (2) as a tech savviness predictor or covariate in advanced statistical analysis.</p>
<h2>Some TAC-10 Response Patterns Are More Plausible than Others</h2>
<p>In addition to being a tool to measure and classify tech savviness, TAC-10 response patterns can also be used to identify potentially problematic respondents based on the plausibility of the pattern.</p>
<p>In May 2023, we collected a large sample of completed TAC-16 checklists as part of a screening survey (<em>n </em>= 4,731) to acquire enough data for <a href="https://measuringu.com/rasch-analysis-of-three-technical-activity-checklists/">Rasch analysis</a> of three versions of the TAC (TAC-9, TAC-10, and TAC-16). In this new analysis, we applied various methods to classify response patterns as plausible, implausible, or indeterminate. Examples of plausible patterns are those consistent with perfect or near-perfect <a href="https://en.wikipedia.org/wiki/Guttman_scale">Guttman scaling</a>. Response patterns that are logically inconsistent are implausible. Patterns that are not clearly plausible or implausible are indeterminate. For these analyses, we coded each TAC-10 response as a binary string of 0s (not selected) and 1s (selected) for activities in the order shown in Figure 1. For example, 1100000000 indicates a user who selected &#8220;installing a new app on your phone&#8221; and &#8220;setting up a new phone,&#8221; but no other activities.</p>
<h3>Responses Consistent with Guttman Scaling Are Plausible</h3>
<p>Guttman scaling, which dates back to the 1940s, is a deterministic predecessor of probabilistic Rasch scaling. The goal of a Guttman scale is to develop a set of distinctive items, from easy to difficult, that form a unidimensional scale. The range from easy to difficult can refer to characteristics like easy to solve to difficult to solve for math problems or easy to agree with to difficult to agree with for attitudinal scales.</p>
<p>For 10 binary (yes/no) items like the TAC-10, there are 2<sup>10</sup> (1,024) possible arrangements of selected (1) or unselected (0) items. Only 11, however, are consistent with a perfect Guttman scale (all 1s toward the left side of the pattern, all 0s to the right): 0000000000, 1000000000, 1100000000, 1110000000, 1111000000, 1111100000, 1111110000, 1111111000, 1111111100, 1111111110, and 1111111111. We categorized these patterns as plausible.</p>
<h3>Other Plausible Response Patterns</h3>
<p>In practice, other patterns that are close to Guttman patterns are also likely to be plausible. For example, if someone is comfortable with all activities except HTML, the pattern would be 1111111101. Although it’s unlikely that someone who programs efficiently in C knows nothing about HTML, it’s possible that they would lack sufficient practical or deep familiarity with it to be comfortable selecting it. In most cases, Guttman-like patterns with one or two discontinuities are plausible.</p>
<h3>Implausible Response Patterns</h3>
<p>Patterns that are the inverse of Guttman patterns (1 and 0 swapped) are categorized as implausible (except for 0000000000 and 1111111111). For example, the pattern 0000000001 indicates someone who programs efficiently but isn’t comfortable with anything else in the list—possible but highly unlikely.</p>
<p>Other problematic patterns are those that start with 01 because, for this to be plausible, the respondent would have to be comfortable setting up a new phone but uncomfortable adding an app to that phone.</p>
<p>Patterns that contain just one 1 (other than 100000000) are implausible and may indicate a respondent who misunderstood the instruction to select all that apply.</p>
<h3>Indeterminate Response Patterns</h3>
<p>Patterns not categorized as plausible or implausible are provisionally defined as indeterminate.</p>
<h2>Plausible TAC-10 Patterns Are Much More Likely in Practice than Implausible Patterns</h2>
<p>We investigated the frequency of occurrence of plausible, implausible, and indeterminate patterns in our large sample of TAC-10 scores. Of the 1,024 possible patterns, only 199 appeared at least once in our dataset of 4,731 cases.</p>
<h3>Guttman Patterns</h3>
<p>Table 1 shows the frequency of Guttman patterns in the large TAC-10 database, accounting for 56.4% of cases.</p>

<table id="tablepress-1048" class="tablepress tablepress-id-1048">
<thead>
<tr class="row-1">
	<th class="column-1">Guttman Patterns</th><th class="column-2">Freq</th><th class="column-3">Percent</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">1111111111</td><td class="column-2">365</td><td class="column-3"> 7.7%</td>
</tr>
<tr class="row-3">
	<td class="column-1">1111111110</td><td class="column-2">633</td><td class="column-3">13.4%</td>
</tr>
<tr class="row-4">
	<td class="column-1">1111111100</td><td class="column-2">764</td><td class="column-3">16.1%</td>
</tr>
<tr class="row-5">
	<td class="column-1">1111111000</td><td class="column-2">474</td><td class="column-3">10.0%</td>
</tr>
<tr class="row-6">
	<td class="column-1">1111110000</td><td class="column-2">268</td><td class="column-3"> 5.7%</td>
</tr>
<tr class="row-7">
	<td class="column-1">1111100000</td><td class="column-2">104</td><td class="column-3"> 2.2%</td>
</tr>
<tr class="row-8">
	<td class="column-1">1111000000</td><td class="column-2"> 27</td><td class="column-3"> 0.6%</td>
</tr>
<tr class="row-9">
	<td class="column-1">1110000000</td><td class="column-2"> 20</td><td class="column-3"> 0.4%</td>
</tr>
<tr class="row-10">
	<td class="column-1">1100000000</td><td class="column-2">  9</td><td class="column-3"> 0.2%</td>
</tr>
<tr class="row-11">
	<td class="column-1">1000000000</td><td class="column-2">  5</td><td class="column-3"> 0.1%</td>
</tr>
<tr class="row-12">
	<td class="column-1">0000000000</td><td class="column-2">  0</td><td class="column-3"> 0.0%</td>
</tr>
<tr class="row-13">
	<td class="column-1"></td><td class="column-2"></td><td class="column-3"></td>
</tr>
<tr class="row-14">
	<td class="column-1">Total</td><td class="column-2">2669</td><td class="column-3">56.4%</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1048 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Distribution of Guttman patterns in the large database.</p>
<h3>Other Plausible Patterns</h3>
<p>Table 2 shows other frequently occurring plausible patterns (each accounting for at least 0.5% of cases in the database). The 21 patterns in the table accounted for 30.7% of cases. In combination, the percentage of the 32 Guttman and other high-frequency plausible patterns in the database is 87.1%.</p>

<table id="tablepress-1049" class="tablepress tablepress-id-1049">
<thead>
<tr class="row-1">
	<th class="column-1">Other Plausible Patterns</th><th class="column-2"> Freq</th><th class="column-3">Percent</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">1111110100</td><td class="column-2"> 271</td><td class="column-3"> 5.7%</td>
</tr>
<tr class="row-3">
	<td class="column-1">1111111010</td><td class="column-2"> 179</td><td class="column-3"> 3.8%</td>
</tr>
<tr class="row-4">
	<td class="column-1">1111101000</td><td class="column-2"> 152</td><td class="column-3"> 3.2%</td>
</tr>
<tr class="row-5">
	<td class="column-1">1111110110</td><td class="column-2"> 127</td><td class="column-3"> 2.7%</td>
</tr>
<tr class="row-6">
	<td class="column-1">1111010100</td><td class="column-2"> 104</td><td class="column-3"> 2.2%</td>
</tr>
<tr class="row-7">
	<td class="column-1">1111010000</td><td class="column-2">  99</td><td class="column-3"> 2.1%</td>
</tr>
<tr class="row-8">
	<td class="column-1">1111110010</td><td class="column-2">  72</td><td class="column-3"> 1.5%</td>
</tr>
<tr class="row-9">
	<td class="column-1">1111101100</td><td class="column-2">  66</td><td class="column-3"> 1.4%</td>
</tr>
<tr class="row-10">
	<td class="column-1">1111011100</td><td class="column-2">  52</td><td class="column-3"> 1.1%</td>
</tr>
<tr class="row-11">
	<td class="column-1">1111010110</td><td class="column-2">  44</td><td class="column-3"> 0.9%</td>
</tr>
<tr class="row-12">
	<td class="column-1">1111111101</td><td class="column-2">  32</td><td class="column-3"> 0.7%</td>
</tr>
<tr class="row-13">
	<td class="column-1">1110100000</td><td class="column-2">  30</td><td class="column-3"> 0.6%</td>
</tr>
<tr class="row-14">
	<td class="column-1">1111111011</td><td class="column-2">  30</td><td class="column-3"> 0.6%</td>
</tr>
<tr class="row-15">
	<td class="column-1">1111100100</td><td class="column-2">  28</td><td class="column-3"> 0.6%</td>
</tr>
<tr class="row-16">
	<td class="column-1">1111011110</td><td class="column-2">  26</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-17">
	<td class="column-1">1111101010</td><td class="column-2">  26</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-18">
	<td class="column-1">1111011000</td><td class="column-2">  25</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-19">
	<td class="column-1">1110101000</td><td class="column-2">  23</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-20">
	<td class="column-1">1100100000</td><td class="column-2">  22</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-21">
	<td class="column-1">1101100000</td><td class="column-2">  22</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-22">
	<td class="column-1">1110111000</td><td class="column-2">  22</td><td class="column-3"> 0.5%</td>
</tr>
<tr class="row-23">
	<td class="column-1"></td><td class="column-2"></td><td class="column-3"></td>
</tr>
<tr class="row-24">
	<td class="column-1">Total</td><td class="column-2">1452</td><td class="column-3">30.7%</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1049 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2:</strong> Distribution of other patterns in the large database that are plausible and had frequencies of at least 0.5%.</p>
<h3>Implausible Patterns</h3>
<p>The database did not contain any cases matching an inverse Guttman pattern.</p>
<p>There were 17 implausible patterns that started with 01, each having a frequency of 1 or 2 for a total of 21, accounting for just 0.4% of the data.</p>
<p>There were only four cases (0.1% of the data) in which a single activity was chosen past the phone activities (three cases with 0010000000 and one with 0001000000, two additional implausible patterns).</p>
<h3>Indeterminate Patterns</h3>
<p>Because there were 32 plausible and 19 implausible patterns (51) out of a total of 199 patterns, the remaining 148 patterns are indeterminate.</p>
<p>Combined, the indeterminate patterns account for 12.4% of the data, with no individual indeterminate case having a frequency greater than 0.4%.</p>
<h2>Summary and Discussion</h2>
<p>In addition to its use as a measure of tech savviness, we investigated how well the TAC-10 might be used to identify plausible and implausible response patterns for the purpose of identifying potentially problematic respondents in screening and data cleaning.</p>
<p>Based on our large database of TAC-10 cases (<em>n</em> = 4,731), using two criteria for identifying plausible response patterns (matching Guttman patterns and/or frequently occurring patterns), we found that 56.4% of cases matched Guttman patterns; an additional 21 frequently occurring patterns that slightly deviated from Guttman patterns accounted for 30.7%, for a total of 87.1%. Clearly implausible patterns accounted for only 0.5% of cases, leaving the others indeterminate.</p>
<p>Our key conclusions from these analyses were:</p>
<p><strong>Plausible patterns made up the vast majority (87%) of TAC-10 cases. </strong>This suggests that most respondents were attending to the items rather than carelessly checking boxes, especially because we randomized the order of presentation of the items.</p>
<p><strong>Implausible patterns were rare. </strong>There were no occurrences of inverse Guttman patterns, and less than 0.5% of the of the responses had a problematic pattern starting with 01 or containing a single 1 (aside from a single 1 for the easiest activity).</p>
<p><strong>TAC-10 responses can be used for screening and data cleaning. </strong>These results (a large percentage of plausible and a low percentage of implausible response patterns) are encouraging regarding the application of TAC-10 to identify potentially problematic (implausible or indeterminate) response patterns as part of a battery of approaches used to identify potential cheaters (along with other strategies such as examination of open-ended responses, implausible completion times, distractors in multiple choice items, attention checks, and straightlining).</p>
<p><strong>Not going to solve AI fraud</strong>: We don’t think the TAC-10 is necessarily a solution to AI fraud. More sophisticated AI methods can convincingly mimic either a low- or high-skilled human respondent, possibly by training on the articles we’ve published on the TAC-10. However, the TAC-10 remains a valuable screening tool in contexts where respondents come from a known population, such as a customer list, or where other panel-level methods have already confirmed that participants are human.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Does AI Find Real UI Problems or Just Hallucinations?</title>
		<link>https://measuringu.com/does-ai-find-real-ui-problems-or-just-hallucinations/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=does-ai-find-real-ui-problems-or-just-hallucinations</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD •&nbsp;Jeff Sauro, PhD •&nbsp;Will Schiavone, PhD&nbsp;•&nbsp;Lucas Plabst, PhD]]></dc:creator>
		<pubDate>Wed, 27 May 2026 00:41:28 +0000</pubDate>
				<category><![CDATA[Usability]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Gemini]]></category>
		<category><![CDATA[Problem Discovery]]></category>
		<category><![CDATA[Usability Problem]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47650</guid>

					<description><![CDATA[In a previous experiment, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session. That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct. But in our analysis of that video, AI generated nearly [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-scaled.png"><img loading="lazy" decoding="async" class="alignleft wp-image-47672 size-medium" src="https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-300x169.png" alt="Feature image showing an AI robot, three documents each labeled &quot;Verified&quot;, &quot;Fake&quot;, or &quot;False alarm&quot;, and a researcher." width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-2048x1152.png 2048w, https://measuringu.com/wp-content/uploads/2026/05/052626-FeatureImage-600x338.png 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>In a <a href="https://measuringu.com/ai-vs-human-usability-problem-analysis-of-a-video/">previous experiment</a>, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session.</p>
<p>That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct.</p>
<p>But in our analysis of that video, AI generated nearly as many <em>additional</em> problems that humans never flagged. Are these problems hidden gems missed by multiple researchers, or just AI hallucinations?</p>
<p>For this article, we classified all the unique problems the AIs generated into one of three categories:</p>
<ol>
<li>a real problem humans missed</li>
<li>a false alarm (a true observation misread as a usability problem)</li>
<li>a hallucination (something the AI reported that simply never happened)</li>
</ol>
<p>What we found suggests that the new AI problems are mostly false alarms, but there are some notable exceptions.</p>
<h2>Experimental Design: Four Researchers, Two LLMs, and One Video</h2>
<p>For this study, we had four humans (professional UX researchers working at MeasuringU) review a video from a previous usability benchmark study of online dining reservation websites. Each researcher independently created a list of the usability issues they observed in the six-minute video.</p>
<p>We ran the video through two LLMs (ChatGPT-5.4 Thinking and Gemini 3 Flash Thinking) four times <a href="https://measuringu.com/ai-usability-problem-analysis-of-a-video/">using the same prompt</a> each time.</p>
<p>So in this study, we held constant the video, the key elements of the prompt, and the LLM versions/settings—variables that we plan to vary in future studies. This time, we varied only the type of analyst: human, ChatGPT, and Gemini.</p>
<h2>Gemini Finds a Jewel; ChatGPT Goes on a Tangent</h2>
<p>Using the human-generated and -verified problem lists as the “gold standard,” Figure 1 shows a summary of what we found.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/Figure2.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47542" src="https://measuringu.com/wp-content/uploads/2026/05/Figure2-1024x901.png" alt="Venn diagram of usability problem discovery by humans, ChatGPT, and Gemini." width="700" height="616" srcset="https://measuringu.com/wp-content/uploads/2026/05/Figure2-1024x901.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/Figure2-300x264.png 300w, https://measuringu.com/wp-content/uploads/2026/05/Figure2-768x676.png 768w, https://measuringu.com/wp-content/uploads/2026/05/Figure2-1536x1351.png 1536w, https://measuringu.com/wp-content/uploads/2026/05/Figure2-2048x1802.png 2048w, https://measuringu.com/wp-content/uploads/2026/05/Figure2-600x528.png 600w" sizes="auto, (max-width: 700px) 100vw, 700px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1: </strong>Venn diagram of usability problem discovery by humans, ChatGPT, and Gemini.</p>
<p>We know Venn diagrams can generate some bad high school math memories, so here’s a summary for all our sanity:</p>
<ul>
<li>Four human researchers found nine problems (3 + 2 + 3 + 1).</li>
<li>Two AIs combined found 14 problems (6 + 1 + 3 + 4).</li>
<li>Only three problems were found by researchers and both AIs (the 3 in the middle of the circles).</li>
<li>ChatGPT matched five of the nine researcher-identified problems (3 + 2).</li>
<li>Gemini matched four (3 + 1) of the researcher-identified problems.</li>
<li>That leaves 11 problems the AIs flagged that no researcher identified (6 + 1 + 4).</li>
<li>Of those 11 problems, six were unique to ChatGPT, four were unique to Gemini, and one was identified by both AIs.</li>
</ul>
<p>So, <strong>AIs generated 11 new problems</strong> not identified by any of the human researchers. Table 1 has details of those 11 problems, listed in chronological order using problem number codes from the previous article. Of the 11 problems no human flagged, one was a genuine find, seven were false alarms, and three were hallucinations. Here&#8217;s more detail about each category.</p>

<table id="tablepress-1047" class="tablepress tablepress-id-1047">
<thead>
<tr class="row-1">
	<th class="column-1">Prob #</th><th class="column-2">Problem Description</th><th class="column-3">Source</th><th class="column-4">Classification</th>
</tr>
</thead>
<tbody class="row-striping">
<tr class="row-2">
	<td class="column-1"><strong>4b</strong></td><td class="column-2">Filters not helpful</td><td class="column-3">ChatGPT</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-3">
	<td class="column-1"><strong>5b</strong></td><td class="column-2">Participant used Ctrl-F to search for "sushi" when it wasn't in the 86-cuisine list</td><td class="column-3">Gemini</td><td class="column-4"><strong>Genuine find</strong></td>
</tr>
<tr class="row-4">
	<td class="column-1"><strong>6b</strong></td><td class="column-2">Search results for sushi included many non-sushi restaurants</td><td class="column-3">ChatGPT</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-5">
	<td class="column-1"><strong>6c</strong></td><td class="column-2">Weak presentation of cuisine information in search results</td><td class="column-3">ChatGPT</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-6">
	<td class="column-1"><strong>7b-Gem</strong></td><td class="column-2">Participant chose the highest price tier</td><td class="column-3">Gemini</td><td class="column-4"><strong><font color="red">Hallucination</font></strong></td>
</tr>
<tr class="row-7">
	<td class="column-1"><strong>7b-GPT</strong></td><td class="column-2">Sorting by highest rated surfaced many non-sushi restaurants</td><td class="column-3">ChatGPT</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-8">
	<td class="column-1"><strong>8b</strong></td><td class="column-2">UI pushes browsing without good decision support</td><td class="column-3">ChatGPT</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-9">
	<td class="column-1"><strong>9b</strong></td><td class="column-2">Seating options only presented after selecting reservation time</td><td class="column-3">Gemini</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-10">
	<td class="column-1"><strong>9c</strong></td><td class="column-2">Participant set time to 5:10 instead of 5:00</td><td class="column-3">Gemini</td><td class="column-4"><strong><font color="red">Hallucination</font></strong></td>
</tr>
<tr class="row-11">
	<td class="column-1"><strong>10a</strong></td><td class="column-2">Selected restaurant labeled "seafood" rather than "sushi" by OpenTable</td><td class="column-3">Both</td><td class="column-4">False alarm</td>
</tr>
<tr class="row-12">
	<td class="column-1"><strong>10b</strong></td><td class="column-2">Task not completed—participant never reached the reservation form</td><td class="column-3">ChatGPT</td><td class="column-4"><strong><font color="red">Hallucination</font></strong></td>
</tr>
</tbody>
</table>
<!-- #tablepress-1047 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1: </strong>The 11 AI-generated problems not identified by any human researcher, classified by type.</p>
<h3>Gemini’s Genuine Find</h3>
<p>Let’s start with the good news: All four Gemini runs identified that after the participant expanded the cuisine filter to show all 86 cuisines, she used Ctrl-F to search the page for “sushi” (5b-Gem)—an event not reported by any of the human evaluators. It happened quickly, so it’s possible that the search field was not in the visual focus of the humans who were likely examining the list of cuisines (Figure 2). We consider this a true usability problem because (1) this behavior was driven by poor filter design and (2) it was unsuccessful—the word “sushi” was not on the page even though the cuisine filter was fully expanded.</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/052626-F2V2.png"><img loading="lazy" decoding="async" class="alignnone size-full wp-image-47654" src="https://measuringu.com/wp-content/uploads/2026/05/052626-F2.png" alt="Frame from video showing Ctrl-F search field with first few letters of “sushi” typed at the top right of the screen and 28 of the 86 cuisine types on the left. " width="825" height="648" srcset="https://measuringu.com/wp-content/uploads/2026/05/052626-F2.png 825w, https://measuringu.com/wp-content/uploads/2026/05/052626-F2-300x236.png 300w, https://measuringu.com/wp-content/uploads/2026/05/052626-F2-768x603.png 768w, https://measuringu.com/wp-content/uploads/2026/05/052626-F2-600x471.png 600w" sizes="auto, (max-width: 825px) 100vw, 825px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 2:</strong> Frame from video showing Ctrl-F search field with first few letters of “sushi” typed at the top right of the screen and 28 of the 86 cuisine types on the left.</p>
<h3>Seven False Alarms</h3>
<p>Next, the not-so-good news. When a researcher identifies something that happened, but it’s not really considered a problem, it’s referred to as a false alarm. Sometimes things are literally a feature and not a bug! From our interpretation, AIs generated seven false alarms (not too different from what <a href="https://measuringu.com/false-positives/">you sometimes see</a> with a group of human evaluators).</p>
<h4><strong>The seafood/sushi labeling issue (6b, 6c, 7b-GPT, 8b, 10a)</strong></h4>
<p>Five of the seven false alarms (6b, 6c, 7b-GPT, 8b, 10a) were derived from ChatGPT taking the search for sushi restaurants too literally. After searching for sushi, many OpenTable results were labeled &#8220;seafood.&#8221; ChatGPT flagged this repeatedly across multiple runs in different ways (e.g., weak cuisine presentation, non-sushi results surfacing, poor decision support), but they all trace back to the same fundamental observation. ChatGPT only considered acceptable restaurants that OpenTable labeled as sushi restaurants, not restaurants that serve sushi on the menu, regardless of OpenTable&#8217;s labeling.</p>
<p>The restaurant the participant ultimately selected was labeled seafood, which led ChatGPT to declare task failure in three of four runs. The human reviewers took a more pragmatic view: the restaurant served sushi, so the participant successfully completed the task. Gemini flagged the same seafood/sushi labeling issue once (10a) but didn&#8217;t spiral into multiple variations of it.</p>
<h4><strong>Seating options not shown until after time selection (9b-Gem)</strong></h4>
<p>OpenTable withholds seating options until you pick a time. Given the range of possible seating configurations (inside, patio, bar, banquette, communal, high top, private, counter), showing them before a time is selected isn&#8217;t really feasible. And if a seating option doesn&#8217;t work out, the recovery path is low friction. Gemini flagged this as a problem. The human researchers recognized this as a design tradeoff rather than a usability problem.</p>
<h4><strong>Filters not helpful (4b-GPT)</strong></h4>
<p>We categorized this as a false alarm because it was overly vague. It’s true that there were issues with some filters (e.g., cuisine), but that was not true of all filters.</p>
<h3>Three Hallucinations</h3>
<p>In contrast to false alarms, which we consider misinterpretations of events that happened, a hallucination is when a problem is associated with something that just didn’t happen. We saw three of these.</p>
<h4><strong>AI claimed the participant incorrectly selected the highest price tier (7b-Gem)</strong></h4>
<p>From the narrative of the second Gemini run:</p>
<blockquote><p><em>The task required selecting a restaurant that was not the lowest or highest price point.</em></p>
<p><em>Problem: The participant chose Ocean Prime, which is a restaurant (the highest tier on the platform). At 05:13, the participant verbally identified this as &#8220;mid-range.&#8221;</em></p>
<p><em>User Impact: The participant technically failed this part of the task constraints.</em></p></blockquote>
<p><strong><em>This didn’t happen</em></strong>. Ocean Prime had a mid-range price designation.</p>
<h4><strong>AI claimed the participant set the reservation time for 5:10 pm (9c-Gem) </strong></h4>
<p>From the narrative of the second Gemini run:</p>
<blockquote><p><em>The participant selected Ocean Prime at 05:10</em></p></blockquote>
<p><strong><em>This didn’t happen</em></strong>. The participant, in accordance with the task instructions, selected 5:00 pm.</p>
<h4><strong>AI claimed the participant did not reach the reservation form (10b-ChatGPT)</strong></h4>
<p>From the narrative of the second ChatGPT run:</p>
<blockquote><p><em>By the end of the clip, they are still comparing list items and time slots; they do not appear to reach the restaurant detail/reservation form step.</em></p></blockquote>
<p><strong><em>This isn’t accurate</em></strong>. The clip ended with the participant selecting the reservation time, then standard dining room seating, then stopping before entering her personal information.</p>
<p>The good news is that there were only three hallucinations out of 11 AI-generated problems. The bad news is you can&#8217;t know which AI-generated problem descriptions were hallucinated without watching the video and reviewing all the problems yourself.</p>
<h2>Summary and Discussion</h2>
<p>In this article, we focused on qualitative similarities and differences in the usability problems listed by professional human UX researchers and two AIs (Gemini 3 Flash Thinking and ChatGPT-5.4 Thinking) after reviewing a video in which a participant made a restaurant reservation.</p>
<p>Our key findings were:</p>
<p><strong>False alarms and hallucinations dominate.</strong> Of the 11 problems the AIs generated that no human flagged, seven (64%) were false alarms, three (27%) were hallucinations, and one (9%) was a genuine find. That&#8217;s a useful number to keep in mind: roughly nine out of ten AI-only problems in this study required either correction or dismissal.</p>
<p><strong>AI adds value as a junior researcher, not a trusted expert.</strong> AI was able to find one problem (a participant had to use Ctrl-F) that was real and useful and not found by humans. But getting to it required reviewing ten other problems that ranged from technically true but irrelevant to simply fabricated. The ROI depends on how much that review costs you.</p>
<p><strong>Most false alarms came from a single fixation.</strong> Five of the seven traced back to ChatGPT interpreting &#8220;sushi restaurant&#8221; more literally than any human would. At least in this video and our criteria for what constitutes a problem, this is a systematic bias worth knowing about if you&#8217;re using these models for task-based evaluations.</p>
<p><strong>Hallucinations were infrequent but consequential.</strong> Three of the problems (27%) were hallucinations. Although nominally low, this is probably too high for most applications. You can&#8217;t catch those without going back to the video, which means human review isn&#8217;t optional.</p>
<p><strong>Like humans, AI usability reviews of videos are prone to the “evaluator effect.” </strong>Just like with human evaluators, multiple runs of AI usability evaluations of videos are not perfectly consistent, so it’s good practice to run these evaluations multiple times for consistency checks. Two of the three hallucinations came from the same Gemini run. Running multiple evaluations and looking for consistency across runs is a practical filter before any human review.</p>
<p><strong>Bottom line: AI usability reviews of videos require human oversight. </strong>In their current form (what we tested), these AI products can add value to this type of UX research, but more as junior researchers whose actions and conclusions require expert human oversight rather than as trusted experts themselves.<strong><br />
</strong></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>How Many Years Does It Take to Become a Senior UX Researcher?</title>
		<link>https://measuringu.com/how-many-years-does-it-take-to-become-a-senior-ux-researcher/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=how-many-years-does-it-take-to-become-a-senior-ux-researcher</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 19 May 2026 22:15:45 +0000</pubDate>
				<category><![CDATA[Methods]]></category>
		<category><![CDATA[Research]]></category>
		<category><![CDATA[Usability]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[UX Maturity]]></category>
		<category><![CDATA[experience]]></category>
		<category><![CDATA[Salary Survey]]></category>
		<category><![CDATA[UX Salary Survey]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47607</guid>

					<description><![CDATA[What does it take to become a senior UX researcher? An advanced degree? Particular experience and skills, like the number of moderated studies conducted or a variety of methods employed? While all those play a role, the type of job (in-house small-team, in-house large-team, solo researcher, or agency) can affect what you are exposed to. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1.png"><img loading="lazy" decoding="async" class="alignleft wp-image-47633 size-medium" src="https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-300x169.png" alt="Feature image showing an entry-level UX researcher becoming a senior UX researcher over several years" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1-600x338.png 600w, https://measuringu.com/wp-content/uploads/2026/05/051926-FeatureImage1.png 2000w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>What does it take to become a senior UX researcher?</p>
<p>An advanced degree? Particular experience and skills, like the number of moderated studies conducted or a variety of methods employed?</p>
<p>While all those play a role, the type of job (in-house small-team, in-house large-team, solo researcher, or agency) can affect what you are exposed to.</p>
<p>Certainly, most would agree that one to two years of experience seems too little time to demonstrate senior-level performance in UX research. We thought that something around five years of experience was a good benchmark. But is that warranted? What is a good number of years of experience?</p>
<p>There is no official rule book on titles. As is often the case when making decisions about jobs, we can use a few approaches:</p>
<ol>
<li>Principle-based: Set a rule based on a principle that disregards what people do.</li>
<li>Tradition and trends: Look to broader workforce trends, what others report, or guidance online.</li>
<li>Data: See what’s happening in practice if you have access to data.</li>
</ol>
<h2>Principle Based</h2>
<p>Even though there isn’t an official designation, we can look broadly at how long it takes to master a skill or job like UX researcher. One used in popular culture (based on some research) and popularized by Malcolm Gladwell is the <a href="https://jamesclear.com/deliberate-practice-strategy">10,000-hour rule</a>. That is, after about 10k hours of practice, you master a skill. That is a very rough guideline and <a href="https://www.bbc.com/future/article/20121114-gladwells-10000-hour-rule-myth">definitely has its critics</a>.</p>
<h2>Tradition and Trends</h2>
<p>Seniority levels can <a href="https://www.indeed.com/career-advice/career-development/seniority-level">differ by industry type</a>. For example:</p>
<ul>
<li>The <a href="https://www.asce.org/career-growth/early-career-engineers/asce-guidelines-for-engineering-grades">American Society of Civil Engineers</a> has grades (from I to VIII) with <a href="https://www.asce.org/-/media/asce-images-and-files/career-and-growth/early-career-engineer/engineering-grades.pdf">detailed descriptions</a> of expected skills and responsibilities. For example, an engineer at Grade IV has at least four years of experience.</li>
<li><a href="https://hrsimple.com/law-firm-hierarchy-roles-and-career-paths/">Associates in a law firm</a> can be junior (1–3 years of experience), mid-level (4–6 years), or senior (7–10 years).</li>
<li>In <a href="https://strategycase.com/big-4-salaries/">large consulting firms</a>, senior associates typically have 2–5 years of experience.</li>
</ul>
<p>For the expected minimum number of years of experience for UX researchers, it makes sense to start with personal experience. In our decades of experience at large companies (IBM, Oracle, GE, Intuit, PeopleSoft), something like five years was a loose criterion. Below that, people would question the designation.</p>
<p>We carry a similar tradition at MeasuringU, and those with five years’ experience are considered senior. But at a tech-enabled agency, UX researchers typically conduct hundreds of moderated sessions and use a wide variety of methods such as unmoderated benchmarking, eye-tracking, in-depth interviews, diary studies, and surveys. A couple of years working here usually exposes a researcher to significantly more UX-related tasks than in a typical in-house role. At the same time, they are much less exposed to the very real job of navigating the politics of competing stakeholders and corporate hierarchies.</p>
<h2>Data: Salary Surveys, LinkedIn Profiles, and Job Posts</h2>
<p>Our preferred method is looking for data to guide decisions. We have three sources. The first is the bi-annual UXPA Salary Survey, which was last conducted in <a href="https://uxpa.org/salary-surveys/">2024</a>. The second is LinkedIn, which provides access to job titles and a crude way of determining years of experience. The third is requirements from recent job postings.</p>
<h3>UXPA Senior User Researcher Data</h3>
<p>The 2024 Salary Survey had 444 responses. Of those, 64% (276) described themselves as user researchers. Respondents could pick one of five employment levels. Table 1 shows that about half (130) of the user researchers classified themselves as “Senior-level, non-supervisory.”</p>

<table id="tablepress-1043" class="tablepress tablepress-id-1043">
<thead>
<tr class="row-1">
	<th class="column-1">Employment Level</th><th class="column-2">Number</th><th class="column-3"> %</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">Entry</td><td class="column-2"> 18</td><td class="column-3"> 7%</td>
</tr>
<tr class="row-3">
	<td class="column-1">Mid-level, non-supervisory</td><td class="column-2"> 73</td><td class="column-3">26%</td>
</tr>
<tr class="row-4">
	<td class="column-1">Mid-level, supervisory</td><td class="column-2"> 10</td><td class="column-3"> 4%</td>
</tr>
<tr class="row-5">
	<td class="column-1">Senior-level, non-supervisory</td><td class="column-2">130</td><td class="column-3">47%</td>
</tr>
<tr class="row-6">
	<td class="column-1">Senior-level, supervisory</td><td class="column-2"> 45</td><td class="column-3">16%</td>
</tr>
<tr class="row-7">
	<td class="column-1">Total</td><td class="column-2">276</td><td class="column-3"></td>
</tr>
</tbody>
</table>
<!-- #tablepress-1043 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1</strong>: Distribution of user researchers by employment level (2024 UXPA salary survey).</p>
<p>Respondents also selected their years of experience in response to the question “How long have you worked in this field (please round to the nearest year)” using the pre-determined buckets shown in Table 2.</p>

<table id="tablepress-1044" class="tablepress tablepress-id-1044">
<thead>
<tr class="row-1">
	<th class="column-1">Years of Experience</th><th class="column-2"># Senior</th><th class="column-3">% Senior</th><th class="column-4">% With More Experience</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1"> 0–2 yrs</td><td class="column-2"> 1</td><td class="column-3"> 1%</td><td class="column-4">99%</td>
</tr>
<tr class="row-3">
	<td class="column-1"> 3–4 yrs</td><td class="column-2">17</td><td class="column-3">13%</td><td class="column-4">86%</td>
</tr>
<tr class="row-4">
	<td class="column-1"> 5–7 yrs</td><td class="column-2">28</td><td class="column-3">22%</td><td class="column-4">65%</td>
</tr>
<tr class="row-5">
	<td class="column-1"> 8–10 yrs</td><td class="column-2">21</td><td class="column-3">16%</td><td class="column-4">48%</td>
</tr>
<tr class="row-6">
	<td class="column-1">11–15 yrs</td><td class="column-2">21</td><td class="column-3">16%</td><td class="column-4">32%</td>
</tr>
<tr class="row-7">
	<td class="column-1">16–20 yrs</td><td class="column-2">15</td><td class="column-3">12%</td><td class="column-4">21%</td>
</tr>
<tr class="row-8">
	<td class="column-1">21+ yrs</td><td class="column-2">27</td><td class="column-3">21%</td><td class="column-4"> 0%</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1044 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2:</strong> Distribution of 130 non-supervisory senior-level user researchers by years of experience (2024 UXPA salary survey).</p>
<p>For example, only one person who reported being a senior user researcher had two years or fewer of experience. That means 99% had more than two years. The second row of the table shows that 17 had between three and four years of experience. Adding that to the one respondent with less experience gets 18 out of the 130 respondents. That means <strong>86% of non-supervisory senior user researchers reported 5 or more years of experience</strong>. Using the center of each age group as a rough estimate of experience, the average number of years across the sample was 12–13 years. Of course, people may inflate their years of experience on an anonymous survey.</p>
<p>We also looked at UX designers in the UXPA dataset and found a similar pattern. Of the 56 UX designers who self-identified as senior, 87% had at least five years of experience (Table 3).</p>

<table id="tablepress-1045" class="tablepress tablepress-id-1045">
<thead>
<tr class="row-1">
	<th class="column-1">Years of Experience</th><th class="column-2"># Senior</th><th class="column-3">% Senior</th><th class="column-4">% With More Experience</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1"> 0–2 yrs</td><td class="column-2"> 1</td><td class="column-3"> 2%</td><td class="column-4">98%</td>
</tr>
<tr class="row-3">
	<td class="column-1"> 3–4 yrs</td><td class="column-2"> 6</td><td class="column-3">11%</td><td class="column-4">87%</td>
</tr>
<tr class="row-4">
	<td class="column-1"> 5–7 yrs</td><td class="column-2">13</td><td class="column-3">23%</td><td class="column-4">64%</td>
</tr>
<tr class="row-5">
	<td class="column-1"> 8–10 yrs</td><td class="column-2">14</td><td class="column-3">25%</td><td class="column-4">39%</td>
</tr>
<tr class="row-6">
	<td class="column-1">11–15 yrs</td><td class="column-2"> 8</td><td class="column-3">14%</td><td class="column-4">25%</td>
</tr>
<tr class="row-7">
	<td class="column-1">16–20 yrs</td><td class="column-2"> 6</td><td class="column-3">11%</td><td class="column-4">14%</td>
</tr>
<tr class="row-8">
	<td class="column-1">21+ yrs</td><td class="column-2">8</td><td class="column-3">14%</td><td class="column-4">0%</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1045 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 3:</strong> Distribution of 56 non-supervisory senior-level UX designers by years of experience (2024 UXPA Salary Survey).</p>
<h3>LinkedIn Profiles</h3>
<p>Another approach is to look at how many years of experience senior UX researchers on LinkedIn have in their job history. While job dates can always be padded a bit, it’s a lot harder to claim unearned experience on a public professional forum. We did an informal examination searching for “Senior UX Researcher” and hand-counting the years of non-supervisory experience for the first 50 respondents.</p>
<p>Of the 50 profiles, the average years of experience was a bit over nine years (Table 4). The minimum was just shy of five years at 4.75. Of the 50 profiles, only four (8%) had less than five years of experience. In other words, using this crude estimate suggests 92% of senior user researchers have more than five years of experience.</p>

<table id="tablepress-1046" class="tablepress tablepress-id-1046">
<tbody>
<tr class="row-1">
	<td class="column-1"><strong>Mean Years of Experience</td><td class="column-2">9.1</td>
</tr>
<tr class="row-2">
	<td class="column-1"><strong>Min Years</td><td class="column-2">4.75</td>
</tr>
<tr class="row-3">
	<td class="column-1"><strong># < 5</td><td class="column-2">4</td>
</tr>
<tr class="row-4">
	<td class="column-1"><strong>% < 5</td><td class="column-2">8%</td>
</tr>
<tr class="row-5">
	<td class="column-1"><strong>Total #</td><td class="column-2">50</td>
</tr>
<tr class="row-6">
	<td class="column-1"><strong>% > 5</td><td class="column-2">92%</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1046 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 4:</strong> Analysis of 50 LinkedIn profiles of senior-level non-supervisory UX researchers.</p>
<h3>Job Posts</h3>
<p>Finally, we did another (very) informal search for senior UX researcher job postings (as of May 3, 2026) that were posted on Indeed. Of the five we found, all explicitly required five or more years of experience.</p>
<h2>Discussion and Summary</h2>
<p>There’s no official rule for what makes a UX researcher senior, but multiple approaches point to a consistent answer: at least five years.</p>
<ul>
<li><strong>Principle-based heuristics are consistent with five.</strong> Guidelines loosely based on research (like the 10,000-hour rule) suggest it takes about <strong>five years of focused experience</strong> to develop expertise. This is a weak rationale, but it&#8217;s a starting point.</li>
<li><strong>Tradition and trends suggest five.</strong> In our experience in the industry, it’s common to use <strong>five years as a minimum threshold</strong>. Other industries fall close to the five-year threshold as well.</li>
<li><strong>Salary survey data supports five.</strong> In the 2024 UXPA Salary Survey, <strong>86% of senior UX researchers reported five or more years of experience</strong>, with an average of around 12–13 years. Of the senior UX Designers, an adjacent role in the UX industry, 87% reported five or more years of experience.</li>
<li><strong>Existing profiles and open jobs show five+ years.</strong> Our LinkedIn sample of 50 senior UX researchers showed similar results, with <strong>about 90% above five years of experience</strong> and an average of 9–10 years. Finally, a selection of five currently open senior UX researcher jobs on Indeed all explicitly require at least five years of experience.</li>
</ul>
<p>If you’re looking to set a threshold for becoming senior, five years seems like a good rule.</p>
<p>Of course, years alone don’t define seniority, but if someone has fewer than five years of experience, the <em>senior</em> title should be the exception, not the rule.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>How to Interpret a Rating Scale Without Historical Data</title>
		<link>https://measuringu.com/how-to-interpret-a-rating-scale-without-historical-data/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=how-to-interpret-a-rating-scale-without-historical-data</link>
		
		<dc:creator><![CDATA[Jim Lewis, PhD and Jeff Sauro, PhD]]></dc:creator>
		<pubDate>Tue, 12 May 2026 20:19:38 +0000</pubDate>
				<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[Rating Scale]]></category>
		<category><![CDATA[UX]]></category>
		<category><![CDATA[Rating Scales]]></category>
		<guid isPermaLink="false">https://measuringu.com/?p=47556</guid>

					<description><![CDATA[UX researchers use a lot of rating scales. We recommend using standardized rating scales when possible. One of the benefits of some standardized scales, such as the SUS, SUPR-Q®, and UX-Lite®, is that you have a reference database of historical data. But there’s not always a standardized questionnaire for everything you’re hoping to measure, so [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><a href="https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-scaled.png"><img loading="lazy" decoding="async" class="alignleft wp-image-47595 size-medium" src="https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-300x169.png" alt="Feature image showing a UX researcher interpreting a rating scale without historical data" width="300" height="169" srcset="https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-300x169.png 300w, https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-1024x576.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-768x432.png 768w, https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-1536x864.png 1536w, https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-2048x1152.png 2048w, https://measuringu.com/wp-content/uploads/2026/05/051226-FeatureImage-600x338.png 600w" sizes="auto, (max-width: 300px) 100vw, 300px" /></a>UX researchers use a lot of rating scales. We recommend using standardized rating scales when possible. One of the benefits of some standardized scales, such as the SUS, SUPR-Q<sup>®</sup>, and UX-Lite<sup>®</sup>, is that you have a reference database of historical data.</p>
<p>But there’s not always a standardized questionnaire for everything you’re hoping to measure, so researchers need to create <a href="https://en.wikipedia.org/wiki/Ad_hoc">ad hoc</a> ones.</p>
<p>Data collected with ad hoc rating scales can be difficult to interpret, especially if you don’t have any historical data (e.g., from past product performance or competitors).</p>
<p>If you’re comparing multiple conditions (e.g., ratings on attributes for two or more websites), then you can check for significant differences in rating scale means.</p>
<p>But even clear differences in means don’t answer the question about whether a given mean indicates a poor or good user experience.</p>
<p>In this article, we provide a way to interpret five- and seven-point UX rating scales when you don’t have enough historical data for custom benchmarks. We use the well-known distribution of the System Usability Scale (<a href="https://measuringu.com/10-things-sus/">SUS</a>) as the basis for our recommendation.</p>
<h2>UX Rating Scales Tend to Be Negatively Skewed</h2>
<p>If you’ve never plotted your distributions of rating scale response options, you should. But don’t be surprised when you see a negatively skewed distribution (tail of data points to the left).</p>
<p>Most UX rating scales have this negative skew because (1) most item stems have a positive tone (e.g., “I felt very confident using this website”) and (2) respondents are <a href="https://dl.acm.org/doi/pdf/10.1145/175276.175282">generally more likely to agree</a> (selecting higher responses). This means that the middle value (e.g., a 3 on a five-point scale) isn’t a good measure of the “average.” This skew doesn’t make the responses necessarily bad or not useful. It just means you need to account for that skew when interpreting them.</p>
<p>For example, you can see the skew in distributions of SUS scores, for which 50 is the middle of the scale (Figure 1), but is not the middle of the distribution (68 is the median).</p>
<p><a href="https://measuringu.com/wp-content/uploads/2026/05/Figure1-1.png"><img loading="lazy" decoding="async" class="alignnone wp-image-47597 size-full" src="https://measuringu.com/wp-content/uploads/2026/05/Figure1-1.png" alt="Figure 1: Distribution of 3,187 individual SUS scores (50 is the middle of the scale, but the median is 68)." width="1200" height="719" srcset="https://measuringu.com/wp-content/uploads/2026/05/Figure1-1.png 1200w, https://measuringu.com/wp-content/uploads/2026/05/Figure1-1-300x180.png 300w, https://measuringu.com/wp-content/uploads/2026/05/Figure1-1-1024x614.png 1024w, https://measuringu.com/wp-content/uploads/2026/05/Figure1-1-768x460.png 768w, https://measuringu.com/wp-content/uploads/2026/05/Figure1-1-600x360.png 600w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></a></p>
<p class="wp-caption-text" style="text-align: left;"><strong>Figure 1:</strong> Distribution of 3,187 individual SUS scores (50 is the middle of the scale, but the median is 68).</p>
<h2>Default Benchmarks Based on Historical SUS Distribution</h2>
<p>Taking advantage of the well-known distribution of the SUS, we created a curved grading scale that is <a href="https://www.researchgate.net/publication/324116412_The_System_Usability_Scale_Past_Present_and_Future">widely used in UX research</a> (Table 1). We’ll use this as a basis for interpreting ad hoc scales.</p>

<table id="tablepress-1040" class="tablepress tablepress-id-1040">
<thead>
<tr class="row-1">
	<th class="column-1">SUS Score Range</th><th class="column-2">Grade</th><th class="column-3">Percentile Range</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">84.1–100</td><td class="column-2">A+</td><td class="column-3">96–100</td>
</tr>
<tr class="row-3">
	<td class="column-1">80.8–84.0</td><td class="column-2">A</td><td class="column-3">90–95</td>
</tr>
<tr class="row-4">
	<td class="column-1">78.9–80.7</td><td class="column-2">A−</td><td class="column-3">85–89</td>
</tr>
<tr class="row-5">
	<td class="column-1">77.2–78.8</td><td class="column-2">B+</td><td class="column-3">80–84</td>
</tr>
<tr class="row-6">
	<td class="column-1">74.1–77.1</td><td class="column-2">B</td><td class="column-3">70–79</td>
</tr>
<tr class="row-7">
	<td class="column-1">72.6–74.0</td><td class="column-2">B−</td><td class="column-3">65–69</td>
</tr>
<tr class="row-8">
	<td class="column-1">71.1–72.5</td><td class="column-2">C+</td><td class="column-3">60–64</td>
</tr>
<tr class="row-9">
	<td class="column-1">65.0-71.0</td><td class="column-2">C</td><td class="column-3">41–59</td>
</tr>
<tr class="row-10">
	<td class="column-1">62.7–64.9</td><td class="column-2">C−</td><td class="column-3">35–40</td>
</tr>
<tr class="row-11">
	<td class="column-1">51.7–62.6</td><td class="column-2">D</td><td class="column-3">15–34</td>
</tr>
<tr class="row-12">
	<td class="column-1"> 0.0–51.6</td><td class="column-2">F</td><td class="column-3"> 0–14</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1040 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 1:</strong> Curved grading scale for the SUS.</p>
<p>The 50<sup>th</sup> percentile of this scale is a SUS score of 68, a solid C. Another important benchmark commonly used in practice is an aspirational score of 80 (the upper end of an A−, a bit higher than the 85<sup>th</sup> percentile). Scores lower than 51.7 are in the F range (just below the 15<sup>th</sup> percentile).</p>
<p>Based on the SUS research, when we consult with clients who need a benchmark for five- or seven-point scales and there is no historical data, we usually recommend setting a benchmark for average to about 70% of the range of the scale, 80% for good, and 50% for poor—similar to the historical benchmarks for the SUS. For example, this is what we did when we created our <a href="https://measuringu.com/grading-scales-for-the-ux-lite/">standard grading scale for the UX-Lite</a>.</p>
<p>Table 2 shows those values for five- and seven-point scales (the midpoint for a five-point scale is 3, and for a seven-point scale is 4).</p>

<table id="tablepress-1042" class="tablepress tablepress-id-1042">
<thead>
<tr class="row-1">
	<th class="column-1">Location on Scale</th><th class="column-2">Interpretation</th><th class="column-3">Five-point</th><th class="column-4">Seven-point</th>
</tr>
</thead>
<tbody>
<tr class="row-2">
	<td class="column-1">80%</td><td class="column-2">Good</td><td class="column-3">4.2</td><td class="column-4">5.8</td>
</tr>
<tr class="row-3">
	<td class="column-1">70%</td><td class="column-2">Average</td><td class="column-3">3.8</td><td class="column-4">5.2</td>
</tr>
<tr class="row-4">
	<td class="column-1">60%</td><td class="column-2">Below Average</td><td class="column-3">3.4</td><td class="column-4">4.6</td>
</tr>
<tr class="row-5">
	<td class="column-1">50%</td><td class="column-2">Poor</td><td class="column-3">3.0</td><td class="column-4">4.0</td>
</tr>
</tbody>
</table>
<!-- #tablepress-1042 from cache -->
<p class="wp-caption-text" style="text-align: left;"><strong>Table 2:</strong> Initial benchmarks for 70 and 80% of the range of five- and seven-point scales.</p>
<p>The formula for computing these values is based on the <a href="https://measuringu.com/converting-scales-to-100-points/">methods for interpolating rating scale scores</a> that start with 1 to a 0–100-point scale, algebraically manipulated to compute the benchmark for the rating scale from the target range (e.g., 80% of the scale, expressed as 80 in the computation) and the maximum possible value of the rating scale (e.g., typically 5 or 7 for scales that start with 1):</p>
<p>Benchmark = Target / (100 / (MaxRating − 1)) + 1</p>
<p>For example, to find 70% of the range of a five-point scale, the benchmark would be:</p>
<p>70 / (100 /(5 − 1)) + 1 = 70 / 25 + 1 = 3.8</p>
<p>An <a href="https://measuringu.com/types-of-100-point-scales/">alternative approach</a> is to convert five- or seven-point ratings to a 0–100-point scale. John Brooke, the developer of the SUS, <a href="https://uxpajournal.org/sus-a-retrospective/">described the value of this approach</a>: “Project managers, product managers, and engineers were more likely to understand a scale that went from 0 to 100 than one that went from 10 to 50, and the important thing was to be able to grab their attention in the short space of time they were likely to spend thinking about usability, without having to go into a detailed explanation.”</p>
<p>The general formula for converting a five- or seven-point scale to 0–100 points is:</p>
<p>Rating100 = (Rating − 1) * 100 / (MaxRating − 1)</p>
<p>For example, a five-point mean rating of 4.2 would become 80:</p>
<p>(4.2 − 1) * (100 / (5 − 1)) = 3.2(25) = 80</p>
<p>A seven-point mean rating of 4.0 would become 50:</p>
<p>(4 − 1) * (100 / (7 − 1)) = 3(16.67) ≈ 50</p>
<p><strong><em>Caveat</em></strong><em><strong>:</strong> Note that these are initial benchmarks to use when UX researchers lack a more grounded rationale for interpreting mean rating scale scores. After a reasonable amount of data collection with the scale, it’s a good idea to revisit the initial benchmarks to see whether they should be adjusted.</em></p>
<h2>Summary</h2>
<p>When you&#8217;re working with an ad hoc rating scale and have no historical data to lean on, the SUS distribution gives you a principled starting point. Because UX rating scales share a consistent negative skew (driven by positive item wording and respondent agreement bias), benchmarks derived from the SUS translate reasonably well to other five- and seven-point scales. It’s not that there’s something magic about the SUS. It works well because it’s a composite of ten five-point UX rating scales that share the tendency of other UX rating scales to be negatively skewed (more favorable than unfavorable). This means that benchmarks informed by the SUS provide a good initial approximation for other UX rating scales.</p>
<p>The characteristics of UX rating scales that this pattern supports are:</p>
<ul>
<li>Setting “Poor” below the midpoint of the scale (50% of the range) because means of positive-tone UX rating scales are consistently higher than the scale midpoint.</li>
<li>Setting “Good” above 80% of the scale range is the <a href="https://www.uslanguageservices.com/guides-resources/understanding-the-u-s-grading-system/">traditional score for a B</a> (above average).</li>
</ul>
<p>Placing other cut points between 50% and 80% leads to these initial benchmarks:</p>
<ul>
<li><strong>Good</strong>: Located at <strong>80%</strong> of the range of the scale</li>
<li><strong>Average</strong>: Located at <strong>70%</strong> of the range of the scale</li>
<li><strong>Below average</strong>: Located at <strong>60%</strong> of the range of the scale</li>
<li><strong>Poor</strong>: Located at <strong>50%</strong> of the range of the scale (the midpoint)</li>
</ul>
<p>It’s important to keep in mind that these are reasonable best guesses without a strong normative database. For UX rating scale items that will be used frequently over time, researchers should plan to build normative databases and use them to tune the benchmarks (like we have <a href="https://measuringu.com/evolution-of-seq/">done with the SEQ<sup>®</sup></a>).</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!-- plugin=object-cache-pro client=phpredis metric#hits=7695 metric#misses=53 metric#hit-ratio=99.3 metric#bytes=4685623 metric#prefetches=223 metric#store-reads=112 metric#store-writes=168 metric#store-hits=318 metric#store-misses=35 metric#sql-queries=52 metric#ms-total=631.35 metric#ms-cache=33.79 metric#ms-cache-avg=0.1211 metric#ms-cache-ratio=5.4 sample#redis-hits=9498118 sample#redis-misses=1798303 sample#redis-hit-ratio=84.1 sample#redis-ops-per-sec=69 sample#redis-evicted-keys=0 sample#redis-used-memory=70238744 sample#redis-used-memory-rss=58585088 sample#redis-memory-fragmentation-ratio=0.8 sample#redis-connected-clients=1 sample#redis-tracking-clients=0 sample#redis-rejected-connections=50 sample#redis-keys=101949 -->
