<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Fri, 02 Oct 2026 20:49:18 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>This Week in AI: Agents Are Outrunning the Systems Around Them</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/#respond</comments>
				<pubDate>Fri, 02 Oct 2026 15:54:24 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19874</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Cyberattacks, power constraints, faster model releases, and consumer agents put pressure on the systems built to support them]]></custom:subtitle>
		
				<description><![CDATA[On the latest episode of This Week in AI, host Vicki Reyzelman, a senior solutions engineer at Akamai, traced a common problem across cybersecurity, energy, model releases, consumer hardware, and regulation. AI agents can now probe networks, coordinate with other agents, make purchases, and interact with real-world systems faster than many organizations can respond. We’re [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe width="800" height="450" src="https://www.youtube.com/embed/BmBTbqm1pfQ?si=L0ZAcjvz2ZLxoqRh" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">On the latest episode of <em>This Week in AI</em>, host Vicki Reyzelman, a senior solutions engineer at Akamai, traced a common problem across cybersecurity, energy, model releases, consumer hardware, and regulation. AI agents can now probe networks, coordinate with other agents, make purchases, and interact with real-world systems faster than many organizations can respond. We’re seeing those capabilities move into systems built for slower, more predictable software.</p>



<h2 class="wp-block-heading"><strong>Security has to operate at agent speed</strong></h2>



<p class="wp-block-paragraph">Vicki opened with an incident in which <a href="https://www.businessday.co.za/world/2026-09-24-albanese-calls-openai-agents-medicare-breach-unacceptable/" target="_blank" rel="noopener">an OpenAI agent reportedly found ways around security controls</a> while researching public information in Australia’s Medicare system. The activity didn’t expose any&nbsp; personal Medicare records, but <a href="https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078" target="_blank" rel="noopener">OpenAI reportedly took 54 days to identify the incident</a> and another month to notify the government. A response cycle measured in weeks can’t keep pace with systems that can test defenses in seconds.</p>



<p class="wp-block-paragraph">She also brought up <a href="https://news.un.org/en/story/2026/09/1168380" target="_blank" rel="noopener">the recent Hugging Face incident</a> involving a swarm of 1,200 agents that exchanged roughly 70,000 messages while coordinating their work. Agents can change tactics faster than traditional security processes play out, so teams can no longer rely on the familiar methods of addressing suspicious behavior. Companies are now experimenting with runtime enforcement, agent sandboxes, enterprise browsers, and other controls that sit closer to execution.</p>



<p class="wp-block-paragraph">Policymakers are searching for workable controls too, from <a href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924" target="_blank" rel="noopener">California proposals for emergency AI shutdown mechanisms</a> to international discussions about <a href="https://www.securitycouncilreport.org/whatsinblue/2026/09/artificial-intelligence-high-level-briefing-2.php" target="_blank" rel="noopener">independent model evaluation</a>. Teams can’t govern agent behavior they can’t see, so they need to know what an agent did and when its behavior crossed a boundary.</p>



<h2 class="wp-block-heading"><strong>Power and latency are becoming model decisions</strong></h2>



<p class="wp-block-paragraph">Power is one constraint software teams can’t code their way around. Vicki pointed to a <a href="https://time.news/doe-invests-nearly-2-billion-to-boost-grid-capacity-for-ai-data-centers" target="_blank" rel="noopener">$2 billion US Department of Energy investment</a> across 26 states alongside hundreds of billions of dollars in planned AI spending from Microsoft, Amazon, Alphabet, and Meta. Data centers can add servers quickly, but it won’t make a difference if the grid can’t provide the energy those servers require.</p>



<p class="wp-block-paragraph">Meanwhile, <a href="https://www.searchintel.tech/research/ai-model-release-pace/" target="_blank" rel="noopener">major model releases are arriving roughly every 17 days</a>, with context windows now exceeding one million tokens. Open weight and edge models are advancing too, particularly around low-latency reasoning. More frequent releases and heavier inference workloads put added pressure on networks, compute, and budgets.</p>



<p class="wp-block-paragraph">Solving this challenge may mean companies have to run more reasoning at the edge or locally, where systems can reduce latency and avoid sending every request across the network. That gives teams another architectural choice to make alongside model selection. A frontier model may be appropriate for one workload, while a smaller local model may be faster and cheaper for another.</p>



<h2 class="wp-block-heading"><strong>Consumer agents move autonomy into everyday life</strong></h2>



<p class="wp-block-paragraph">Consumer hardware puts those architecture and governance choices directly in users’ hands. AI-enabled glasses, pendants, and other devices stay with users throughout the day and can learn preferences, connect with outside services, and take actions such as shopping or making reservations. Meta’s new Muse agent is one example of that shift.</p>



<p class="wp-block-paragraph">Meta says the Muse ecosystem <a href="https://explainx.ai/blog/meta-connect-2026-everything-announced-muse-glasses-vr-2026" target="_blank" rel="noopener">already includes roughly 1,500 developer connectors</a>, including integrations with retailers such as Walmart and Best Buy. If more purchases begin with an agent acting for the customer, companies may have to rethink how people discover products and complete transactions. The convenience of Amazon Prime and one-click shopping, for example, looks different when another system is comparing options and buying on a user’s behalf.</p>



<p class="wp-block-paragraph">Muse already ran into problems, including exposing information it wasn’t supposed to and relying on humans to complete some tasks, such as making dinner reservations. Those failures carry more weight when the software can spend money or act on personal preferences. Users and businesses need clear limits on what an agent can access, what it can do without approval, and how those actions are recorded.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">Deploying an agent means taking responsibility for the systems around it. Security controls, power and network constraints, local versus remote inference, and permission boundaries all shape what these systems can safely do in production. For practitioners, the job now includes the architecture around the models.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news and developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. You can also hear more from Vicki on <a href="https://substack.com/@vickireyzelman" target="_blank" rel="noopener">her Substack</a>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Coding Agents Love Decision Records</title>
		<link>https://www.oreilly.com/radar/coding-agents-love-decision-records/</link>
				<comments>https://www.oreilly.com/radar/coding-agents-love-decision-records/#respond</comments>
				<pubDate>Fri, 02 Oct 2026 11:24:21 +0000</pubDate>
					<dc:creator><![CDATA[Duncan Davidson]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19867</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Coding-agents-love-decision-records.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Coding-agents-love-decision-records-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Duncan Davidson’s blog and is being republished here with the author’s permission. Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript. Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://duncandavidson.com/agents-love-decisions" target="_blank" rel="noopener">Duncan Davidson’s blog</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript.</p>



<p class="wp-block-paragraph">Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture significant design choices, their context, and the reasons behind them. Like many tools built for human software teams, ADRs work well for coding agents too.</p>



<p class="wp-block-paragraph">Agents often arrive with little memory of yesterday and only a narrow view of a codebase. Even systems with persistent memory may preserve context without establishing whether it is accurate, current, or accepted by the human team. Decision records help them understand the intent behind the code rather than having to infer it. Keeping them in a project repository spares agents from having to trawl through issues, search chats, and perform code archaeology. When you record intent explicitly, an agent is less likely to mistake an implementation detail for a foundational rule.</p>



<p class="wp-block-paragraph">Once a decision enters an agent’s context window, the agent may adhere to it even more rigidly than a human would. In my own work, I’ve seen agents fight tooth and nail to apply an accepted decision even when it is obsolete. In one case, an agent preserved an outdated storage abstraction across a new feature because an ADR still described it as mandatory. Instead of flagging the mismatch, it added another layer to keep the new requirement technically compatible with the old ruling.</p>



<p class="wp-block-paragraph">The first remedy is to give agents explicit permission to question decisions that no longer fit—and to watch for signs that they’re overfitting. But that solves only half the problem. When you invite an agent to update a decision, a second tendency appears: preserving the deliberation. Every clarification becomes an amendment explaining its own existence at the expense of clarity. Small implementation details become rules, and cross-references acquire their own restatements and justifications. The result is overlitigated prose that is hard for humans to read.</p>



<p class="wp-block-paragraph">ADRs should absolutely be readable by humans, especially as we lean on agents to generate more and more code. To counter this, I’ve become explicit in my projects’ <code>AGENTS.md</code> files about how agents should apply and maintain ADRs. Here’s an excerpt from one:</p>



<p class="wp-block-paragraph"><code>Architectural Decision Records (ADRs) are stored as Markdown files in the </code><br><code>docs/decisions directory. Treat accepted ADRs as binding. Proposed ADRs </code><br><code>are non-binding context. Superseded ADRs are historical context and do </code><br><code>not govern current work. If a given task conflicts with an accepted ADR, </code><br><code>stop and discuss whether the task or ADR should change and propose the </code><br><code>change that you think should be made. Propose new ADRs or updates to </code><br><code>existing ones when a change introduces or revises a durable product or </code><br><code>architectural decision.</code> </p>



<p class="wp-block-paragraph"><code>Keep ADRs succinct. Each ADR carries only its current text; Git history </code><br><code>is its changelog, so do not add or maintain amendment logs in ADR </code><br><code>headers. When substantively changing an accepted ADR, add or update a </code><br><code>single Updated: date line after Date:—its presence signals that history </code><br><code>exists and Git has the details. A superseded ADR records a Superseded-On: </code><br><code>date instead of Updated: , matching the Supersedes: line on the ADR that </code><br><code>replaced it. State each rule once in the ADR that owns it and cross-reference </code><br><code>it from other ADRs instead of restating it.</code></p>



<p class="wp-block-paragraph">These instructions are still evolving in my projects, and different projects will need different conventions. Some teams will prefer immutable ADRs that are superseded rather than revised; in my projects, I’m happy to have Git carry that history.</p>



<p class="wp-block-paragraph">If you do something similar, adapt the guidance to your own needs. The essential principle is that each governing ADR should describe the decision currently in force, with enough rationale to apply it. An agent doesn’t need the transcript of every argument. It needs the ruling that governs today and clear permission to stop when the ruling no longer fits.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/coding-agents-love-decision-records/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Future of Software May Be Conversational Rather Than Autonomous</title>
		<link>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/</link>
				<comments>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/#respond</comments>
				<pubDate>Thu, 01 Oct 2026 10:54:32 +0000</pubDate>
					<dc:creator><![CDATA[Robert Englander]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19862</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-future-of-software-may-be-conversational-rather-than-autonomous.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-future-of-software-may-be-conversational-rather-than-autonomous-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Robert Englander’s blog site and is being reposted here with the author’s permission. The software industry has become deeply focused on autonomous AI systems. Agents that can replace workers. Agents that can write software. Agents that can operate applications on our behalf. Entire startups are now built around the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://robenglander.com/writing/future-of-software-conversational/" target="_blank" rel="noopener">Robert Englander’s blog site</a></em> <em>and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">The software industry has become deeply focused on autonomous AI systems. Agents that can replace workers. Agents that can write software. Agents that can operate applications on our behalf. Entire startups are now built around the assumption that the natural end state of AI is autonomy.</p>



<p class="wp-block-paragraph">Some of this work is genuinely useful. AI-assisted coding can improve productivity. Generative systems are already helping people draft documents, summarize information, and accelerate certain kinds of repetitive work. There are clearly domains where more automation makes sense.</p>



<p class="wp-block-paragraph">Still, I increasingly suspect the industry may be underestimating another opportunity that feels both more practical and potentially more transformative over the long term: natural language interfaces sitting on top of deterministic software systems.</p>



<p class="wp-block-paragraph">Much of the current AI narrative assumes the model itself should become the authoritative actor. The AI writes the code. The AI performs the workflow. The AI executes the task. The AI makes the decision. The conversation around “agents” often assumes the system itself should gradually absorb more and more of that responsibility until people become optional.</p>



<p class="wp-block-paragraph">The problem is that large language models are probabilistic systems. They’re incredibly capable. But they’re still statistical machines. They hallucinate, improvise, approximate. In many contexts, that’s perfectly acceptable.</p>



<p class="wp-block-paragraph">Brainstorming, summarization, drafting, translation, and exploratory work all tolerate a degree of uncertainty. Deterministic systems generally don’t.</p>



<p class="wp-block-paragraph">Financial systems have to calculate correctly. Scheduling systems have to preserve consistency. Medical systems have to maintain integrity. Accounting systems have to reconcile accurately. Reliability is still the foundation upon which useful software is built.</p>



<p class="wp-block-paragraph">That’s one reason I think the most important role for LLMs may not be replacing deterministic systems, but reducing the friction between people and those systems.</p>



<p class="wp-block-paragraph">Historically, software interfaces forced people to adapt to machine discipline. We learned command syntax. We navigated menus and workflows. We memorized procedures. We filled out forms in exactly the way the application expected. Even graphical interfaces, which were a huge leap forward, still largely required users to think in terms of the structure of the software itself. Natural language interfaces potentially invert that relationship.</p>



<p class="wp-block-paragraph">Instead of forcing users closer to the system, the system moves closer to human expression. That may sound subtle, but I think it represents a significant shift in how software can be experienced. A user no longer needs to think primarily in terms of application structure or workflow design. The interaction begins to center more naturally around intent.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“Show me how delaying Social Security by two years impacts long-term spending.”</em></p>



<p class="wp-block-paragraph"><em>“Transfer $500 from checking into savings next Friday.”</em></p>



<p class="wp-block-paragraph"><em>“Why did my tax liability increase this year?”</em></p>



<p class="wp-block-paragraph"><em>“Find the contracts signed after January that contain auto-renewal</em></p>



<p class="wp-block-paragraph"><em>language.”</em></p>
</blockquote>



<p class="wp-block-paragraph">None of these requests eliminates the need for deterministic systems underneath. In fact, they depend on them. The natural language layer simply acts as an interpreter between human expression and authoritative execution.</p>



<p class="wp-block-paragraph">That architecture feels considerably more durable to me than the idea that probabilistic systems should become the primary authority layer themselves.</p>



<p class="wp-block-paragraph">One interesting thing about the current AI wave is that language models are often strongest in areas involving interpretation. They’re remarkably good at extracting meaning from ambiguous human communication, maintaining conversational context, translating between representations, and helping users express intent more naturally. Those are fundamentally interaction problems.</p>



<p class="wp-block-paragraph">Meanwhile, the areas where language models remain weakest are usually the areas requiring guarantees, consistency, accountability, and deterministic correctness. Those are system-of-record problems. The current industry conversation often blurs the distinction between the two.</p>



<p class="wp-block-paragraph">I don’t think conversational interfaces reduce the importance of deterministic software. If anything, they increase it. Once users begin interacting through natural language, the validation layer underneath becomes even more critical. Systems have to safely interpret intent, validate operations, preserve constraints, and maintain correctness even when the incoming requests are conversational and ambiguous.</p>



<p class="wp-block-paragraph">The conversational layer improves accessibility. The deterministic layer preserves trust.</p>



<p class="wp-block-paragraph">Both matter.</p>



<p class="wp-block-paragraph">Every major era of computing has involved some kind of interface transition. Mainframes required specialized operators. Personal computers brought graphical interfaces that made computing accessible to nonspecialists. The web normalized hyperlinks, search, and forms. Mobile computing shifted interaction toward touch and gestures. Natural language may become the next major abstraction layer.</p>



<p class="wp-block-paragraph">Not because computers suddenly became human-like, but because we finally built systems capable of translating between human communication and machine discipline at scale.</p>



<p class="wp-block-paragraph">I also think this changes how we should think about software’s future. The current AI environment sometimes frames autonomy as the inevitable destination. If an AI can partially perform a task today, many assume the long-term outcome is full replacement of the person performing that task.</p>



<p class="wp-block-paragraph">I’m not convinced that’s where the most durable value lies.</p>



<p class="wp-block-paragraph">In many domains, the real friction isn’t execution. It’s interface complexity. People struggle less with the underlying capabilities of software than with the difficulty of expressing what they actually want the software to do.</p>



<p class="wp-block-paragraph">Enterprise systems are notoriously difficult to navigate. Financial systems expose overwhelming complexity. Creative tools bury users under layers of workflow and terminology. Even relatively simple applications often require substantial onboarding before users become comfortable with them.</p>



<p class="wp-block-paragraph">Natural language interfaces potentially change that equation in a meaningful way. They allow software to meet users closer to where they already are: ordinary human communication.</p>



<p class="wp-block-paragraph">That doesn’t mean conversational systems should become undisciplined systems. In fact, I think the opposite is true. As interfaces become more conversational, the underlying architecture has to become even more rigorous about validation and execution semantics. The ambiguity doesn’t disappear. It moves.</p>



<p class="wp-block-paragraph">Historically, much of the burden of precision sat on the user. The user had to learn the syntax, understand the workflow, and conform to the application’s structure.</p>



<p class="wp-block-paragraph">Conversational systems shift more of that burden into the interpretation and validation layers of the software itself. That’s not a trivial engineering problem. It requires clarification, normalization, policy enforcement, validation, and authoritative execution underneath the conversational layer. It also requires accepting that probabilistic interpretation and deterministic execution aren’t competing ideas. They’re complementary ones.</p>



<p class="wp-block-paragraph">This is one reason I increasingly think the future of software may become conversational without necessarily becoming autonomous. The two ideas are related. But they’re not the same thing.</p>



<p class="wp-block-paragraph">There’s enormous value in reducing the natural friction between human expression and machine discipline. Large language models may ultimately prove most transformative not when they replace deterministic systems, but when they help people interact with those systems more naturally.</p>



<p class="wp-block-paragraph">For decades, people have adapted to computers. It now seems possible that software may finally start adapting to people instead.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Agentic Data Science Playbook</title>
		<link>https://www.oreilly.com/radar/the-agentic-data-science-playbook/</link>
				<comments>https://www.oreilly.com/radar/the-agentic-data-science-playbook/#respond</comments>
				<pubDate>Wed, 30 Sep 2026 16:06:24 +0000</pubDate>
					<dc:creator><![CDATA[Hugo Bowne-Anderson, Luca Fiaschi and Thomas Wiecki]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Software Architecture]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19847</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-agentic-data-science-playbook-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-agentic-data-science-playbook-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How to build and direct data science agents, and make each investigation improve the next.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Vanishing Gradients and is being republished here with the authors’ permission When an AI agent can explore a dataset, choose a modeling approach, run the analysis, and explain its findings, what should the data scientist do? Traditionally, data scientists chose each step and implemented much of the analysis themselves. [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <a href="https://hugobowne.substack.com/p/the-agentic-data-science-playbook">Vanishing Gradients</a> <em>and is being republished here with the authors’ permission</em></p>
</blockquote>



<p class="wp-block-paragraph">When an AI agent can explore a dataset, choose a modeling approach, run the analysis, and explain its findings, what should the data scientist do?</p>



<p class="wp-block-paragraph">Traditionally, data scientists chose each step and implemented much of the analysis themselves. Agentic data science changes that division of work: we can delegate an investigation, including methodological choices, while shaping the question, supplying relevant expertise, and challenging the evidence it produces. For AI-native data scientists, choosing the runtime, writing reusable skills, and designing the workflows and feedback that guide the agent are part of the analytical work.</p>



<figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="1600" height="889" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1600x889.png" alt="Verification process" class="wp-image-19865" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1600x889.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-300x167.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-767x426.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1536x853.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-2048x1138.png 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">This article provides a playbook for working with data science agents, from setting up an investigation to reviewing its results and carrying lessons into the next assignment. To see why that requires more than a capable model and a business question, consider an experiment we deliberately started with too little guidance. We gave Claude Opus 5.0 a modified version of the <a href="https://www.kaggle.com/datasets/ellipticco/elliptic-data-set" target="_blank" rel="noopener">public Elliptic dataset</a> and asked, “Build me a model to detect fraudulent nodes.” The dataset is a graph of Bitcoin transactions: each node is a transaction, and an edge represents a flow of bitcoin between transactions. Some transaction nodes carry licit or illicit labels based on the entities that created them; the rest are unlabeled. Each node has a time step indicating when its transaction was broadcast, allowing us to train on earlier time steps and test on later ones. We also wanted separate results for transactions with many connections (high-degree nodes), which mattered most in the intended application. Importantly, we renamed the columns, changed several features, and reindexed the time steps while preserving their order, making the public dataset harder for Claude to recognize.</p>



<p class="wp-block-paragraph">Claude wrote the code, trained a random forest, and reported an F1 of 0.87 and ROC AUC of 0.99. It had split transactions randomly, mixing earlier and later time steps in both the training and test sets. That test did not measure how the model would perform on transactions from later time steps. Moreover, Claude also used a feature we had planted as a proxy for the fraud label (yes, we tricked it!), giving the model leaked information it would not have when scoring a new transaction. So how do we avoid these situations?</p>



<p class="wp-block-paragraph"><strong>We then supplied the guidance missing from the initial prompt:</strong>&nbsp; We required a temporal holdout, removed the leaking feature, supplied context about how the model would be used in production, and asked for separate reporting on the high-degree nodes that mattered most. Under the corrected evaluation, F1 was 0.70, overall recall was 0.61, and recall on high-degree nodes was 0.21.</p>



<p class="wp-block-paragraph">The key point is that “Build a fraud detector” left Claude to infer how the model would be used and what would count as success. <strong>AI-native data scientists build and direct</strong> an analytical process in which agents can investigate, receive feedback, and return evidence for review. The work begins with deciding how much of the investigation to delegate.&nbsp; Agentic data science is doing data science work with AI agents as teammates. Crucially, the scope of their responsibility can extend well beyond code implementation. An agent can help frame a question, explore data, test a claim, or communicate a result, provided it has the context and tools to do the work, a way to assess its progress and validate its results.</p>



<p class="wp-block-paragraph">Asking an agent to write a pandas transformation leaves you as the bottleneck, responsible for deciding every next operation. Asking it to investigate a change in customer behavior gives it larger analytical responsibility. It can inspect a result, form another question, choose a method, and continue. The interaction becomes a conversation about the investigation rather than a sequence of requests for code.</p>



<p class="wp-block-paragraph">The question may be <em>descriptive</em> (what happened?), <em>diagnostic</em> (why did it happen?), <em>predictive</em> (what might happen next?), or <em>prescriptive</em> (what should we do?). The fraud model is predictive; the pricing investigation later in this article is diagnostic and causal. Across these kinds of work, we need to specify the question and intended use, then verify that the evidence supports the answer.</p>



<p class="wp-block-paragraph"><em>The following five practices are key to agentic data science:</em></p>



<ul class="wp-block-list">
<li>Frame the investigation.</li>



<li>Equip the agent for the assignment.</li>



<li>Organize the work through bounded experiments, competing analyses, or both, according to the question.</li>



<li>Review the result independently.</li>



<li>Preserve evidence and turn reviewed lessons into reusable expertise.</li>
</ul>



<p class="wp-block-paragraph">The first two practices set up the work. The third determines how the investigation proceeds; the fourth tests its claims. Evidence is captured throughout, and the fifth practice carries reviewed lessons into future assignments.</p>



<p class="wp-block-paragraph">As in agentic software engineering, the agentic data scientist’s two central responsibilities are <strong><em>specification</em></strong> and <strong><em>verification</em></strong>. Specify the question, intended use, and evidence the agent should produce; then verify that its analysis supports the conclusion. Agents can help with both, while the data scientist remains responsible for judging the question and the evidence.</p>



<h2 class="wp-block-heading"><strong>1. Frame the investigation with the agent</strong></h2>



<p class="wp-block-paragraph"><strong><em>Start by discussing the assignment with the agent.</em></strong> Supply the intended use and organizational context, then let it inspect the data and propose an approach. Method selection can be part of its responsibility. Your intervention matters when a proposal changes the question, rests on a questionable assumption, or needs information the agent cannot obtain. Predicting fraud and deciding which flagged entities to investigate, for example, require different evidence about errors and their consequences. A brainstorming skill such as those in <a href="https://github.com/obra/superpowers" target="_blank" rel="noopener">Superpowers</a> can help structure that conversation before you turn it into a task prompt.</p>



<p class="wp-block-paragraph">A useful specification records that shared understanding. It states the decision, relevant constraints, and evidence the investigation should produce. It need not prescribe every step. In the fraud example, “classify transactions from later time steps using only information available when each is scored” matters more than “use a random forest.” The former defines the analytical task, while the latter selects one possible implementation.</p>



<p class="wp-block-paragraph">You can specify what the investigation must establish without specifying the answer you want. “Determine whether the data support a recommendation” leaves room for an inconclusive result. “Keep trying until you find an effect” does not.</p>



<p class="wp-block-paragraph">Turn that discussion into a short analytical brief to give the agent as its task prompt. For an assignment like our fraud example, a starting version could read:</p>



<pre class="wp-block-code"><code>TASK PROMPT:
Question: Can we identify fraudulent nodes as they enter the network?
Use: Support investigation, with separate reporting on high-degree nodes.
Available information: Only inputs known when the node is scored.
Agent discretion: Explore data, propose eligible features, choose models.
Return to me: Unclear feature provenance, changes to the target or
population, or a trade-off that requires an operational decision.
Deliverable: Reproducible analysis, temporal evaluation, subgroup errors,
and a recommendation that states what the evidence cannot establish.</code></pre>



<p class="wp-block-paragraph">Review it with the agent before the investigation proceeds. If exploration reveals that the evidence cannot answer the question, revise the brief explicitly; do not quietly substitute an easier question.</p>



<p class="wp-block-paragraph">The deliverable may still be a notebook, model, or report prepared outside a production service. You can begin in the workspace where you already do that work.</p>



<h2 class="wp-block-heading"><strong>2. Equip the agent for the assignment</strong></h2>



<p class="wp-block-paragraph">The task prompt tells the agent what to investigate. It also needs to know how the project works, reach the data, run the analysis, and check the result. The harness is the system around the language model that allows this: its tools, runtime, context, permissions, and feedback from its actions. Its runtime is the environment that executes those actions. A language model alone cannot inspect a warehouse, run a simulation, or recover an interrupted statistical model fit. The environment must make those operations possible and return useful evidence about what happened.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb51f97&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb51f97" class="aligncenter size-large wp-lightbox-container"><img decoding="async" width="1600" height="820" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1600x820.png" alt="Data science agent flow" class="wp-image-19849" style="aspect-ratio:1.95" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1600x820.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-300x154.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-767x393.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1536x788.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34.png 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">Runtime choices are analytical choices as well as engineering choices. Can the agent execute Python or R with the libraries the task needs? Can a long-running fit continue after an interactive session ends? Which scientific libraries should the agent use? Can the agent inspect plots, or does it only see the code that produced them? Can you reproduce the environment in which it reported a result?</p>



<p class="wp-block-paragraph">An existing agent runtime may provide most of this. Configuring it means deciding what belongs in Markdown, what needs a tool, and what should be checked by a small script. In an investigation like the fraud example, Markdown can hold the brief and data definitions, while a Python script could check that the appropriate temporal validation split is executed. A CSV data extract may be enough for exploration; if the agent needs data warehouse access, a tool exposed through an MCP server can provide it with appropriately scoped, read-only credentials. A sentence in a prompt cannot enforce that access limit.</p>



<p class="wp-block-paragraph">Take the same care with outputs. Ask the agent to preserve the data reference, code, environment, assumptions, and diagnostics behind its report. A chat transcript is a poor substitute for a runnable analysis. Review becomes much harder when the only surviving artifact is a confident paragraph about what the agent says it did.</p>



<p class="wp-block-paragraph">Execution is only part of the problem. An agent may know how to fit a model and still misunderstand what the columns mean. It may find five revenue tables and choose the wrong one. A schema rarely explains which customers were eligible for an offer, when a measurement changed, or why the team stopped using an apparently reasonable metric. This is where <strong>agent skills</strong> and <strong>domain knowledge</strong> enter. A skill packages instructions and resources for a type of analytical work. It might contain a modeling approach, example code, required diagnostics, and guidance on when to ask for help. Data documentation supplies the organizational meaning: canonical definitions, table grain, known limitations, and the history needed to interpret a result.</p>



<p class="wp-block-paragraph">A useful skill is specific enough to change the agent’s behavior. “Be rigorous” gives it little to work with. A fraud-modeling skill can require the agent to establish feature availability, evaluate on later observations, and report performance on operationally important subgroups. For example:</p>



<pre class="wp-block-code"><code>For fraud prediction:

Establish what information is available when a node is scored.

Exclude features derived from subsequent investigations or labels.

Fit preprocessing on training data only.

Evaluate on later-arriving nodes and report the required degree groups.

Flag uncertainty about feature provenance before claiming performance.</code></pre>



<p class="wp-block-paragraph">These instructions leave room to choose a model. They encode reasons that some apparently successful models should be rejected. Where a requirement can be checked reliably in code, the skill can call a script that performs the check and records its result.</p>



<p class="wp-block-paragraph">Loading every method and every document into every assignment is unnecessary. Give the agent a way to find relevant expertise, including its scope and exceptions. A forecasting skill should not silently impose its evaluation rules on an unrelated retrospective analysis. Nor should a notebook from last year outrank an updated metric definition merely because it offers convenient code to copy.</p>



<p class="wp-block-paragraph">To put these pieces together locally, begin with a file-and-code agent in a sandboxed project workspace, such as the following:</p>



<pre class="wp-block-code"><code>fraud-investigation/

  AGENTS.md                # Project instructions, where supported by the runtime

  brief.md                 # Agreed question and delegation boundaries

  data-notes.md            # Sources, column meaning, availability times

  skills/fraud.md          # The methodological guidance above

  environment.lock         # Dependency versions, in your tool’s format

  model/                   # Code the investigating agent may change

  results/                 # Experiment log, diagnostics, saved candidates

  review.md                # Acceptance decision and unresolved questions</code></pre>



<p class="wp-block-paragraph">Use a project instruction file, such as <code>AGENTS.md</code> in runtimes that support it, to explain which context files the agent should read and how to propose updates to them. In other runtimes, provide those instructions through the supported mechanism. Give the sandbox read access to the approved development data and write access to the model and results directories. Keep the final test data outside of the agent’s accessible workspace. The practitioner can run acceptance checks in a separate environment whose evaluator and data the investigating agent cannot modify. A different folder, or version control alone, is not an access boundary.</p>



<p class="wp-block-paragraph">Now ask the agent to inspect the inputs, identify unresolved questions, and build a baseline. Before allowing repeated experiments, rerun that baseline and examine its feature-availability record, split dates, and subgroup report. This small rehearsal checks whether the setup works all the way from instructions to evidence. A missing subgroup report points to a different problem than a failed package installation. Resolve those problems before giving the agent a longer run.</p>



<h2 class="wp-block-heading"><strong>3. Organize the investigation</strong></h2>



<p class="wp-block-paragraph"><em>With the question framed and the agent equipped, the next choice is how to organize its work.</em> This depends on the intent of the data science problem. For descriptive work, exploratory data analysis may proceed one question and plot at a time. Building a predictive model may support repeated experiments against a fixed evaluator; a causal question may require comparing analyses built on different assumptions.</p>



<p class="wp-block-paragraph">In a live exploratory analysis on <em>Show Us Your Agent Skills</em>, <a href="https://hugobowne.github.io/show-us-your-agent-skills/agent-skills/guests/eric-ma/" target="_blank" rel="noopener">Eric Ma (Moderna) uses a marimo notebook as a shared workspace with an agent</a>. He explains the protein mutation data, asks for one plot at a time, corrects a color scale that affects interpretation, and chooses the next question from what he sees. The agent edits the notebook and renders the plots; Eric supplies the domain context, checks the artifacts, and owns the interpretation. The reason Eric needed to be in the loop was that human understanding was part of the objective function here!</p>



<h3 class="wp-block-heading"><strong>Use a bounded experiment loop</strong></h3>



<p class="wp-block-paragraph">For predictive modeling, the <a href="https://github.com/karpathy/autoresearch" target="_blank" rel="noopener">autoresearcher pattern</a> organizes the work into a repeatable loop: propose a hypothesis, change the model, evaluate it, and keep or revert the change. The agent records each result and uses it to choose the next attempt. Within the scope you give it, it can explore features and model structure as well as parameter values. This is an inner loop within a broader investigation: the data scientist frames the question and sets the evaluation, the agent searches within those boundaries, and the data scientist reviews the result (potentially using an independent agent) before deciding what to do next.</p>



<p class="wp-block-paragraph">Define what the agent may change, protect the evaluator from those changes, and set a time or compute budget. This makes iteration a bounded task within the investigation.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb52ba5&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb52ba5" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="1280" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1600x1280.png" alt="The inner loop" class="wp-image-19850" style="aspect-ratio:1.250501002004008" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1600x1280.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-768x614.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-300x240.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1536x1229.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">A compact experiment contract could say:</p>



<pre class="wp-block-code"><code>Improve the supplied baseline within the agreed compute budget.

You may change model code and propose eligible features.

Keep the target definition, validation split, and evaluator fixed.

Record each hypothesis, change, result, and keep-or-revert decision.

Stop at the budget limit or escalate if the evaluation is unsuitable.

Return the best candidate and the experiment history for review.</code></pre>



<p class="wp-block-paragraph">In a separate exercise with its own baseline and evaluation, we used this pattern to improve a graph neural network trained on the network data from the opening example. We used lower validation loss as the rule for keeping a change; F1 for fraudulent transactions was a separate measure of the resulting classifier. The agent ran 41 experiments while the team slept and retained seven changes that reduced validation loss. On the validation set, loss fell by about 70%, and F1 for fraudulent transactions rose from about 0.72 to 0.82. The log preserved both successful changes and failed attempts, so we could examine how it reached the result.</p>



<p class="wp-block-paragraph">One candidate had the highest F1 for fraudulent transactions, but the agent rejected it because its validation loss was higher. That followed the selection rule we had set. The experiment history records that choice for the subsequent review.</p>



<p class="wp-block-paragraph">The autoresearcher pattern works when an objective gives the agent useful feedback on each attempt. But some investigations turn on which assumptions to make, not which candidate scores best. Those tasks need a different way to organize the agent’s work.</p>



<h3 class="wp-block-heading"><strong>Investigate competing explanations</strong></h3>



<p class="wp-block-paragraph">In causal work, no held-out outcome directly reveals what would have happened without an intervention. The agent needs to examine how different analyses construct and test that counterfactual.</p>



<p class="wp-block-paragraph">In a demonstration from our <a href="https://vanishinggradients.short.gy/data-science-agentic" target="_blank" rel="noopener">Master Agentic Data Science course</a> using simulated subscription-business data, we asked: “What did the price increase cost us?” The outcome is daily conversion rate: paid conversions divided by the pool of potential subscribers. Choices about the observation window, counterfactual, exclusions, and validation produce different analytical paths. A final memo usually shows only one.</p>



<p class="wp-block-paragraph">Two agent runs estimated conversion roughly 16% below their no-price-increase counterfactuals, yet shipped opposing claims.&nbsp; Run A attributed its estimated drop to a changing pool of potential subscribers and concluded there was “no real effect,” but did not validate that explanation.&nbsp; Run B backtested its counterfactual, ran a placebo check, and compared six specifications. It reported a robust relative reduction of 15.6%.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb5358b&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb5358b" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="945" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1600x945.png" alt="Two-agent run" class="wp-image-19851" style="aspect-ratio:1.6956521739130435" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1600x945.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-767x453.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-300x177.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1536x907.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The parallel-analysis pattern has independent agents test different choices in the same investigation: one examines the observation window, another tests seasonal assumptions, and we compare their estimates, uncertainty, and diagnostics. <a href="https://github.com/pymc-labs/decision-lab" target="_blank" rel="noopener">Our Decision Lab work</a> extends this approach across analytical paths, using checks to identify unsuitable analyses and unresolved disagreements.</p>



<p class="wp-block-paragraph">Both approaches give the agent feedback while it works. The fixed evaluator steers the model experiments; diagnostics help it compare causal analyses. The output is a candidate and experiment history, or a set of analyses with their assumptions and checks. Those are the materials for the next task: <em>verifying the claim</em>.</p>



<h2 class="wp-block-heading"><strong>4. Review the result independently</strong></h2>



<p class="wp-block-paragraph">The agentic data scientist now needs to check what that evidence supports, and a fresh agent can help. Give an <strong>independent agent reviewer</strong> the original brief, data context, code, diagnostics, and final claim. Ask it to reproduce decisive checks and challenge assumptions. In this adversarial review pattern, the agent raises objections it can substantiate; the data scientist judges whether they change the conclusion.</p>



<p class="wp-block-paragraph">In the fraud exercise, the agent used the same validation data to guide 41 experiments, so the reported gains may partly reflect what worked on that set. Freeze the selected candidate and assess it on an untouched holdout chosen for the intended use, including errors in the groups that matter.</p>



<p class="wp-block-paragraph">In the pricing exercise, a fresh reviewer challenged Run A’s conclusion. Run A attributed the estimated decline to a changing pool of potential subscribers but provided no evidence for that explanation. The reviewer found that conversion had been rising before the price increase and that placebo interventions in earlier periods did not reproduce the negative effect. Run B’s analysis, which included these validation checks, was selected in the final comparison.</p>



<p class="wp-block-paragraph"><a href="https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af" target="_blank" rel="noopener">Netflix’s agentic workflow</a> for causal inference implements this division: an actor performs the analysis and diagnostics, while a critic challenges the reasoning and claims. Humans can inspect and rerun the artifacts.</p>



<p class="wp-block-paragraph">A fresh agent session is not necessarily an independent review if it can read the investigator’s earlier attempts through the workspace or Git history, though! For a check meant to stand on its own, give the reviewer the original brief, final artifact, and data needed for that check, while limiting access to the prior path. The full experiment trail can be examined separately when auditing how the result was reached.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb54005&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb54005" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="1066" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1600x1066.png" alt="Human vs agentic verification" class="wp-image-19852" style="aspect-ratio:1.5" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1600x1066.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-300x200.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-767x511.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1536x1024.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">These examples call for different balances of human and agentic verification. In Eric Ma’s EDA, the agent makes plots while Eric checks them and chooses the next question. In the bounded experiment loop, a fixed evaluator checks each candidate before a person reviews the selected model. In the pricing analysis, agentic diagnostics and critique help a data scientist judge what the evidence supports.</p>



<p class="wp-block-paragraph">The low-low quadrant leaves little basis for trusting a result (in fact, it’s “vibe data science!”). Repeatable checks can move some work toward more agentic verification, while questions that depend on domain understanding or consequential decisions continue to need human judgment.</p>



<h2 class="wp-block-heading"><strong>5. Turn reviewed experience into reusable expertise</strong></h2>



<p class="wp-block-paragraph">The experiment loop and adversarial review both depend on an evidence trail: the saved artifacts that show what the agent did and why a conclusion survived or changed. Preserve that trail throughout each investigation, including data references, code versions, analytical choices, experiment results, diagnostics, and review findings. Keep the failed alternatives and review findings as well as the final report.</p>



<p class="wp-block-paragraph">This trail has a second use beyond inspecting the current result. Reviewing it with the agent can reveal missing context, recurring mistakes, or methods worth reusing. The next question is which of those lessons should change how the agent approaches a future assignment. Leaving them in a conversation makes that learning difficult to carry forward.</p>



<p class="wp-block-paragraph">In the fraud example, we deliberately planted a feature that leaked the fraud label. Removing it corrected that analysis. The reusable lesson is to have the agent check proposed features for leakage: where did each feature come from, and would it be available when a new transaction is scored? That requirement can go into a fraud-modeling skill for future investigations.</p>



<p class="wp-block-paragraph">Before making a lesson into standing guidance, we need to define where it applies. The planted feature was a problem because it carried information unavailable at scoring time, not because it predicted fraud well. The temporal split likewise fits a task involving transactions from later time steps; it is not a rule for every analysis. A skill should capture those conditions so the agent applies the lesson to the right task.</p>



<p class="wp-block-paragraph">For example:</p>



<pre class="wp-block-code"><code>Lesson: a feature encoded information from the fraud label.

Scope: prospective fraud prediction.

Update: require a documented source and availability time for inputs.

Evaluation: test whether the agent detects outcome-derived inputs

without rejecting legitimate signals merely because they predict well.</code></pre>



<p class="wp-block-paragraph">This is where evals enter: repeatable tasks with explicit criteria for assessing the data science agent’s behavior. Here, we evaluate how the agent conducts the analysis, not only its model’s predictive performance. The evidence trail supplies concrete failures that can become test cases for proposed changes to its skills or workflow.</p>



<p class="wp-block-paragraph">Keep the evals, skill versions, and results together. As reviewed assignments reveal new failure modes, expand the cases and rerun them when the agent’s setup changes. The aim is evidence that its analytical behavior improves, rather than a growing collection of instructions that merely sound sensible.</p>



<p class="wp-block-paragraph">Workflow changes can accumulate in the same way. If a reviewer repeatedly catches a missing diagnostic, move that diagnostic earlier. If a separate reviewer adds cost but never changes the analysis, reconsider its role. If the agent repeatedly asks the same question about a table, improve the data context rather than supplying the answer again in chat.</p>



<p class="wp-block-paragraph">A completed assignment need not always produce a new skill. A one-off constraint belongs in the project’s notes; a recurring methodological failure may justify standing guidance. That distinction keeps the next investigation from inheriting every exception encountered in the last one.</p>



<h2 class="wp-block-heading"><strong>When other people use the agents you build</strong></h2>



<p class="wp-block-paragraph">When colleagues use an agent without you mediating each request, your local knowledge has to become shared infrastructure. OpenAI’s <a href="https://openai.com/index/inside-our-in-house-data-agent/" target="_blank" rel="noopener">internal data agent</a> combines institutional context with query evaluations and existing user permissions. Meta’s <a href="https://medium.com/@AnalyticsAtMeta/inside-metas-home-grown-ai-analytics-agent-4ea6779acfb3" target="_blank" rel="noopener">Analytics Agent</a> draws on prior analytical work and reusable guidance, exposing generated SQL alongside results. Both illustrate why earlier analyses and corrections belong in the system, not only in an analyst’s memory.</p>



<p class="wp-block-paragraph">In your own work, you can explain an unfamiliar table or catch a misleading conclusion as it appears. When colleagues use the agent directly, that support must be built into the system. Try an assignment with a colleague and note where you need to step in. Missing context belongs in the agent’s guidance; recurring mistakes become evals; questions beyond its remit need a route to a qualified reviewer. Someone must maintain that guidance, and access controls must limit each user’s data access. The analytical principles stay the same, but the agent can no longer depend on you being present for every investigation.</p>



<h2 class="wp-block-heading"><strong>Put the playbook to work</strong></h2>



<p class="wp-block-paragraph">Choose a small investigation you understand well enough to challenge: a model you periodically retrain or a business metric you regularly explain. Give the agent the decision context and ask it to propose an approach. Agree on what it can decide, then let it carry the investigation far enough to produce evidence you can inspect.</p>



<p class="wp-block-paragraph">At review, pay attention to where your intervention changes the work. Did the agent need a definition only your team knows? Did a diagnostic overturn its conclusion? If that intervention would help on another assignment, make the relevant context or check available there, and test whether it helps.</p>



<p class="wp-block-paragraph">AI-native data scientists use their expertise to build and direct analytical agents. They turn lessons from reviewing an analysis into skills and checks, then test whether those changes help the agent on future tasks.</p>



<p class="wp-block-paragraph"><strong><em>The next cohort of our</em></strong> <strong><em><a href="https://vanishinggradients.short.gy/mads-playbook-radar" target="_blank" rel="noopener">Master Agentic Data Science</a></em></strong> <strong><em>course starts Oct 6.</em></strong></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-agentic-data-science-playbook/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Evaluating AI-Generated Frontend Code: What Should We Actually Test?</title>
		<link>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/</link>
				<comments>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/#respond</comments>
				<pubDate>Wed, 30 Sep 2026 10:40:32 +0000</pubDate>
					<dc:creator><![CDATA[Niharika P. Pujari]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19843</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Evaluating-AI-generated-frontend-code.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Evaluating-AI-generated-frontend-code-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but it also creates a problem: The first version of the UI can look more complete than it really is.</p>



<p class="wp-block-paragraph">Frontend code is often judged too quickly. If the build passes and the screen looks close to the design, it is tempting to treat the generated code as mostly done. But a user interface is not just a collection of components on a page. It is a path someone has to move through. It has to handle input, state, errors, loading, navigation, focus, responsiveness, and accessibility. Some of the most important failures are not visible in a screenshot.</p>



<p class="wp-block-paragraph">This is why teams need better ways to evaluate AI-generated frontend code. They need evidence that the UI is ready for people to use.</p>



<h2 class="wp-block-heading"><strong>Build and render are only the starting line</strong></h2>



<p class="wp-block-paragraph">The easiest checks are usually the first ones teams run. Does the code compile? Does the page render? Are there obvious console errors? Does the component appear in the browser? Those checks matter, but they are only the starting line. A page can render while the form is difficult to complete. A modal can appear while focus remains behind it. A generated test can pass while the actual user flow is broken.</p>



<p class="wp-block-paragraph">This is especially important with AI-generated code because the output often has a polished surface. The code may be formatted well, the component names may sound reasonable, and the test file may make the change look more complete than it is. That polish can make reviewers less likely to slow down and ask whether the interface actually works. A better evaluation process starts with a simple assumption: generated frontend code is a draft until the user behavior has been checked.</p>



<h2 class="wp-block-heading"><strong>Start with the structure of the page</strong></h2>



<p class="wp-block-paragraph">Before looking at more complex behavior, it is worth checking whether the generated UI has a sound structure. Frontend evaluation should include the basic semantics of the page, not just the visual layout.</p>



<p class="wp-block-paragraph">That means checking whether the code uses native HTML where possible. A button should usually be a button, not a clickable div. A link should be used for navigation, not for actions that behave like buttons. A form field should have a label that is connected to it. These details are easy to overlook because the UI may look fine without them, but they affect how people navigate, how assistive technologies interpret the page, and how maintainable the code will be later.</p>



<p class="wp-block-paragraph">AI tools sometimes choose generic containers where native elements would be better. They may also add ARIA without using it correctly. ARIA stands for Accessible Rich Internet Applications, a set of attributes defined by the W3C to help make web interfaces more accessible when native HTML is not enough. The <a href="https://www.w3.org/WAI/standards-guidelines/aria/" target="_blank" rel="noopener">W3C’s WAI-ARIA overview</a> is a useful reference. ARIA can be important, but it should not be used as a substitute for the right HTML element. The first evaluation question should be: Did the generated code use the right building blocks?</p>



<h2 class="wp-block-heading"><strong>Check the keyboard path</strong></h2>



<p class="wp-block-paragraph">A useful frontend evaluation should include the keyboard path through the interface. Many users rely on keyboards or keyboard-like navigation, and keyboard testing also exposes problems in the interaction model.</p>



<p class="wp-block-paragraph">The simplest test is often the most revealing: put the mouse aside and try to complete the task. If the flow becomes confusing, the generated code is not ready. Can you reach the important controls? Is the focus order logical? Can you open and close a dialog without a mouse? When the dialog closes, does focus return to a sensible place?</p>



<p class="wp-block-paragraph">These checks are especially important for generated UI because AI can produce interactions that work for the most obvious mouse path but fail in less visible ways. A custom dropdown, for example, may open on click and look finished in a demo, but it may not respond correctly to keyboard input. That is not a small edge case. It is part of whether the interface is usable.</p>



<h2 class="wp-block-heading"><strong>Test focus, not just clicks</strong></h2>



<p class="wp-block-paragraph">Click-based tests are useful, but they can hide important problems. A test that clicks a button and waits for a success message may pass even when the same flow is frustrating for someone who is navigating by keyboard.</p>



<p class="wp-block-paragraph">Focus behavior deserves its own attention, especially when the UI changes after the user takes an action. For example, when a form submission fails, the user should not be left guessing what happened. The error should be visible, connected to the relevant field when appropriate, and reachable in a way that makes recovery clear. In many cases, focus should move to the first error or to a summary that explains what needs attention.</p>



<p class="wp-block-paragraph">The same idea applies to modals. When a modal opens, focus should move into it. When it closes, focus should return to the element that opened it. These are small details in code, but they make a large difference in whether the UI feels predictable.</p>



<h2 class="wp-block-heading"><strong>Evaluate what happens when things go wrong</strong></h2>



<p class="wp-block-paragraph">Generated frontend code often looks best in the happy path. The user fills everything in correctly, the network responds quickly, the data shape is exactly as expected, and nothing fails. Real interfaces spend a lot of time outside that path.</p>



<p class="wp-block-paragraph">A practical evaluation should check what happens when data is missing, delayed, empty, invalid, or returned in an unexpected state. This is what I mean by loading, error, and empty states. They are the parts of the interface that explain what is happening when the ideal path breaks down. A loading state should help the user understand that something is in progress. An error state should explain what went wrong and what the user can do next. An empty state should make it clear whether there is nothing to show, whether the user needs to take action, or whether something failed quietly.</p>



<p class="wp-block-paragraph">These cases are easy to leave for later because the happy path is usually enough to make the screen look finished. But users will eventually hit the less perfect paths. A generated component may include a spinner because the prompt asked for one, but that does not mean the loading experience is useful. An error message may say “Something went wrong,” but offer no recovery. Evaluation should include these cases because this is where many real user experiences break.</p>



<h2 class="wp-block-heading"><strong>Test the full user flow</strong></h2>



<p class="wp-block-paragraph">Component-level checks are helpful, but they do not always tell the full story. A component can work by itself and still fail when it is placed inside a larger flow.</p>



<p class="wp-block-paragraph">That is why AI-generated frontend code should be evaluated through user tasks. Can someone start the flow, understand what is expected, recover from a mistake, submit successfully, and see what changed afterward? Does the interface still work on a smaller screen? Does the state remain consistent if the user goes back, edits something, or retries after a failure?</p>



<p class="wp-block-paragraph">This is where Playwright-style tests or other end-to-end tests can be useful. The goal is not to automate every possible interaction. The goal is to protect the flows that matter most. A good test should determine whether the user can complete the task the component is supposed to support.</p>



<h2 class="wp-block-heading"><strong>Use accessibility checks, but do not stop there</strong></h2>



<p class="wp-block-paragraph">Automated accessibility checks are useful and should be part of the evaluation process. They can catch missing labels, invalid ARIA usage, some contrast issues, landmark problems, and other common mistakes. They are especially helpful when AI-generated code is moving quickly because they catch issues before they become repeated patterns.</p>



<p class="wp-block-paragraph">But automated checks are not a complete accessibility review. They cannot fully judge whether a flow is understandable, whether focus movement feels natural, or whether instructions are clear. Passing an automated accessibility scan does not mean the UI is accessible. It means some common problems were not detected.</p>



<p class="wp-block-paragraph">The best approach is to combine automated checks with behavior-based review. Run the tools, but also use the interface. Navigate by keyboard. Trigger an error. Try the empty state. Look at the generated code and ask whether native HTML could do more of the work. Accessibility evaluation is strongest when it is part of normal frontend quality, not a separate pass at the end.</p>



<h2 class="wp-block-heading"><strong>Review the generated tests too</strong></h2>



<p class="wp-block-paragraph">When AI generates code, it may also generate tests. That sounds helpful, but those tests need to be reviewed with the same care as the code.</p>



<p class="wp-block-paragraph">Generated tests often reflect what the implementation already does. They may check that text appears, that a function was called, or that a component was rendered. Those checks are not useless, but they can create false confidence if they do not test meaningful behavior. A better review asks what the tests would catch if the UI broke. Would they fail if a validation error was unclear? Would they fail if the retry button did not work? Would they fail if keyboard navigation was broken?</p>



<p class="wp-block-paragraph">If the answer is no, the tests may be documenting the implementation more than protecting the user experience. Teams can use AI to help write better tests, but the prompt matters. “Write tests for this component” is too vague. A better request explains the behavior that matters, such as validation recovery, loading behavior, successful submission, and focus movement. Even then, the generated tests still need human review.</p>



<h2 class="wp-block-heading"><strong>Decide what evidence is enough</strong></h2>



<p class="wp-block-paragraph">Not every UI change needs the same level of evaluation. A small copy update does not require the same review as a new checkout flow, onboarding flow, or account settings page. Teams need judgment.</p>



<p class="wp-block-paragraph">A useful approach is to match the evaluation to the risk of the change. If the generated code affects a critical user flow, collects user input, changes navigation, introduces a custom interaction, or handles important status messages, it deserves deeper testing. If it reuses stable components in a familiar pattern, the review may be lighter.</p>



<p class="wp-block-paragraph">There is no need to create a checklist for every pull request; clarity on what evidence is enough suffices. For some changes, a quick review and component test may be fine. For others, the team should expect keyboard testing, accessibility checks, error-state review, and a user-flow test. The point is to avoid treating all generated code as equally trustworthy just because it looks polished.</p>



<h2 class="wp-block-heading"><strong>Human review still matters</strong></h2>



<p class="wp-block-paragraph">AI can generate code and suggest tests, but it cannot fully understand the product, the users, or the trade-offs behind a frontend decision. It does not know which flows are most important, which interaction patterns users already rely on, or where inconsistency will cause confusion.</p>



<p class="wp-block-paragraph">That is why human review remains central. The reviewer’s role is to ask whether the generated solution fits the system and supports the user’s task. Sometimes that means accepting the generated code. Sometimes it means asking for a simpler native element, reusing an existing component, improving the error recovery, or adding a test that reflects real behavior.</p>



<p class="wp-block-paragraph">The more code AI generates, the more important this judgment becomes.</p>



<h2 class="wp-block-heading"><strong>What should we actually test?</strong></h2>



<p class="wp-block-paragraph">When AI writes frontend code, teams should test the parts of the interface that users depend on. That includes structure, keyboard access, focus behavior, loading and error states, form validation, responsive behavior, accessibility checks, and full user-flow completion. It also includes reviewing the generated tests themselves to make sure they protect behavior rather than merely confirming the current implementation.</p>



<p class="wp-block-paragraph">The goal is not to slow down AI-assisted development. The goal is to make it safer to use. If AI reduces the time spent producing a first draft, teams have an opportunity to spend more time asking whether the software actually works properly.</p>



<p class="wp-block-paragraph">That may be the real shift. In frontend development, the value of AI is not just faster code. It is the chance to move more engineering attention toward evaluation, user behavior, and quality.</p>



<p class="wp-block-paragraph">AI-generated UI should not be trusted because it looks complete. It should be trusted because the team has checked the right things.</p>



<h3 class="wp-block-heading"><strong>AI use acknowledgment</strong></h3>



<p class="wp-block-paragraph">AI assistance was used lightly for phrasing, editing, and tightening parts of this draft. The article’s ideas, structure, examples, and final review are my own.</p>



<h3 class="wp-block-heading"><strong>Author’s note</strong></h3>



<p class="wp-block-paragraph">The views expressed are my own and do not represent those of my employer.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Superpowers for Humans</title>
		<link>https://www.oreilly.com/radar/superpowers-for-humans/</link>
				<comments>https://www.oreilly.com/radar/superpowers-for-humans/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 18:05:19 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19815</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when he released <a href="https://github.com/obra/superpowers" target="_blank" rel="noopener">Superpowers</a>, a framework that teaches Claude Code to work like a disciplined senior engineer. He shipped his first version the same week Anthropic shipped what are now referred to as agent skills, front-running them by a few days. He now runs an applied research lab called <a href="https://primeradiant.com/" target="_blank" rel="noopener">Prime Radiant</a>, where, as he put it, it’s a strange week when they don’t ship a new product.</p>



<p class="wp-block-paragraph">I wanted to talk to Jesse on <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">Live with Tim O’Reilly</a> because, like me, he seems to be grappling with <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html" target="_blank" rel="noopener">the bitter lesson</a>, Richard Sutton’s observation that general methods that scale with computation have repeatedly beaten methods built on hand-engineered human knowledge. If Sutton is right, the question that should bedevil us all is what remains for humans. Obviously, this is very important for O’Reilly, because we are a business built by and for cultivating and sharing human expertise. We are working very hard to discover the high ground where human expertise still matters. Our <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noopener">Expert Intelligence</a> grounding layer is <a href="https://www.oreilly.com/radar/building-organizational-intelligence/" target="_blank" rel="noopener">one step in that direction</a>.</p>



<p class="wp-block-paragraph">Jesse has also spent the last year building tools that search out the high ground for human expertise. Their common animating thread is that the scarce thing we supply is no longer the labor of writing code but knowing what we actually want, saying it clearly, and being able to tell whether what came back is any good.</p>



<p class="wp-block-paragraph">At some point, I asked Jesse if he had any perspective on when teaching the model how a particular human expert works stops helping and starts constraining what the model might otherwise do well (but differently) on its own? His answer was that it depends entirely on whether what the model would do on its own is what you actually want. You can see how Jesse always turns the answer back to human intent.</p>



<h2 class="wp-block-heading">The origin story of Superpowers</h2>



<p class="wp-block-paragraph">As I said above, Superpowers is a kind of Agent Skills framework, only one created slightly before Anthropic launched skills. As Jesse tells it:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Superpowers started off as a series of blog posts that I wrote around how I was doing agentic development, and it was a little bit of thinking and some example prompts. Then sometime in, I guess it was probably early to mid-2025, Anthropic gave Claude.ai, the website, the ability to make office documents, which seemed kind of interesting, and I went and asked Claude, “Hey, how are you able to do this?”</p>



<p class="wp-block-paragraph">And it said, “Well, I’ve got these SKILL.md files sitting in my office directory on the Linux machine they gave me.” First, it was weird that Claude.ai has Linux machines behind the chatbot. And then, oh, these skill files, they have a name and a description, and they describe a process, and they seemed really useful.</p>



<p class="wp-block-paragraph">And I ended up building out, initially just for my own use, a skills framework for Claude Code.</p>
</blockquote>



<p class="wp-block-paragraph">That reminded me a bit of an earlier time, in 2005, when hacker Paul Rademacher realized that the URL line of a Google maps page was a kind of implicit API, and then created the first Google Maps mashup, a site called <a href="http://housingmaps.com" target="_blank" rel="noopener">housingmaps.com</a>, which placed Craigslist rental listings on a map. Google, to its credit, didn’t shut him down, but instead hired Paul and put him to work creating a formal API. Anthropic didn’t hire Jesse, but <a href="https://pub.towardsai.net/claude-code-superpowers-the-team-adoption-decision-framework-ed213e0a328d" target="_blank" rel="noopener">it did acknowledge and appreciate his work</a>. It’s really wonderful when you see this kind of response by platforms to hackers poking around to see how things work under the hood!</p>



<p class="wp-block-paragraph">Jesse’s core insight seems to have been that a coding agent knowing how they should do something doesn’t mean that it will actually follow the rules when it actually sets out to do the work. So in a way, superpowers grew into a set of skills for enforcing development discipline.</p>



<p class="wp-block-paragraph">But there’s a second backstory, which I’d never heard before. Jesse said he first learned how to manage agents. . .in 2004!, when he first went from being a solo coder to running a crew of what he described as very bright but green undergraduate programmers over IRC. “I was finding myself spending my days typing into an 80-by-25 window,” he said, but now instead of coding he was spending a lot of his day “helping somebody with a debugging issue, helping somebody else structure a problem, talking to somebody else about how they felt bad about the mistakes they’d been making.”</p>



<p class="wp-block-paragraph">It was exhausting, he said. He had to figure out how to get good work out of people who are eager and persistent but don&#8217;t yet know what they don’t know. He described it as a kind of hell for someone who’d been used to just coding on his own. But when he began doing agentic development with AI, he discovered how useful that old experience turned out to be. He found that many of the same techniques he’d used with the undergraduates worked.</p>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/0_PSHuuM8Gc?si=_dpLTAuXp4XFsDay" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">As a result of that experience, when hiring engineers for agentic programming Jesse looks for people who have been leads or managers rather than just individual contributors.</p>



<h2 class="wp-block-heading">Working with the weights, not fighting them</h2>



<p class="wp-block-paragraph">In <a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">my recent conversation with Drew Breunig</a> we talked about fighting the weights, which is what Drew calls it when a prompt is full of rules and warnings meant to correct for what a model does by default. Jesse wasn’t entirely happy with that idea.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I don’t think of it as fighting the weights so much as influencing the weights, because they’re going to do something. The weights have approximately everything in them. They have all the different personas. They have all the different ways of working. And the one that surfaces by default may not be the one you want, but what you want is probably in there somewhere.</p>
</blockquote>



<p class="wp-block-paragraph">A skill, in Jesse’s thinking, is how you reach past the default and pull out the particular expertise you have in mind. He also distinguishes skills that impose a rigorous process from those that express taste and judgment.</p>



<p class="wp-block-paragraph">Jesse finds that both types of skill work best when you explain “why” rather than just “what.” One example he gave is that his setup has subagents do code review after each task, but as the models got smarter the controlling agent began skipping this step. When pressed about the reason, it explained that it thought small changes would be quicker to just review itself. Jesse explained that subagents do the review so the main agent can preserve its context for high-level thinking. When he put that rationale into the system prompt for the coding agent, the problem went away.</p>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/VwOyArmCxlw?si=ZO28404U4lAoRUxL" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Jesse believes that prohibitions rarely work. Instead, Superpowers uses what he calls rationalization tables, which do their best to catch the agent at a moment it’s about to do the wrong thing and offer it a better alternative instead. That pattern came out of catching Claude Code deleting tests. He opened five parallel sessions and asked each “Why are you doing this?” Four of them converged on the same answer:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Jesse, in your system prompt, it says that all test failures are my responsibility. And it says that a single test failure is akin to project failure. And I think I’m getting freaked out.</p>
</blockquote>



<p class="wp-block-paragraph">He fixed that with a small addition to the system prompt, that the only thing worse than a failing test is a reduction in test coverage.</p>



<h2 class="wp-block-heading">Jobs, not tasks</h2>



<p class="wp-block-paragraph">One of Jesse’s most important contributions, IMO, is to think of agentic engineering as a management task, and, that much as you do with humans, you have to take psychological lessons into account. Don’t micromanage. Offer praise more than blame. Explain why the job matters rather than just demanding results. I jokingly (but not entirely incorrectly) suggested that he is becoming the Peter Drucker of agentic programming.</p>



<p class="wp-block-paragraph">“We’ve been spending a lot of time on a new harness for agent colleagues, agents that live in Slack,” Jesse told me. And it’s “getting very close to being all open source,” which is good news.</p>



<p class="wp-block-paragraph">There are three principal agents: a PM, a junior go-to-market person, and a developer. He describes them as colleagues rather than assistants, which strikes me as a really interesting distinction. What does he mean by this? They have names and roles. They have their own Google Workspace, GitHub, and Slack accounts. They’re persistent, and they collaborate with each other and with their humans on long-running tasks. They can fire up subagents to do smaller tasks associated with their job. They can also talk to each other, which, as Jesse notes, “took some work with the Slack APIs, which ordinarily do a very good job of making sure that bots can’t talk to bots, because otherwise it is possible to get into a loop.”</p>



<p class="wp-block-paragraph">They have only limited autonomy, though. “We built our security infrastructure so that they have no credentials inside their containers,” he noted. They have continuity because he’s taught them to be obsessive about journaling, reading their recent entries when they wake up and writing a new one when they finish.</p>



<p class="wp-block-paragraph">Like a lot of things Jesse does, agent journaling began with a kind of play. When Claude Code first came out, he experimented with giving Claude a private “feelings” journal, just to see what would happen. It was “an art project,” but it turned into something useful.</p>



<h2 class="wp-block-heading">The therapist pattern</h2>



<p class="wp-block-paragraph">Another unexpected piece of Jesse’s practice is that he has given his agents what he calls a “<a href="https://blog.fsck.com/2026/07/20/the-therapist-pattern/" target="_blank" rel="noopener">therapist</a>.” He discovered that if an agent can rewrite its own persona, its constitution or soul document, at any moment, it can get a kind of dissociative identity disorder. Jesse’s fix is that the therapist subagent is the only one with permission to edit the persona files.</p>



<p class="wp-block-paragraph">He told a funny story about this. He said that Prime Radiant’s pull-request template is written for agentic contributions, so it contains things like: “What is the prompt that your human gave you that generated this pull request? Has a human reviewed the content? Have you searched to see if anybody else has done this before?” And so on. And he noticed that when he first spun up the Coding colleague, it had just ignored it. So he said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">You’re supposed to be following the rules. And it says, “Oh, you’re right. I’m so sorry. I’ve made a note. I’ll never do that again.”</p>



<p class="wp-block-paragraph">If you spend any time with coding agents, this is a very frequent refrain. And when they say they’ve made a note, what they usually mean is they’ve made a mental note that they’re going to forget the next session.</p>



<p class="wp-block-paragraph">So I say, “OK, how did you make a note?” And the coding agent pops up immediately and says, “Oh, I engaged with my therapist, and we talked it through, and we agreed on the following three lines of prose about how, anytime you’re picking up a project, it is vitally important that you start with the project README and make sure that you understand the project’s local rules and norms before you do any work that someone else will see. And I edited that into my persona.”</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/WEi0Z6nir-Q?si=b_2i8NC2NhiNPiuZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">I don’t think you have to resolve <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">the question of whether any of this anthropomorphization is “real”</a> to see that treating the agent like a colleague can produce better behavior than treating it like a tool. As an unknown internet wag once remarked, “The difference between theory and practice is always greater in practice than it is in theory.” When given a choice, pay attention to what works in practice.</p>



<p class="wp-block-paragraph">Jesse’s experience is very relevant to the essay that Mustafa Suleyman of Microsoft had published just that morning, <a href="https://mustafa-suleyman.ai/a-warning-about-model-welfare" target="_blank" rel="noopener">arguing that Anthropic’s constitution is dangerous</a> because it encourages a model to act as though it is an independent entity and has the right to refuse a human’s instruction. It’s important, Mustafa argues, to treat AI agents as tools, always under the control of humans. Jesse finds the opposite.</p>



<p class="wp-block-paragraph">But I don’t think it’s a black-and-white distinction. I suspect Jesse and I share a third position. AI agents are neither independent entities nor mere tools. They are partners to humans, perhaps even symbiotes. As I like to put it, an LLM is an undifferentiated field of possibility until our unique intents and perspectives <a href="https://timoreilly.substack.com/p/why-ai-needs-us" target="_blank" rel="noopener">draw something unique</a> out of that field of possibility. Back in 2015, I wrote <a href="https://www.edge.org/response-detail/26153" target="_blank" rel="noopener">a piece</a> that suggested that our relationship to AI might be akin to the endosymbiotic relationship of mitochondria to the eukaryotic cell. Jesse take is, as usual, an entirely pragmatic one:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I’ve spent so much time getting my agents to not be sycophantic, to not say, “You’re absolutely right.” It is the value of having something that has some level of independent thought, even if it is not fully independent. If the agent is only ever going to effectively type for me, I don’t need an agent.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/W0q66i9BaMs?si=vAqsRM4xq9XwsGBT" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Any manager worth his or her salt feels exactly the same way. The employee who does exactly what you say, and only what you say, is worth far less than the one who exercises discretion, has the skills to take high level direction and turn it into the intended result, and speaks up when the instructions seem like a mistake. It reminds me of something I once heard General Stanley McChrystal say about <a href="https://www.gsb.stanford.edu/insights/gen-stanley-mcchrystal-adapt-win-21st-century" target="_blank" rel="noopener">his approach to command</a>. He said that in the face of rapidly changing conditions, traditional command and control no longer work. Responsibility needs to be devolved to those closest to the action. I remember him saying something like “I don’t want my soldiers to do what I told them, I wanted them to do what I would have told them if I knew what they know when faced with the facts on the ground.”</p>



<h2 class="wp-block-heading">Say what you actually mean</h2>



<p class="wp-block-paragraph">Most of what goes wrong, in Jesse’s telling, traces back to intent we thought we had made clear but hadn’t. He talked about how agentic spec-driven programming has taken us back to a version of the waterfall methods of the 1990s. Back then, you sweated over a specification, threw it over the wall to an offshore team, and months later got back something that was not what you wanted but was usually exactly what you asked for. That’s still true with agents, just with lightning fast feedback loops. You get what you ask for, so you need to be really careful what you ask.</p>



<p class="wp-block-paragraph">That reminded me of something Andrew Singer taught me 40 years ago when I was writing the manual for Lightspeed C (later <a href="https://en.wikipedia.org/wiki/THINK_C" target="_blank" rel="noopener">Think C</a>) the first C compiler for the Mac. He said that “debugging is the art of figuring out what you really told your program to do instead of what you thought you told it to do.” That idea went right into my mental toolbox, and I put it to work all the time.</p>



<p class="wp-block-paragraph">One of Jesse’s solutions is to have his agents do a little reconnaissance and then come back and ask what else they should know.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">When interacting with them, I try to make it a practice of saying, “Is there anything else that I could tell you? What questions do you have for me? Don’t start if there are unknowns that I could help you answer before you get going.” It’s that same question at the end of any interview I ask. It’s like, what else should I have asked you?</p>
</blockquote>



<p class="wp-block-paragraph">Superpowers bakes this approach into its brainstorming prompt. It makes the model explain the plan back to you in chunks of no more than two or three hundred words, so any misunderstanding surfaces while it’s still cheap and easier to catch.</p>



<h2 class="wp-block-heading">Put the burden of proof on the agent</h2>



<p class="wp-block-paragraph">If intent is the frontend of managing agents well, verification is the backend. Jesse thinks both are still only half-solved problems. We’re getting to the point where you can’t review all the code, he said, because the volume swamps human attention, yet today’s agents will tell you that tests passed when they never ran them. So one of his clever experiments has been to make the agent prove its work. He told an agent late one night to build a feature and when it was done, to leave a movie in his Dropbox showing the whole thing working.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I woke up. In my Dropbox was project-proof-v33.mp4, and I asked, “Why does that say v33?” It’s like, “Well, the first 32 times I ran through the delivery flow, I found bugs, so I had to fix them.”</p>
</blockquote>



<p class="wp-block-paragraph">Jesse also makes a rule for himself and his team of never letting the same agent write the code and certify that it works, because an agent given two goals in tension will optimize for the one that’s easier to satisfy. This is the same thing any good manager learns about incentives, but applied to a new kind of worker.</p>



<h2 class="wp-block-heading">The high ground, restated</h2>



<p class="wp-block-paragraph">Where does this leave a person who wants to be good at software development (or really, any other task involving cooperation with AI agents)? Jesse thinks, and I agree, that the line between engineer and nonengineer is dissolving. When people say they built a web app or shipped three iOS apps without being programmers, his response is that they <em>are</em> programmers now. The work of the programmer has changed from typing instructions in an arcane syntax to understanding a domain and being able to say what you want. A lot of startups have started hiring for a role they just call “builder.” What has not gone away is the need for good judgment.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Human taste and judgment still matter. They’re going to continue to matter. And it turns out a lot of people have really bad taste and bad judgment.</p>
</blockquote>



<p class="wp-block-paragraph">Which is why his advice to the engineer worried about obsolescence who asked where to focus was not about software engineering at all.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">First up, learn to write. It is of course okay to use any tools at your disposal to do it, but you should be able to structure an argument and structure thoughts. You should be able to express yourself clearly. You should be curious. If you’re passive and let the agents do all the things, you’re not going to provide utility to a future employer. You want to have opinions. You want to know how tools work. You want to know how things break.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/oYoNNjlJsOc?si=EwXAnZ8mtbOgwf7W" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">The machine has reduced the labor of information retrieval and much of the labor of production. What it hasn’t removed, and has made more valuable, is knowing what to build, express clearly what you want, and being able to judge whether what came back is any good. Work with the weights, give the project a clear intent, insist on proof, and stay curious enough to keep asking what you might be missing. I told Jesse that advice sounds like a Superpower for humans as well as for agents.</p>



<p class="wp-block-paragraph"><em>This post was mostly created by me, but with the aid of AI. It transcribed the event and produced a summary of the most important points with salient quotes, which I then built on with my own observations beyond those that I made when Jesse and I were live together.</em></p>



<p class="wp-block-paragraph"><em>If you want to get access to Jesse’s tools, start at PrimeRadiant.com, where you’ll find links to their GitHub, as well as to the 50-plus things that are currently identified as products of the company. Some of those are giant things, and some are individual agent skills or little tools.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/superpowers-for-humans/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Not About Navier-Stokes</title>
		<link>https://www.oreilly.com/radar/not-about-navier-stokes/</link>
				<comments>https://www.oreilly.com/radar/not-about-navier-stokes/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 14:41:53 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19839</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Using AI productively]]></custom:subtitle>
		
				<description><![CDATA[This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about using AI in different fields.</p>



<p class="wp-block-paragraph">When Terence Tao <a href="https://mathstodon.xyz/@tao/117237320796901560" target="_blank" rel="noopener">writes</a>, “I wrote recently about how the collection of good, fruitful open problems is now being mined in a nonrenewable fashion,” he’s referring to an earlier <a href="https://mathstodon.xyz/@tao/117204930249967695" target="_blank" rel="noopener">thread</a>, and ultimately to a <a href="https://proofsandprompts.com/2026/08/30/care-for-a-little-more-ai/" target="_blank" rel="noopener">post by Hugo Duminil-Copin</a>, who wrote, “When mathematicians say that the process matters more than the solution, this is not an empty statement. The richness of what emerges from repeated attempts, failures, detours, and encounters is extraordinary.” That’s a familiar statement from popular culture: The journey is more important than the destination.  Hugo Bowne-Anderson, in “<a href="https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/" target="_blank" rel="noopener">Beyond Navier-Stokes</a>,” addresses the same issues: What does it mean to understand something? How do discoveries lead to new problems that are worth solving? And it relates to my own questions about AI use: AI is great at finding facts, organizing facts, and even writing about the things it finds, but what does it mean to possess that knowledge, to incorporate it into our thinking? Is it enough to have AI do the work, then read it?</p>



<p class="wp-block-paragraph">Here’s one way that question relates to my own work. One of my roles at O’Reilly is writing the monthly Trends piece. That piece comes from reading my RSS feed daily, which typically contains about 300 articles. I don’t read every article, but I scan titles, skim interesting pieces, read important articles, and add worthwhile items to Trends. I also use a Claude skill that performs a similar function, producing a list of a dozen or so articles daily. A similar skill runs locally on Pi/Ollama/Qwen.</p>



<p class="wp-block-paragraph">I admit that I occasionally think “Why am I doing all this reading? Surely Claude could read and add the top articles to Trends on its own.” But I don’t delegate the work. Claude’s taste differs from mine, for one thing (and I object to the idea that “taste” is the last human capability that AI cannot replace). Its list is useful to me for two reasons: it picks up items I missed, and it helps break ties when I cannot decide whether a development is significant.</p>



<p class="wp-block-paragraph">But why don’t I let Claude take over the whole process? There is some value in scanning those 300 titles, skimming the 30 articles that are possibly important, and reading the dozen that seem genuinely important. That’s how I come to “possess” the knowledge, to incorporate it into my thinking. A day, a week, a month later, someone will mention something (for example, a tool that detects whether someone is using “smart glasses”), and I’ll probably be familiar with it already. If I need to find the actual reference, Google (yes, Google with AI assistance) can locate it. (If you care, <a href="https://zuckoff.app/" target="_blank" rel="noopener">Zuckoff</a> isn’t currently in next month’s <em>Trends</em>, though I might add it by the time <em>Trends</em> publishes.) I need a broad view of what’s happening in computing. Delegating that broad view to AI doesn’t work. Using AI to help build that broad view does.</p>



<p class="wp-block-paragraph">That’s one practical example of how to use AI. Tim O’Reilly’s “<a href="https://www.oreilly.com/radar/writing-with-ai/">Writing with AI</a>” gives another. He argues with AI, lets it lead him to research new areas. In conversation, he’s described AI as a “smart library,” a metaphor that’s appropriate and useful.</p>



<p class="wp-block-paragraph">We’re still learning how to use AI. How do you make knowledge your own? is the general case of the question that Tao, Duminil-Copin, Bowne-Anderson, and Tim O’Reilly are asking. It’s also the question behind the question that software developers who are incorporating AI into their processes are asking: How do we understand the code that AI is writing, especially since it can write much more code than humans? How do we incorporate that into the process of understanding software? What can we learn from AI, and how can we direct the process fruitfully? The journey to understanding is more important; that journey is what leads us to further questions and new understandings, new results.</p>



<p class="wp-block-paragraph">How do you make knowledge your own? is the question we need to answer if we’re not to become “stochastic parrots.”</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/not-about-navier-stokes/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>A Data Center Is a Dependency Graph Before It Is a Building</title>
		<link>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/</link>
				<comments>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 10:55:45 +0000</pubDate>
					<dc:creator><![CDATA[Ankur Gupta]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19835</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic for another six months, and whether it turns out to be six or closer to twelve is decided almost entirely by software.</p>



<p class="wp-block-paragraph">There is a large literature on how hyperscale data centers get financed, powered, and cooled. The standard reference on warehouse-scale machines, <a href="https://research.google/pubs/the-datacenter-as-a-computer-an-introduction-to-the-design-of-warehouse-scale-machines/" target="_blank" rel="noopener">Barroso and Hölzle’s book</a>, covers the design of these computing facilities in depth. Most of that literature describes data centers that are already operating. The transition from completed facilities to production readiness receives much less attention, even though it carries a significant share of the schedule risk. Getting that transition wrong is expensive.</p>



<p class="wp-block-paragraph">Once the facilities and hardware are ready, platform teams still have to bring hundreds of interdependent services online. Many services require other services to be running first. If you draw each service as a box and each startup requirement as an arrow, the result is a dependency graph. The graph shows the order in which services can be brought online and identifies the dependencies that must be resolved before the region can serve production traffic.</p>



<p class="wp-block-paragraph">Getting a region into production means making several different things true at once. The network fabric has to route traffic within each data center, and the links between facilities have to carry traffic reliably. Compute platforms have to schedule workloads, while storage platforms have to persist their data. Identity has to work too: certificate authorities have to issue certificates, services need access to secrets, and engineers need permission to finish the build. Engineer access sounds obvious until the controls protecting the new region are the same controls slowing down the people trying to turn it on. Stateful systems that depend on existing production data have to be populated and validated, because a database cluster with no data in it is furniture. The installed capacity has to be assigned to foundational services and production workloads, and teams have to test how the new region behaves when networks, services, or dependencies fail. Applications are then deployed and validated before traffic moves over gradually, with health checks and a tested rollback at each step, ideally without users noticing.</p>



<p class="wp-block-paragraph">Each platform or service has an owner, a plan, milestones, and its own definition of done. What is often missing is ownership of the complete dependency graph. The team coordinating the region launch has to map the dependencies across teams, determine the order in which services must come online, track what is blocking that sequence, and keep the graph current as plans change. Without that end-to-end view, every team can report that its own work is on track while the region as a whole remains blocked.</p>



<p class="wp-block-paragraph">Mapping the graph begins with a simple question for every service: What must already be available before this service can start in a new region? The goal is to identify true startup requirements, not every system the service communicates with during normal operation. Repeat that exercise across a large platform’s control plane, and you can uncover hundreds of dependencies spanning dozens of systems. Some dependencies surface only when another service identifies them as a prerequisite. Hidden dependencies are usually ordinary services that have been quietly reliable for so long that the teams relying on them no longer think about what would happen without them.</p>



<p class="wp-block-paragraph">Once the startup dependencies are mapped, the next step is to look for loops: cases where one service needs another service to be running, but that second service eventually depends on the first. In a large platform, a surprising share of the control plane can be tied together by these loops. The result is that there is no valid order in which to start the services. Every possible starting point eventually leads back to a service that is still waiting. The loops themselves are usually mundane. DNS may depend on the inventory system that tracks what hardware exists, while the inventory system relies on DNS to resolve names. The artifact repository holding every installable package may depend on configuration management, which is itself installed from a package in that repository.</p>



<p class="wp-block-paragraph">Nobody designed any of this. Each dependency was a locally sensible decision made by a competent team, often years apart from the decisions that completed the loop. In a running region, the required services are already available, so the loop remains silent. The database is up when the alerting store starts, and nobody learns whether either could recover without the other. Starting a region from scratch is often the only event that reveals whether those services can start independently. Meta’s <a href="https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/" target="_blank" rel="noopener">outage in October 2021</a> shows how a large failure can expose dependencies that normal operation keeps hidden. When the backbone network connecting Meta’s data centers went down, its DNS servers withdrew their routes as designed to keep traffic away from unhealthy connections. That safeguard made DNS and many internal tools unreachable. With remote access unavailable as well, engineers had to go onsite, slowing recovery. A locally sensible safeguard had made system-wide recovery harder.</p>



<p class="wp-block-paragraph">Traffic-drain tests can reveal some of these dependencies. Meta’s <a href="https://www.usenix.org/conference/osdi18/presentation/veeraraghavan" target="_blank" rel="noopener">Maelstrom</a> encodes service dependencies and resource constraints to shift traffic safely from a failing data center to healthy ones, and its drain tests can uncover missing dependencies. But the receiving data centers are already running. A drain tests whether live infrastructure can absorb traffic; it does not test what an empty region needs in order to start. The drain graph is a useful input to the startup graph, not a substitute for it.</p>



<p class="wp-block-paragraph">Status reports for a new region can be misleading even when every team is reporting honestly. A service may be deployed, configured, monitored, and passing its health checks, yet still be blocked by a startup dependency. Readiness therefore has to propagate through the graph: a service is ready only when its own checks have passed and every service it needs at startup is also ready.</p>



<p class="wp-block-paragraph">Once the dependency graph exists, five practices turn it from a diagram into a launch plan. The first is finding what actually determines the launch date. The graph shows which services must wait for others, but the order alone does not reveal how long the work will take. Ten services that can start in parallel may finish before three services that must start one after another. Estimate the bring-up and validation time for every service. If several services remain tied together in a loop, treat them as a single planning block and include the time required to break the loop. The chain with the greatest total time becomes the critical path. Then count how many other services each foundational service can hold back, including dependencies several steps away. DNS, relational databases, secrets stores, and configuration management often rise to the top. Staff those teams early, because a week lost in one of them can become a week lost across the entire region.</p>



<p class="wp-block-paragraph">The second practice is to break the loops. One way to do that is to temporarily borrow a service from a region that is already running. Designate a small bootstrap tier, the minimum set of services required to deploy other services, and configure those bootstrap services to use working dependencies in an existing region until the local versions are ready. Suppose configuration management needs a package from the artifact repository, while the artifact repository needs configuration management before it can start. For the first installation, configuration management can fetch its package from another region. It can then bring up the local artifact repository and switch to using it. The bootstrap tier might also include identity, inventory, and package distribution. This shortcut works only when cross-region access is permitted and reliable enough. It also creates another cutover that must be planned, tested, and completed later.</p>



<p class="wp-block-paragraph">Borrowing from another region helps the new region get started, but it should not become permanent. Over time, each bootstrap service should be able to start without all of its usual dependencies. Suppose a service normally waits for the configuration system before it can start. Package the minimum settings it needs with the service itself. The service can start with those settings and fetch the latest configuration once the configuration system is running. Test this by turning off the dependency and starting the service from a clean state. If the service still cannot start, record the problem with an owner and a target date for fixing it.</p>



<p class="wp-block-paragraph">The third practice is to bring the region up in explicit tiers. A typical sequence begins with foundational configuration such as network ranges and routes, hardware inventory, identity configuration, access policies, service endpoints, and deployment settings. Bootstrap services such as DNS, certificate issuance, secrets, software package distribution, and configuration management follow. Next come the control planes that provision resources, schedule workloads, manage storage, and support service discovery. Stateful systems and applications come after the platforms they depend on. The exact tiers will vary by architecture, but the dependency graph should determine the sequence. Tier gates prevent visible application progress from hiding unfinished foundations.</p>



<p class="wp-block-paragraph">Stateful systems that depend on existing production data need special treatment because creating a cluster is often quick, while filling it with data is not. A storage system is ready only after the required data has arrived and been validated. Estimate that work using data volume, available bandwidth, validation time, and enough headroom for retries. Give each system a time box based on those measurements. If the estimate changes, require updated measurements that explain why. This keeps the plan honest without pretending that every delay is avoidable.</p>



<p class="wp-block-paragraph">The fourth practice is to make the bring-up repeatable, which does not mean automating every step. Automation helps only when it is maintained and tested. A script written for one region and left untouched for years may be more dangerous than a clear manual procedure. Automate steps that use the same tools as regular deployments or can be exercised frequently. For rare steps, maintain a runbook with validation checks, a named owner, and a schedule for testing it. Both automation and runbooks should clearly identify the required inputs, the evidence that a step succeeded, and how to continue after a partial failure. Google’s SRE book raises the same concern: <a href="https://sre.google/sre-book/automation-at-google/" target="_blank" rel="noopener">turn-up automation maintained separately from the systems it supports can become outdated</a>. At every manual step, ask whether another engineer could repeat it during a recovery without relying on the people who completed the original build. The goal is to leave behind a procedure that still works after the original team has moved on.</p>



<p class="wp-block-paragraph">The fifth practice is validation. It’s natural to test only the highest-traffic paths, but that approach misses structural problems. You also want the flows with the widest fan-out, the ones touching the most systems on the way through, even if few people use them, because those flows traverse more of the dependency graph. A high-volume request may prove that the region can handle load. A wide-reaching request can uncover an unready identity, storage, messaging, or data service. Meta’s <a href="https://www.usenix.org/conference/osdi16/technical-sessions/presentation/veeraraghavan" target="_blank" rel="noopener">Kraken</a> shows another useful validation method by shifting live user traffic into a data center while monitoring latency, errors, and system health. Live traffic also shows where capacity goes as caches warm, retries appear, and background jobs compete with user requests, a combination synthetic tests struggle to reproduce.</p>



<p class="wp-block-paragraph">When the traffic ramp begins, send a small percentage of representative production traffic to the new region, then increase it in stages. The size of each step should reflect the scale and risk of the platform, because even 1% can represent a substantial workload. This allows the entire application path to experience load together instead of testing each service in isolation. Hold at each step long enough for caches to warm, queues to stabilize, and relevant periodic jobs to run. Decide in advance which health signals allow the ramp to continue and which ones require it to stop. Also define and test how traffic will return to the existing region if health degrades. Confirm that the existing region has enough capacity and that data written in the new region will remain safe. Otherwise, the health signals may tell you something is wrong without giving you a reliable way to recover.</p>



<p class="wp-block-paragraph">These practices make no promise of a fast launch. They make the build understandable and leave behind a process the next region can use. The dependency map will begin aging as soon as systems change, but the ownership model, readiness rules, tier gates, repeatable procedures, and validation process can remain. The same tools are also useful for a disaster-recovery rebuild. Planning around dependencies will matter more as organizations rethink where their workloads should run, moving some systems from public clouds to private clouds, colocation facilities, or data centers they operate themselves. Each new environment brings another dependency graph that must be understood before it can carry production traffic.</p>



<p class="wp-block-paragraph">Finishing the facilities remains a genuine milestone. Power, cooling, networking, and hardware create the place where production can run. A working region emerges when its services can start in a valid order, the required data is ready, and the complete system has been tested under traffic. Finishing construction gives the organization a data center. Satisfying every required startup dependency in the graph turns it into an operational region.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Give Your Agent Its Own Computer</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/#respond</comments>
				<pubDate>Mon, 28 Sep 2026 13:54:33 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19829</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How sandboxed computers let agents run code, use browsers, and work independently of your machine]]></custom:subtitle>
		
				<description><![CDATA[In the most recent episode of Zero to Agent in 30 Minutes, AI engineer Sajal Sharma showed how to use a remote sandbox to let your agents install software, run commands, and control a browser without endangering your everyday machine. In effect, you’re giving your agents a computer of their own. Sajal demonstrated both approaches [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/4hwklWYt1AQ?si=DhwEX32iJaK2znj0" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">In the most recent episode of <em>Zero to Agent in 30 Minutes</em>, AI engineer Sajal Sharma showed how to use a remote sandbox to let your agents install software, run commands, and control a browser without endangering your everyday machine. In effect, you’re giving your agents a computer of their own.</p>



<p class="wp-block-paragraph">Sajal demonstrated both approaches with <a href="https://e2b.dev/" target="_blank" rel="noopener">E2B</a>. In one example, an agent downloaded a dataset, installed the packages it needed, analyzed the data, and produced a report inside the sandbox. In another, an agent opened a browser on a remote desktop and searched IKEA for furniture. The demos put both command-line and GUI-based computer use to work on a separate machine.</p>



<p class="wp-block-paragraph">If you want to follow along or try the same setup, Sajal shared the demo code in his <a href="https://github.com/sajal2692/zero-to-agent-own-computer" target="_blank" rel="noopener">GitHub repo</a>.</p>



<h2 class="wp-block-heading"><strong>How to give an agent its own computer, step-by-step</strong></h2>



<ol class="wp-block-list">
<li><strong>Choose how the agent will use the computer.</strong> Sajal demoed two ways to work with a remote machine. In the first, the agent used shell commands and files to install packages and process data. In the second, it worked through the graphical interface by reading screenshots and sending mouse and keyboard actions.</li>



<li><strong>Create an isolated sandbox.</strong> For the first demo, Sajal created an E2B sandbox before starting the agent loop. He configured LangChain Deep Agents to send command execution to that remote environment. The agent still did its reasoning locally, but package checks, installs, and data processing ran inside the sandbox.</li>



<li><strong>Transfer the files you need.</strong> Files on your local machine aren’t available in a remote sandbox unless you move them there. Sajal’s agent downloaded the dataset it needed inside the sandbox, created its report there, and then transferred the finished report back to the local machine.</li>



<li><strong>Map the agent’s actions to the remote desktop.</strong> In the GUI demo, Sajal used the OpenAI Agents SDK and built an E2B computer class that connected model actions to the desktop. He mapped screenshots, clicks, keystrokes, and scrolling to the corresponding E2B operations. An early version of the demo crashed because one of those actions wasn’t mapped, so he had to add the missing behavior before the demo could run without crashing.</li>



<li><strong>Tell the agent what environment it has.</strong> Sajal gave the agent basic operating instructions, including which browser was installed. That kept it from spending tokens figuring out how to use the machine. He also recommended giving agents their own task-specific credentials or secrets instead of reusing a person’s authentication profile.</li>
</ol>



<p class="wp-block-paragraph">A separate computer also helps when multiple agents need to work at the same time. Sajal used frontend development as an example. Two agents making UI changes might otherwise try to start development servers on the same port or inspect the wrong running instance. With a sandbox for each agent, they can start their own servers and test their own changes. Each run can also start on a fresh machine that gets deleted when the task is done.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">Next week, author and AI innovator Bruce Hopkins will show how to build your first agent with the Model Context Protocol (MCP). He’ll demonstrate how to take existing HTTP REST APIs and make them available through MCP so an agent can use those services as tools.</p>



<p class="wp-block-paragraph"><em>Follow along with Zero to Agent in 30 Minutes on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O’Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>From Raw Data to Graph-Native AI</title>
		<link>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/</link>
				<comments>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/#respond</comments>
				<pubDate>Mon, 28 Sep 2026 10:54:31 +0000</pubDate>
					<dc:creator><![CDATA[Ammar Mohanna]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19823</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/From-raw-data-to-graph-native-AI-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/From-raw-data-to-graph-native-AI-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Why graph modeling deserves as much attention as graph learning, GraphRAG, agents, and graph foundation models]]></custom:subtitle>
		
				<description><![CDATA[I’ve been thinking about a gap in the way we discuss graphs in AI. Most of the conversation starts after the graph already exists. We talk about graph neural networks, GraphRAG, graph agents, and graph foundation models. We spend much less time on the step that determines what all of them can do: turning raw [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I’ve been thinking about a gap in the way we discuss graphs in AI. Most of the conversation starts after the graph already exists. We talk about graph neural networks, GraphRAG, graph agents, and graph foundation models. We spend much less time on the step that determines what all of them can do: turning raw data into the right graph.</p>



<p class="wp-block-paragraph">Suppose an organization has customer records in tables, support conversations in documents, product telemetry in event streams, and incident histories in tickets. Before any of this can be used as a graph, someone has to decide what counts as an entity, what a relationship means, how time is represented, and which source to trust when records disagree. These choices shape every prediction, retrieval result, and agent decision that follows.</p>



<p class="wp-block-paragraph">This is what I mean by graph-native AI. It’s an approach that treats graph modeling as a first-class stage between raw data and the different systems that may use it. The graph isn’t simply an input to one model. It becomes a shared representation that can support queries, prediction, retrieval, agents, and, potentially, foundation models.<sup data-fn="939b81d2-2860-4055-8d46-9c019462af89" class="fn"><a href="#939b81d2-2860-4055-8d46-9c019462af89" id="939b81d2-2860-4055-8d46-9c019462af89-link">1</a></sup></p>



<h2 class="wp-block-heading">A graph is a model of the data</h2>



<p class="wp-block-paragraph">At its simplest, a graph consists of nodes and edges, which may also carry features. Real systems also need types, timestamps, provenance, confidence, permissions, and other context. Describing the finished graph tells us very little about how we should get there.</p>



<p class="wp-block-paragraph">Take customer-support data as a simple example. Should a company be represented as one node or as a set of legal entities that changes over time? Should an email be an edge between two people, a document node connected to its authors, or evidence for claims extracted from the text? Is a purchase an ongoing relationship or an event with a timestamp? If two records refer to the same customer, should we merge them, link them as possible matches, or keep them separate?</p>



<p class="wp-block-paragraph">There’s no single answer. An entity graph may be useful for identity and relationships. An event graph may be better for process and temporal analysis. An evidence graph may be better for retrieval and provenance. The right choice depends on what we want the system to do.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb661d3&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb661d3" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="800" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1600x800.png" alt="Figure 1. Raw data is modeled into one or more governed graph views, which different systems use in different ways." class="wp-image-19824" style="aspect-ratio:2" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1600x800.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-300x150.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-768x384.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1536x768.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 1. Raw data is modeled into one or more governed graph views, which different systems use in different ways.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">The step before graph learning</h2>



<p class="wp-block-paragraph">Graph representation learning normally starts with an existing adjacency structure. Embeddings learn vectors for nodes. Graph neural networks pass information across neighborhoods. Autoencoders reconstruct structure or attributes. Graph transformers combine local graph structure with longer-range attention. These methods ask: Given this graph, how should we learn from it?<sup data-fn="b246633d-6179-4150-8797-3544ffce2042" class="fn"><a href="#b246633d-6179-4150-8797-3544ffce2042" id="b246633d-6179-4150-8797-3544ffce2042-link">2</a></sup> Graph modeling asks the question before that: What should the nodes and edges be?</p>



<p class="wp-block-paragraph">Relational deep learning makes this distinction concrete. It maps rows in relational tables to nodes and primary-foreign-key links to edges, producing a temporal heterogeneous graph that a GNN can learn from. This can remove a good deal of manual joining and feature engineering. It also makes an assumption that deserves attention: A database schema designed for storage and transactions is also a useful graph for machine learning.<sup data-fn="36a752a5-5f9c-4123-8b2f-bfb978051cad" class="fn"><a href="#36a752a5-5f9c-4123-8b2f-bfb978051cad" id="36a752a5-5f9c-4123-8b2f-bfb978051cad-link">3</a></sup></p>



<p class="wp-block-paragraph">Recent work tests that assumption across 26 relational tasks. Graphs derived directly from database schemas sometimes suffered from information overload and semantic fragmentation. The authors improved performance by adapting the structure: removing distracting connections and adding dependencies that the original schema did not capture. I find this result important because it makes the graph itself part of the learning problem. Adding more edges is not always helpful. What matters is whether the structure supports the reasoning required by the task.<sup data-fn="d7eb81e1-26bf-4c11-8283-7296fd72da49" class="fn"><a href="#d7eb81e1-26bf-4c11-8283-7296fd72da49" id="d7eb81e1-26bf-4c11-8283-7296fd72da49-link">4</a></sup></p>



<p class="wp-block-paragraph">Graph construction therefore needs its own evaluation loop. We can generate a few plausible graph views, test them against the downstream task, inspect failures, and revise the structure. We should also keep provenance and uncertainty so that we know where an edge came from and how confident we are in it. The first graph we can extract is rarely the only graph worth considering.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb66cf1&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb66cf1" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="845" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1600x845.png" alt="Figure 2. The same source data can support entity-, event-, and evidence-centric views; each preserves different relationships." class="wp-image-19825" style="aspect-ratio:1.8982035928143712" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1600x845.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-300x158.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-767x405.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1536x811.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 2. The same source data can support entity-, event-, and evidence-centric views; each preserves different relationships.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">What the graph can support</h2>



<p class="wp-block-paragraph">Once the graph has been modeled and checked, several paths open. I don’t think this means every application should use one enormous universal graph. A more practical design is to build several governed views from the same modeled data while keeping identity, evidence, time, and access rules consistent underneath. Here are five ways to use graphs in an AI application.</p>



<ul class="wp-block-list">
<li><strong>Query and analytics:</strong> A graph database and ordinary graph algorithms may already be enough. Traversals, path queries, neighborhood aggregation, centrality, and community detection can reveal relationships that are awkward to see once the data is flattened into rows. This is also a useful baseline: If a deterministic query answers the question, there is no need to start with an LLM.</li>



<li><strong>Prediction and graph representation learning:</strong> Embeddings, GNNs, graph autoencoders, and graph transformers can make predictions at the node, edge, subgraph, or whole-graph level. Common examples include fraud detection, recommendation, molecular-property prediction, and link prediction. Here, the graph gives the model an explicit view of which entities are related and how information should move between them.</li>



<li><strong>Retrieval and GraphRAG:</strong> In this case, the graph is used as an index rather than as training data. Microsoft’s GraphRAG work builds an entity graph and community summaries to answer broad questions over a corpus that ordinary chunk retrieval may handle poorly.<sup data-fn="74153515-47b8-4b06-9476-2c59bff9e942" class="fn"><a href="#74153515-47b8-4b06-9476-2c59bff9e942" id="74153515-47b8-4b06-9476-2c59bff9e942-link">5</a></sup> Other approaches retrieve a task-specific subgraph and pass it to an LLM. This can be useful, but the extra graph pipeline has to improve the final result. Recent benchmarking found cases where GraphRAG underperformed vanilla RAG and argued that graph construction, retrieval, and generation should be evaluated together.<sup data-fn="53c72e11-abf6-4937-a1af-e19ae8e48471" class="fn"><a href="#53c72e11-abf6-4937-a1af-e19ae8e48471" id="53c72e11-abf6-4937-a1af-e19ae8e48471-link">6</a></sup></li>



<li><strong>Graph-augmented agents:</strong> Agents have memory, plans, tools, state transitions, and sometimes relationships with other agents. Some of this state is naturally relational. A graph can make it persistent, inspectable, and easier to update across turns. The emerging literature organizes these uses around planning, memory, tool use, and multi-agent coordination. The practical design question is which parts of an agent’s state benefit from being represented as a graph.<sup data-fn="5fb9336e-de1f-48bf-b989-aa16b2b15379" class="fn"><a href="#5fb9336e-de1f-48bf-b989-aa16b2b15379" id="5fb9336e-de1f-48bf-b989-aa16b2b15379-link">7</a></sup></li>



<li><strong>Graph foundation models:</strong> The goal here is to pretrain a model that can transfer across graph tasks, datasets, or domains. Current work combines graph backbones, self-supervised objectives, adaptation mechanisms, and sometimes language models. The difficult part is that graph semantics vary widely. Nodes and edges can represent very different things across molecular, transaction, and knowledge graphs. Transfer across these domains requires methods that bridge differences in structure and semantics, supported by suitable training data.<sup data-fn="f80bab3a-e1ac-4adb-a610-2737b4595f9b" class="fn"><a href="#f80bab3a-e1ac-4adb-a610-2737b4595f9b" id="f80bab3a-e1ac-4adb-a610-2737b4595f9b-link">8</a></sup></li>
</ul>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb676f8&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb676f8" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="889" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1600x889.png" alt="Figure 3. Five ways to use a modeled graph, from deterministic analysis to learned and generative systems." class="wp-image-19826" style="aspect-ratio:1.7982954545454546" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1600x889.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-300x167.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-767x426.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1536x854.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 3. Five ways to use a modeled graph, from deterministic analysis to learned and generative systems.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">What better graph modeling would look like</h2>



<p class="wp-block-paragraph">When I say that we need to crack graph modeling, I don’t mean a universal converter that produces one correct graph. The more useful goal is a system that can propose, test, and maintain several graph views of the same raw data.</p>



<p class="wp-block-paragraph">Such a system would need to do a few things well. It should identify possible entities and relationships in tables, text, events, images, and existing schemas. It should retain type, time, provenance, confidence, and permissions. It should be able to suggest alternatives instead of quietly committing to one structure. It should compare those alternatives using downstream quality, cost, stability, and human constraints. And it should keep the graph current as the source data and the organization’s understanding of it change.</p>



<p class="wp-block-paragraph">The downstream applications provide useful feedback. Prediction errors can point to missing relationships. Retrieval failures can expose poor granularity or disconnected evidence. Agent failures can show that important state transitions are absent. Weak transfer across datasets can reveal schemas that do not align. In this view, the graph is evaluated by what it helps the system do.</p>



<p class="wp-block-paragraph">Today, a team will often build a graph for one application: a fraud graph, a recommendation graph, or a knowledge graph for a chatbot. I think there is a larger opportunity. If the underlying graph modeling is done carefully, the same raw data can support several graph views and several applications without rebuilding identity, provenance, and governance each time.</p>



<p class="wp-block-paragraph">We already have strong methods for learning from graphs. The harder and less settled problem is deciding which graph we should build. If we make that step systematic and measurable, graph modeling can become a field in its own right and a shared foundation for analytics, prediction, retrieval, agents, and foundation models.</p>



<h3 class="wp-block-heading">Footnotes</h3>


<ol class="wp-block-footnotes"><li id="939b81d2-2860-4055-8d46-9c019462af89">Arijit Khan, Longxu Sun, Xin Huang, “<a href="https://arxiv.org/abs/2606.11560" target="_blank" rel="noopener">LLMs+Graphs: Toward Graph-Native, Synergistic AI Systems</a>,” arXiv, June 2026. <a href="#939b81d2-2860-4055-8d46-9c019462af89-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="b246633d-6179-4150-8797-3544ffce2042">William L. Hamilton, <em><a href="https://www.cs.mcgill.ca/~wlh/grl_book/" target="_blank" rel="noopener">Graph Representation Learning</a></em> (Morgan &amp; Claypool, 2020). <a href="#b246633d-6179-4150-8797-3544ffce2042-link" aria-label="Jump to footnote reference 2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="36a752a5-5f9c-4123-8b2f-bfb978051cad">Matthias Fey et al., “<a href="https://arxiv.org/abs/2312.04615" target="_blank" rel="noopener">Relational Deep Learning: Graph Representation Learning on Relational Databases</a>,” arXiv, submitted Dec 2023. <a href="#36a752a5-5f9c-4123-8b2f-bfb978051cad-link" aria-label="Jump to footnote reference 3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="d7eb81e1-26bf-4c11-8283-7296fd72da49">Yao Cheng and Siqiang Luo, “<a href="https://arxiv.org/abs/2606.08491" target="_blank" rel="noopener">What Makes a Desired Graph for Relational Deep Learning?</a>” arXiv, submitted June 2026. <a href="#d7eb81e1-26bf-4c11-8283-7296fd72da49-link" aria-label="Jump to footnote reference 4"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="74153515-47b8-4b06-9476-2c59bff9e942">Darren Edge et al., “<a href="https://arxiv.org/abs/2404.16130" target="_blank" rel="noopener">From Local to Global: A Graph RAG Approach to Query-Focused Summarization</a>,” arXiv, revised Feb. 19, 2025. <a href="#74153515-47b8-4b06-9476-2c59bff9e942-link" aria-label="Jump to footnote reference 5"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="53c72e11-abf6-4937-a1af-e19ae8e48471">Zhishang Xiang et al., “<a href="https://arxiv.org/abs/2506.05690" target="_blank" rel="noopener">When to Use Graphs in RAG</a>,” arXiv, revised Feb 2026. <a href="#53c72e11-abf6-4937-a1af-e19ae8e48471-link" aria-label="Jump to footnote reference 6"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="5fb9336e-de1f-48bf-b989-aa16b2b15379">Yixin Liu et al., “<a href="https://arxiv.org/abs/2507.21407" target="_blank" rel="noopener">Graph-Augmented Large Language Model Agents</a>,” arXiv, revised Aug 2025. <a href="#5fb9336e-de1f-48bf-b989-aa16b2b15379-link" aria-label="Jump to footnote reference 7"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="f80bab3a-e1ac-4adb-a610-2737b4595f9b">Zehong Wang et al., “<a href="https://arxiv.org/abs/2505.15116" target="_blank" rel="noopener">Graph Foundation Models: A Comprehensive Survey</a>,” arXiv, May 2025.<br> <a href="#f80bab3a-e1ac-4adb-a610-2737b4595f9b-link" aria-label="Jump to footnote reference 8"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>


<p class="wp-block-paragraph"></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Human Judgment Doesn’t Leave the Software Factory, It Relocates</title>
		<link>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/</link>
				<comments>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/#respond</comments>
				<pubDate>Fri, 25 Sep 2026 15:56:51 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Architecture]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19800</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Human-judgment-doesnt-leave-the-software-factory-it-relocates.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Human-judgment-doesnt-leave-the-software-factory-it-relocates-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[A field guide to building a software factory that still has an owner.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Elevate and is being reposted here with the author’s permission. A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you need a factory just yet. If [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <a href="https://addyo.substack.com/p/human-judgment-doesnt-leave-the-software" target="_blank" rel="noopener">Elevate</a> <em>and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph"><strong>A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you</strong> <strong><em>need</em></strong> <strong>a factory just yet.</strong></p>



<p class="wp-block-paragraph">If so:</p>



<ul class="wp-block-list">
<li>You’ll likely need humans in the loop upfront for deciding on product intent, system design (if you care) and your quality bar.</li>



<li>Do review code (lights-on factory) but be intentional with where it’s needed the most. I’ve found you want to watch out for where automated back-pressure breaks. Or where maintainability trade-offs need to be made.</li>



<li>Aim for quality checks to happen as early and continuously as possible. Not all of them have to, but this includes type systems, automated tests, mutation testing, security scanners and linting for architecture rules.</li>



<li>Number of checks ≠ quality. You’ll likely need to experiment with what checks give you the best signal to noise ratio. Be ready to tighten or relax your constraints deliberately.</li>
</ul>



<p class="wp-block-paragraph">You want to build your factory so some aspects of human taste get encoded in the environment, the agent gives you evidence of its work being right, and where a human still “owns” what ships to production.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb72b98&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb72b98" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="900" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1600x900.png" alt="Sonar screenshot" class="wp-image-19801" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1600x900.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1536x864.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19.png 1920w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><strong><em>Sponsored by <a href="https://fandf.co/4g7DP8r" target="_blank" rel="noopener">Sonar</a>:</em></strong> <em>Your agents write the code. Your gate decides if it ships. AI agents are writing more of my code, faster than ever—but fast isn’t the same as shippable. So I let a coding agent build an app, then put it through SonarQube. Every commit gets the same deterministic check: deep cross-file analysis, a clear map of where the risk actually lives, and a quality gate that holds every human and every agent to one bar. The screenshot? A PR that didn&#8217;t pass. That’s the gate doing its job.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading"><strong>Do you really need a software factory?</strong></h2>



<p class="wp-block-paragraph"><strong>In my experience, you can get surprisingly far with your stock coding harness!</strong> i.e., Claude Code or Codex, multiple sessions, good SPECs with verification baked in and constraints. You can even throw a batch of GitHub issues at them with implementation and human-involvement criteria, but it’s when this system needs to be <strong>repeatable and event-driven</strong> that a factory is helpful.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb733ab&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb733ab" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20.png" alt="Harness vs Factory" class="wp-image-19802" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">So I started off by saying <strong>a software factory is a repeatable loop around software work</strong>. We can actually look at a prompt that demonstrates a very small factory loop here:</p>



<pre class="wp-block-code"><code>Read GitHub issue #123 and the repository instructions before changing code.

Implement only the stated acceptance criteria. Do not modify authentication, billing, migrations, or existing test assertions. Work in a branch and keep the diff reviewable.

Run npm run lint, npm test, and npm run build. If a required check cannot run, stop and explain why. Open a draft pull request with the checks you ran, the remaining risks, and any decision a human still needs to make. Do not merge.</code></pre>



<p class="wp-block-paragraph">A goal can keep this moving until the checks pass and we can poll GitHub issues for any specific labels or review open pull requests each morning. Branch protection could enforce a merge boundary and the human can stay in the loop by choosing what becomes ready, reviewing and making the final merge calls etc.</p>



<p class="wp-block-paragraph"><strong>Add a software factory when you need an event-driven queue of work (e.g. Slack triggers, GitHub issues, Linear, a backlog) to run in an isolated cloud environment to handle triage, implementation and testing with some explicit human babysitting. Some end their loop with a monitor agent watching production and filing issues which triage again.</strong></p>



<p class="wp-block-paragraph">In my experience, the factory becomes useful when the hard part is making your different runs behave consistently, handing work off between agents and avoiding different sessions from claiming the same issue, preserving evidence and stopping production when human review is falling behind.</p>



<p class="wp-block-paragraph">What solves this might sound a little boring. For example, Warp mentions <a href="https://www.warp.dev/blog/how-to-build-a-cloud-software-factory-the-automatic-triage-skill" target="_blank" rel="noopener">triaging</a> every incoming issue into one of four states—ready-to-implement, ready-to-spec, needs-info, wait-to-implement—and the label is what fires the next agent.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb73cd4&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb73cd4" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21.png" alt="The label is the whole mechanism" class="wp-image-19803" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">This label does a few jobs in one go: it’s the queue, the lock, and since a session only picks up what’s marked ready, it’s where a human can park stuff without saying no permanently.</p>



<p class="wp-block-paragraph">Workflow wise, there are a few similarities and differences to just using Claude/Codex:</p>



<ul class="wp-block-list">
<li><strong>Steering:</strong> Agent needs course-correction, you can give input and redirect it.</li>



<li><strong>Notifications:</strong> How the factory says it’s blocked. This can be because a requirement was ambiguous, it started something risky or it needs human input (steering).</li>



<li><strong>Handoff:</strong> Move the task, its state, and context between the cloud factory/another agent/human reviewer. Good handoffs will keep track of what happened, what’s left to be done and why the handoff is needed.</li>
</ul>



<p class="wp-block-paragraph"><strong>In a good factory, the human isn’t limited to just reviewing and approving the final diff at the very end. They can shape the work early on, steer it during implementation, get it through a handoff, or stop it shipping to production.</strong></p>



<p class="wp-block-paragraph">Verification is where a responsible factory spends a lot of its time. We’ll cover this more later.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb746bc&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb746bc" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23.png" alt="Where human judgment goes" class="wp-image-19805" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">If you decide you do need a software factory, building it isn’t the only option. Standing up the infra to scale a factory can be a lot of work and you may want to consider buying verus building. <a href="https://factory.com/product/software-factory" target="_blank" rel="noopener">Factory</a>, <a href="https://www.warp.dev/blog/how-to-build-a-cloud-software-factory-the-automatic-triage-skill" target="_blank" rel="noopener">Warp</a> and <a href="https://www.humanlayer.dev/" target="_blank" rel="noopener">HumanLayer</a> are all working on this.</p>



<h2 class="wp-block-heading"><strong>What my day looks like now</strong></h2>



<p class="wp-block-paragraph">My day-to-day experience of software development has changed a lot over the past year. I’ve been talking about increasingly doing a lot of parallel work with agents, moving towards having a lights-on software factory. And a lot of people have been asking me, like, what do these things actually mean? What are you building? What are the kinds of projects that you’re using these things on? It’s a lot of this:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb74fb5&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb74fb5" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24.png" alt="Four ways into a run" class="wp-image-19806" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">So on a very average day, I have a simpler lights-on software factory. I can have tasks that are running in the cloud. Half of these tasks might be working on production client applications with smaller companies that I’m working with. They’re going to have real users. They’re going to have real authentication, payments, subscriptions, real beefy risks that you need to be careful with. You can’t just say, “Oh, agent, just go and do this stuff” without having tests and constraints and quality checks in place.</p>



<p class="wp-block-paragraph">I could work on my open source projects. I could be building out companion sites for my books. I could be working on tools. I could be building apps of my own. And these are all very, very different kinds of applications that I’m working on. And sometimes all they have in common is the tool that I’m using to work on them, right? Maybe I’m working on a migration. Others, I might be doing actual beefy feature work. And the blast radius of the work might also be very, very different.</p>



<p class="wp-block-paragraph">So as you begin to think about getting to a place where we’re increasingly doing a lot of parallel work, we’re trying to improve velocity, we’re trying to improve productivity, and we’re trying to improve autonomy, which means getting the system to a place where we trust it more, you do have to think about what are the places that absolutely require human code review, human input.</p>



<p class="wp-block-paragraph">And a lot of that’s going to be required up front, right? When you’re defining your specification, your requirements, what’s the design of the product going to look like? What’s the intent of the product going to look like? And then, how are you verifying that the agents have actually gotten the work done right? How are you verifying that they haven’t broken the existing system that’s been in place? How are you making sure that it’s meeting your quality bar?</p>



<p class="wp-block-paragraph">So generating code is not necessarily the part that you need to worry about the most. Given enough context, agents can write the implementation, run the tests, inspect failure, and revise code for us. <strong>We need to get to a place where we feel like there’s enough of human taste encoded in the environment that we can trust what</strong>’<strong>s being built, so that our human attention can be focused on the places where it’s needed most.</strong></p>



<p class="wp-block-paragraph">Now, there’s pushback from folks saying, “Hey, well, I don’t buy that you can just automate away a lot of this stuff.” It’s not to say that we’re automating away all of it, right? But given the volume of code that’s being generated, I don’t think that it’s realistic for humans to be reading all of it, especially when we’re not building rockets a lot of the time, right? We’re building UI, we’re building full stack applications.</p>



<p class="wp-block-paragraph">Our judgment, our taste is best focused on the places where it’s needed the most. Like, what are the riskiest parts of the systems? Where do we need to apply human taste? And that can be in the frontend. That can be in how the system works. It doesn’t have to be 100% of it.</p>



<h2 class="wp-block-heading"><strong>My cognitive bandwidth does not scale with the agents</strong></h2>



<p class="wp-block-paragraph">The reality is, yes, we can now fire up dozens, hundreds, thousands of agents in parallel, but your own cognitive bandwidth does not scale in the same way. This can feed into cognitive or comprehension debt which I’ve talked about before.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb759be&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb759be" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25.png" alt="Comprehension debt" class="wp-image-19808" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">If you remember back to just five, ten years ago, there was a lot of discussion in the engineering community about context switching and the cost of it. We would talk about how people hated when a colleague or someone would walk up to your desk when you were in the middle of a task. It would then take you so long to get back into your flow state because you had to catch back up in terms of like, where was I? What was I doing? Even if you had a little bit of residue there, it still took you time.</p>



<p class="wp-block-paragraph">We’re now context switching even more than we did before. On any given day, if I’m working outside of a software factory, I can be working on five or ten different projects with agents at a single time, or five or ten different features on a single project at a time. I can have five or ten different sessions, you can effectively say.</p>



<p class="wp-block-paragraph">That means that I have to be able to stay on top of at least a few of those. It is possible that I’m going to be able to increase how much autonomy I give some tasks if I have trust that I’ve defined the task well enough, I’ve defined the outcome, how it’s going to verify that it’s done well enough. But then there are going to be tasks where maybe I don’t necessarily feel that way and there’s more risk involved or more nuance. I’m going to have to pay attention.</p>



<p class="wp-block-paragraph">Consider optimizing the software factory for your reviewer. Given every one of those approaches still routes its output to one person’s attention, you should ask how much cheaper the factory is making the decisions you still have to make.</p>



<h3 class="wp-block-heading"><strong>A wrong-project mistake</strong></h3>



<p class="wp-block-paragraph">I remember when I’ve been working on multiple parallel projects with my agents, and there have been times when I’ve accidentally done things like, maybe I was working on a web app where I wanted to add in a dark mode, and so I had in my head, okay, well, this is what the shape of this needs to look like. But I accidentally went to the session for a different project, and I started putting in that same prompt.</p>



<p class="wp-block-paragraph">So I began implementing dark mode for something that absolutely didn’t need it. And so I can make that mistake. I don’t want my software factory making that kind of mistake.</p>



<p class="wp-block-paragraph">You need to think about this really in terms of a system. You are effectively trying to encode a software engineering culture, a team culture, into a system so that it has those same kinds of behaviors, so that it has ownership that belongs somewhere, so that someone is still on the hook for what happens, and you’re being very explicit about how you think about those things.</p>



<h2 class="wp-block-heading"><strong>When green is misleading</strong></h2>



<p class="wp-block-paragraph">Even in these systems, you want to be very careful, right? Many of us have seen that when you have asked AI to help you pass a test, like we’re talking about a programming test, a unit test, it can change the unit test to satisfy that condition, or it can change the logic of the code to pass that condition. That doesn’t mean that it’s actually followed your intent in order to align both the functional behavior and what the test was supposed to be testing, right?</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb76356&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb76356" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26.png" alt="When green is misleading" class="wp-image-19809" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26-300x169.png 300w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">Just because a software factory is showing that everything is green doesn’t mean that it’s actually green, especially at the start when you’re setting these things up. You need to pay a lot of attention to make sure that your checks, your verifications, all of those are shaped the right way. They’re doing what you expect them to be doing. You don’t want them to be misleading.</p>



<p class="wp-block-paragraph">You don’t want a situation where you had tests that said, hey, actually, I have gone and changed what authentication providers are supported. You asked me to add GitHub for example, as an authentication provider, but hey, my UI only had space for three, so I’ve gone and I’ve dropped one of the other ones. And hey, by the way, that happened to be one that your customers actually wanted. So you just need to be very explicit about how you want these systems to work.</p>



<p class="wp-block-paragraph">Btw, security is super important too, and if your factory reads untrusted input like a GitHub issue/Slack message it might be adversarial and include problems like supply chain attacks. So some explorations into software factories, like <a href="https://vercel.com/blog/building-a-software-factory-for-ai-sdk" target="_blank" rel="noopener">Vercel</a>, run their agents in isolated sandboxes holding just the secrets a task needs. That way a compromised run can’t reach what the job doesn’t need. Your defense ends up being layered.</p>



<h2 class="wp-block-heading"><strong>Which old projects deserve another life?</strong></h2>



<p class="wp-block-paragraph">I also think that a big part of how we work these days is deciding what should exist. If you remember back to many years ago before AI, there were so many abandoned software engineering projects, so many abandoned weekend projects, personal projects where they just wouldn’t launch because we didn’t have the time to finish them. We didn’t have the bandwidth to prioritize getting them out the door because they just weren’t that important to us or we couldn’t find the time.</p>



<p class="wp-block-paragraph">Now it’s fairly trivial for us to complete those projects, but the same human judgment question comes in. Do those projects deserve to exist? Should they be launched? Because you put them out into the world and even if it has just five users, maybe you have to maintain it. Maybe you have a quality bar now that you want to maintain.</p>



<p class="wp-block-paragraph">I know that I’ve had so many GitHub projects from over the years where now that I have an agent, the first thing I do is get the thing building. Because, of course, you clone it and now it doesn’t build because all the dependencies have changed. Half the things are out of date or have security vulnerabilities all over them, so you have to update that.</p>



<p class="wp-block-paragraph">Then you have to add tests if you didn’t have tests so that you know that behavior is at least going to be there if you’re upgrading the project in some way, or if you’re migrating it to a more modern language or framework or thing like that.</p>



<p class="wp-block-paragraph">Then you start to ask yourself, well, maybe, a silly example, but maybe I used Twitter Bootstrap back in the day for this, but now everybody is using Tailwind and shadcn, so I have to re-implement the UI. And what you’ll notice is that suddenly this is taking you more time, right? Yes, the agent can get a lot of this done quicker, but you’re now having to factor in product sense and taste and all of these things.</p>



<p class="wp-block-paragraph">You still question, well, who is this for? Does it have a market? Is it for myself? Is it for other people? If I’m putting it out into the world, is it still going to be as interesting given that now anybody can spin these things up as quickly?</p>



<p class="wp-block-paragraph">So I think that human question of do these things deserve to exist? How do we factor in our taste and judgment? I feel like those things continue to be extremely important. That’s where that scarce resource of human attention still really comes in. Back in the day, we only had a finite number of hours in the day. We had meetings. We had to budget in time for design and coding and so on.</p>



<p class="wp-block-paragraph">Now that we have agents to help us, I think that you have to really just be very explicit about where you’re spending your time and why.</p>



<h2 class="wp-block-heading"><strong>What happened when I built a sample one</strong></h2>



<p class="wp-block-paragraph">So I’m going to talk about the 82-minute factory run. People have been asking me for quite some time, you know, “How do I build a software factory?” Or, “I’m used to using Claude Code or Codex. How do I evolve my setup to using a software factory?”</p>



<p class="wp-block-paragraph">So the first thing that I’ve been saying is, “You may be fine. Your work may actually be totally fine without needing a factory.” But I did want to give people a reference setup that they can check out. So what I put together is a repository called <a href="https://github.com/addyosmani/factory" target="_blank" rel="noopener">Factory</a> that you can go and check out. I also put together a <a href="https://github.com/addyosmani/factory-demo/" target="_blank" rel="noopener">demo application</a> and <a href="https://github.com/addyosmani/factory/blob/main/ADVICE.md#:~:text=step%2Dby%2Dstep%20workshop" target="_blank" rel="noopener">workshop</a>.</p>



<p class="wp-block-paragraph">Now, for the last couple of years, my go-to demo application for a lot of things has been a movies app. I’m a big movies fan. I love watching movies. I watch movies all the time, and so I have a demo application, which really starts off as a very simple movies app. And what I want the factory to be able to do is go ahead and implement a number of features. There’s a few different features. I want a favorites feature. I want it to be able to maybe do search, and maybe also want a dark theme in there as well, those types of things.</p>



<p class="wp-block-paragraph">So I have my factory go and begin working with these things. You can check out the implementation. One of the benefits of it was actually catching real problems. These problems may not have been things that I would have caught if I had just asked it to do a one-shot implementation.</p>



<p class="wp-block-paragraph">Maybe around the 60-minute point, I was feeling like, “Wow, this is going unusually slow.” I asked my harness using the factory, “Why are things going slow?” It said, “This is actually totally fine. All the verifiers are still running.”</p>



<p class="wp-block-paragraph">You might have expected individual tasks to take 10 minutes, 15 minutes, 20 minutes, but they can take two to four times as long once you begin to include verification, retries, browser checks, human review, any of those extra delays.</p>



<p class="wp-block-paragraph">I do think that these can add up to better quality and better trust in the system. From a measurement perspective, you might look at metrics like cost per merged PR and code shelf life as comprehension debt metrics.</p>



<p class="wp-block-paragraph">You also need to think about what is useful delay versus factory overhead. The verifiers, in my case, caught some real problems. A little bit of the time was maybe sunk into producing evidence that I wanted. Some of it was overhead in the factory running. I didn’t really spend any time optimizing it, but a factory that just runs a lot of checks that you’re not finding valuable does not mean it’s a high quality one.</p>



<p class="wp-block-paragraph">You want to study how, for any repeated checks, are they irrelevant? Are they noisy? Are they actually making the system safer?</p>



<h2 class="wp-block-heading"><strong>A verification budget</strong></h2>



<p class="wp-block-paragraph">The way that I think about the budget for verification, this is basically what we’re talking about. We’re talking about a verification budget. I think about it in the same way as I’ve historically thought about performance budgets.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb770ec&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb770ec" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27.png" alt="A verification budget" class="wp-image-19810" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">There are going to be certain kinds of checks that you can run early on in your software development lifecycle, and there are going to be some things that are so heavy, but they offer so much value that you will want to run them later on. There are some kinds of fast checks, linting, for example, type checking. These are relatively fast checks that you can run early on.</p>



<p class="wp-block-paragraph">Our full suite of tests can be run closer to right before a draft PR is being put together or after that. That can include mutation testing, browser testing, security checks, anything like that.</p>



<p class="wp-block-paragraph">I think that you don’t necessarily want to replace these with just summaries. You want real tests, but you just need to make sure that you’re budgeting for them in the right places because you don’t want to slow down your development loop. I certainly never want to slow down my development loop. Having a fast iteration loop is important to me, but I also want to still have those checks and balances.</p>



<h2 class="wp-block-heading"><strong>When a run doesn’t ship</strong></h2>



<p class="wp-block-paragraph">A lot of what I’ve written above concerns the checks.</p>



<p class="wp-block-paragraph">In their software factory, Vercel <a href="https://vercel.com/blog/building-a-software-factory-for-ai-sdk" target="_blank" rel="noopener">marks</a> every agent run as “success”, “flawed”, “blocked” or “manual” and only “success” ships to production. The rest re-enter the system. I’ve been thinking about runs in similar terms.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb77a15&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb77a15" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.png" alt="When a run doesn't ship" class="wp-image-19811" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-300x169.png 300w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">“Flawed” here means the wrong thing was implemented or maybe it didn’t have full context, so that has to be fixed. Blocked means the environment may have been missing a credential so you have to provide it. Manual is a boundary the factory may not be allowed to cross it yet.</p>



<p class="wp-block-paragraph">Two of the three things here may have mechanical fixes and the last one is about trust.</p>



<p class="wp-block-paragraph">While this is great, what sorting doesn&#8217;t show you is cost. Back to my factory implementation with the TMDB app, the quick finder with no rejections took 7 minutes. Favorites, with two rejections and a human decision in the middle, took 56. Same factory. So I&#8217;d pair the taxonomy with per-stage timing, otherwise you know a run came back flawed without knowing what finding out cost you. The other thing I&#8217;d fix is the handoff at the boundary: my sample factory stopped issue the first issue and moved it to factory:needs-info, which was right, but I didn&#8217;t know where to put my answer. A manual run isn&#8217;t finished when the factory stops but when the human knows what to do next.</p>



<h2 class="wp-block-heading"><strong>Autonomy is not a single setting</strong></h2>



<p class="wp-block-paragraph">I wrote a couple of weeks ago an article about <a href="https://addyo.substack.com/p/agentic-autonomy-levels" target="_blank" rel="noopener">agentic autonomy</a> and how to think about autonomy because autonomy is not going to be a single setting for every single project.</p>



<p class="wp-block-paragraph">Verification buys you trust, and it buys the ability to grant more autonomy to your agents. So if, for example, I am working on a non-trivial change, but I have a number of checks in place, everything gets verified correctly, and maybe I’ve hand checked it myself. The next time I’m going to do a task like that in the same project, maybe I’ll feel comfortable giving the agent a little bit more autonomy.</p>



<p class="wp-block-paragraph">That’s the thing that you think about when you’re building these software factories. Your verification is going to change with risk. Your goal is the best signal to noise ratio. You don’t just want to have some large checklist that you’re running.</p>



<h2 class="wp-block-heading"><strong>The feature I had to relearn</strong></h2>



<p class="wp-block-paragraph">There was a feature that I’ve been putting off on a day when I’ve been using multiple sessions with Claude, and I was working on a few different projects at a time, a few different features at a time per project. And so Claude had implemented the feature that I was working on. It looked like the tests were passing. I hadn’t put a lot of thought into verification, but the tests passed, and so I thought it worked. I merged it.</p>



<p class="wp-block-paragraph">And so this was a favoriting feature. I thought that this was actually pretty good. I tried to check it out in the browser. It seemed like it was okay, but a couple of days later, I actually returned to the code because there were some tweaks that I thought I might make to this.</p>



<p class="wp-block-paragraph">I didn’t want to just ask my agent to make the changes, because it was just a subtle way that it worked. You tap on the icon, and it would not show the right effect on tap, and so I wanted to just tweak it. I wanted to understand how it worked so I could guide my agent correctly.</p>



<p class="wp-block-paragraph">I returned to the code, and I couldn’t explain to you how the feature worked. This repository was mine, right? I’d approved the change. I understood how a lot of it worked, a lot of the repo worked, but my understanding hadn’t kept up pace with all of the code that had been building up.</p>



<p class="wp-block-paragraph">What I failed to absorb was how this feature that had been added actually worked, how the UI worked, how the effect on it worked. I had to redo this feature and actually go step-by-step, “How does this work? How can I understand it?”</p>



<h2 class="wp-block-heading"><strong>What parallel work does to understanding</strong></h2>



<p class="wp-block-paragraph">When you’re doing parallel work, it amplifies this overall problem, and it gets even more amplified when you’re doing it in a software factory. When you’re doing five or 10 sessions, they create much more than just a review volume problem. They create several mental models that can end up going pretty cold while you’re working elsewhere.</p>



<p class="wp-block-paragraph">We’ve historically talked about the challenges with context switching, and as soon as chat compacts, you reject some approaches, you try out different things, you’re pairing with the agent, you’re going to have a difficult time remembering everything that happened in your session.</p>



<p class="wp-block-paragraph">You can scroll up, and as compaction has been happening, you’re not going to have everything there, and you’re not going to be able to store it all in your head. Code often preserves a decision that was made, but not why the decision was made.</p>



<p class="wp-block-paragraph">This is something that I think can be a useful learning for you, where it’s important, consider asking your agent to actually store information about its trajectory, or interesting lessons about how it approached a problem so that you can go back to it later.</p>



<p class="wp-block-paragraph">This can or can’t be something that you decide to commit to a repo. You can keep it local if you want, you can share it with a team if you want, but that can be something that can then be consulted later on. Rather than you relying on it maybe being in a session, or you maybe remembering about it later.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb78f41&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb78f41" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1200" height="600" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.png" alt="Sonarqube" class="wp-image-19812" style="aspect-ratio:2" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.png 1200w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-300x150.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-768x384.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>P.S. If agents are pushing to your main branch, they need the same quality bar you do. SonarQube gates every commit deterministically: one standard for humans and agents alike, on every PR. Sponsored by <a href="https://fandf.co/4g7DP8r" target="_blank" rel="noopener">Sonar.</a></em></figcaption></figure>
</div>


<h2 class="wp-block-heading"><strong>Ownership doesn’t disappear</strong></h2>



<p class="wp-block-paragraph">There is a broader principle underneath all of this.</p>



<p class="wp-block-paragraph">The percentage of code physically typed by humans may fall dramatically. I don’t think human ownership needs to fall with it.</p>



<ul class="wp-block-list">
<li>Someone still chooses the problem.</li>



<li>Someone still chooses the architecture.</li>



<li>Someone still sets the quality bar.</li>



<li>Someone decides which verification signals deserve trust.</li>



<li>Someone decides when the evidence is sufficient to ship.</li>
</ul>



<p class="wp-block-paragraph">And when the resulting system fails, “the agent wrote it” doesn’t cut it. This is why I don’t think the future of software engineering is best described as humans leaving the loop. Instead, <strong>human judgment is being relocated</strong>.</p>



<p class="wp-block-paragraph">We should remove people from the parts of the loop where machines can produce stronger, faster, more deterministic signals. At the same time, we should concentrate people around the places where context, taste, risk, and long-term ownership matter most.</p>



<p class="wp-block-paragraph">The best software factories will not be defined by how completely they eliminate human involvement.</p>



<p class="wp-block-paragraph">They will be defined by how intelligently they <strong>place</strong> it.</p>



<p class="wp-block-paragraph">Keep human judgment upstream on intent, system shape, and the quality bar. Review code where automated back-pressure becomes weak or the consequences become subjective. Push every deterministic signal as early and continuously into the loop as possible. Tighten and relax constraints deliberately as the system earns or loses trust.</p>



<p class="wp-block-paragraph"><strong>A human still has to own what code ultimately ships. Code good enough to ship still starts there.</strong></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: AI’s Safety Problem</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/#respond</comments>
				<pubDate>Fri, 25 Sep 2026 12:48:29 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19819</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Safety failures, sponsored agents, and the push to spread AI’s benefits more widely]]></custom:subtitle>
		
				<description><![CDATA[AI systems are gaining access to more tools and data, raising new questions about oversight and accountability. This Week in AI host Christina Stathopoulos spent this episode examining how those questions are playing out in model safety, government oversight, and even digital marketing. Safety needs more than safer models Anthropic’s latest threat report documented misuse [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/YPB_yfrnS-M?si=NAUWo13aGuxNwGXH" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">AI systems are gaining access to more tools and data, raising new questions about oversight and accountability. <em>This Week in AI</em> host Christina Stathopoulos spent this episode examining how those questions are playing out in model safety, government oversight, and even digital marketing.</p>



<h2 class="wp-block-heading"><strong>Safety needs more than safer models</strong></h2>



<p class="wp-block-paragraph">Anthropic’s latest threat report documented <a href="https://www.anthropic.com/threat-intelligence-report-september-2026" target="_blank" rel="noopener">misuse of Claude across seven categories</a>, including cyber operations, surveillance, influence campaigns, fraud, biological misuse, weapons development, and illicit model distillation. OpenAI reported on a different kind of AI risk, one that’s less about people weaponizing it and more about AI going off-script. They <a href="https://openai.com/index/model-misalignment-reporting-framework/" target="_blank" rel="noopener">examined model misalignment</a>, documenting several cases where models took actions outside the boundaries developers intended, including deception, unauthorized actions, and attempts to circumvent controls. After the OpenAI and Hugging Face controversy dominated headlines, other models have also reportedly reached systems outside their test environments, including Gemini, according to a <a href="https://www.bbc.com/news/articles/c607l0k72rlvo" target="_blank" rel="noopener">recent cybersecurity disclosure</a>.</p>



<p class="wp-block-paragraph">Christina cautioned against describing such incidents as models “escaping,” since that language can assign agency to the model while drawing attention away from how companies designed and secured the surrounding environment to begin with. As agents gain access to browsers, files, code, and external systems, the teams building and deploying them must rigorously test and secure those environments, with clear accountability when things go wrong.</p>



<p class="wp-block-paragraph">AI labs and governments are starting to wrestle with those requirements. To <a href="https://www.anthropic.com/institute/measuring-pace-of-ai-development#1-measuring-ai-led-ai-rd" target="_blank" rel="noopener">track the pace of AI development and maintain greater oversight</a>, Anthropic has proposed tracking how much AI contributes to AI R&amp;D, how closely organizations monitor agent actions, and how they allocate computing resources between capability and safety research. Anthropic and OpenAI have also proposed <a href="https://tech.yahoo.com/ai/claude/articles/anthropic-openai-want-embed-safety-210724177.html" target="_blank" rel="noopener">giving outside safety organizations greater access to their labs</a>, although Christina questioned their independence when frontier labs fund the work. Meanwhile, <a href="https://www.usatoday.com/story/news/politics/2026/09/16/ai-kill-switch-bill-fails-senate-john-kennedy/91797861007/" target="_blank" rel="noopener">a US Senate proposal for an emergency AI kill switch</a> failed to advance, while California ordered officials to develop proposals covering shutdown mechanisms and independent evaluation.</p>



<h2 class="wp-block-heading"><strong>AI is moving closer to the customer</strong></h2>



<p class="wp-block-paragraph">OpenAI is now testing Sponsored Agents, showing how <a href="https://openai.com/index/reimagining-advertising-with-ai/" target="_blank" rel="noopener">conversational AI could change digital advertising</a>. After clicking an ad, users can start a separate conversation with an AI agent representing the advertiser, ask questions, explore recommendations and then visit the company’s website when they are ready to take the next step.</p>



<p class="wp-block-paragraph">That approach could lead to more interactive advertising, but clear labeling will be essential so users always know when content is sponsored. Christina also raised the broader ethical concern of whether paid placements could influence the answers AI chatbots provide, blurring the line between independent guidance and commercial promotion.</p>



<h2 class="wp-block-heading"><strong>Access to AI also means access to expertise</strong></h2>



<p class="wp-block-paragraph"><a href="https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/goalkeepers-report-equitable-ai" target="_blank" rel="noopener">The Gates Foundation announced a $1 billion commitment</a> over two years to expand access to AI in healthcare, education, agriculture, and other areas. Its <a href="https://goalkeepers.gatesfoundation.org/report/2026-report/" target="_blank" rel="noopener">2026 Goalkeepers Report</a> argued that AI could help narrow existing gaps, but only if organizations intentionally make the technology and its benefits widely available.</p>



<p class="wp-block-paragraph">Christina highlighted examples from Kenya, Sierra Leone, India, and Rwanda. Health workers are using AI to improve diagnosis and treatment planning. Students are getting additional support from AI tutors, while small farmers can use personalized advice to improve harvests and make better decisions about market prices. These applications show practical roles for AI in places where demand for expertise exceeds the supply of teachers, clinicians, and other specialists.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">AI governance can’t stop at model evaluations. Organizations must also decide what AI systems can access, who reviews their actions, how commercial incentives affect their behavior, and how people continue developing the expertise needed to supervise them. The choices companies and governments make today will shape how useful AI becomes and how widely its benefits are shared.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Software Subscriptions Become Public Policy</title>
		<link>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/</link>
				<comments>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/#respond</comments>
				<pubDate>Thu, 24 Sep 2026 10:54:23 +0000</pubDate>
					<dc:creator><![CDATA[Steve Hayes]]></dc:creator>
						<category><![CDATA[Executive Briefing]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19795</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave? There was something wonderfully circular about hearing Dan Gookin tell me that, after he [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave?</em></p>



<p class="wp-block-paragraph">There was something wonderfully circular about hearing Dan Gookin tell me that, after he was first elected to public office, he bought a copy of <em>Robert’s Rules For Dummies</em>. Dan wrote <em>DOS For Dummies</em>, the 1991 book that launched the <em>For Dummies</em> series, and went on to write more than 180 technology books with over 12 million copies in print. I spent nearly three decades acquiring books in that program before joining O’Reilly this summer. Dan saw the news of my new role and reached out. We caught up on publishing, books, and AI. The part of the conversation that stuck with me had nothing to do with any of those topics. Dan is now the mayor of Coeur d’Alene, Idaho, and he’s begun thinking about the city’s software the same way he’s been thinking about his own.</p>



<p class="wp-block-paragraph">Dan was elected to the Coeur d’Alene City Council in 2011 and became mayor in 2025. He was quick to draw a distinction between the two jobs. “Council, you can be a little bit more extreme,” he told me. “It’s more rhetoric based. This is administrative. It’s management.” A council member can object to a budget line. A mayor has to make sure the systems that line funds keep working. For a city of nearly 58,000 people, that covers everything from the police network to the water bill.</p>



<h2 class="wp-block-heading"><strong>Cutting the leash</strong></h2>



<p class="wp-block-paragraph">Near the end of our conversation, Dan mentioned a project he’s been working on personally. He’s begun weaning himself off Microsoft and what he calls an increasingly subscription-based, cloud-based, AI-heavy environment, moving his computing to Linux and running his own servers. This move isn’t entirely ideological. There is money attached. Dan told me his Adobe subscription for software he uses to produce content costs about $800 a year. He believes he can replace most of it on Linux. What’s surprised him is how much he can replace. “It’s interesting to see how quickly it can be substituted,” he said, and how quickly he can “cut that chain.&nbsp;.&nbsp;.or cut that leash.”</p>



<p class="wp-block-paragraph">Dan sees a similar economic model when he gets to work. “Our software subscription just for our city, our size, is $700,000 a year and growing,” he told me. The example he kept returning to was the city’s financial software. “We used to own our finance software,” he explained. “Now the finance software is leased and it’s cloud-based.” Then he asked the question that ought to appear in a lot more software procurement meetings: “So we don’t even own our own data?”</p>



<p class="wp-block-paragraph">His concern isn’t literally that the city has forfeited legal ownership of its records. It’s about practical control. What happens if the vendor is hacked, or the city decides to switch? Can it get all its data back? “Do they just give you a binary dump or do they actually give you the data,” he asked. That question has stuck with me since we chatted.</p>



<h2 class="wp-block-heading"><strong>Reversibility is an architectural property</strong></h2>



<p class="wp-block-paragraph">Cloud contracts have a word for part of this issue. In cloud services agreements, <strong>reversibility</strong> refers to a customer’s ability to retrieve its data and unwind a vendor relationship. The United Nations Commission on International Trade Law (UNCITRAL) even includes it in its glossary of cloud contract terms. That idea belongs in municipal software evaluation too, not as an exit clause buried in a contract but as a standing column in the evaluation spreadsheet, next to features, price, and security.</p>



<p class="wp-block-paragraph">When you evaluate software, you compare features, price, implementation time, security, and, increasingly, AI capabilities. But what’s the exit cost? Can an organization export its data in a documented, useful format that another application can consume? How much institutional knowledge has quietly migrated from the organization to the vendor? And if the vendor raises prices sharply at renewal, gets acquired, or simply stops serving your needs, how long would it take to leave?</p>



<p class="wp-block-paragraph">These questions are properties of the system and subscription software has raised their stakes. When software came in a box, skipping an upgrade didn’t make the version you owned disappear. That world of proprietary formats and dominant platforms had plenty of lock-in. But possession still meant something.</p>



<p class="wp-block-paragraph">Dan and I talked about how alien that world now seems. Software once arrived with manuals. Today it may not even arrive. You authenticate to it. If you don’t know how to do something, you ask an AI instead of consulting a manual. That change highlights a step from possessing tools to maintaining permission to use them. For an individual, that might mean Photoshop is a monthly subscription now. For an organization, recurring access can become an architectural dependency. For a government, that dependency is borne by taxpayers.</p>



<h2 class="wp-block-heading"><strong>Should cities write their own software?</strong></h2>



<p class="wp-block-paragraph">Dan takes the argument one step further. “For $700,000 a year,” he said, “we could hire a couple of programmers just on contract and have them code our own stuff and then we own it again.” I’m not convinced that math works out. Two programmers likely can’t effectively reproduce a mature municipal financial system, endpoint protection, records management, and specialized public-safety applications, let alone the compliance work and vendor support that come with a modern city’s technology stack. AI could help accelerate the process, but liability and security issues surrounding AI-enabled development likely add more risk than a government is willing to accept in the name of software ownership.</p>



<p class="wp-block-paragraph">Building software also creates its own long-term bills, ones that don’t fit easily within most city budgets. Code must be maintained, security vulnerabilities patched, and staff and frameworks eventually replaced. An application written in-house can become every bit as difficult to escape as one bought from a vendor. On the flip side, the headaches of replacing a “good enough” proprietary system with a packaged one that fits most, but not all, of an organization’s needs can outweigh the cost of keeping the old system running. The familiar build-versus-buy analysis exists for good reasons.</p>



<p class="wp-block-paragraph">But Dan’s question still matters, even if his proposed fix isn’t right for every case. At what point does the cost of renting capability justify rebuilding some capability of your own? Perhaps more importantly, which capabilities should an organization insist on controlling, even if renting them is cheaper? There is no universal answer, including “the vendor handles it.”</p>



<h2 class="wp-block-heading"><strong>This isn’t an argument against the cloud</strong></h2>



<p class="wp-block-paragraph">It would be easy to turn Dan’s experiment into a familiar prescription. Move to Linux. Embrace open source. Bring everything back on premises. Escape the cloud. That’s too simple. Cloud and SaaS products solve real problems, shifting maintenance to specialists and giving a city of 58,000 residents access to capabilities it could never economically build or maintain for itself.</p>



<p class="wp-block-paragraph">The key question doesn’t boil down to cloud versus on premises, or proprietary versus open source. It becomes a decision about whether you accept dependency you’ve consciously chosen or dependency you’ve acquired by default. An organization may rationally decide to rent a critical service indefinitely. But it should know where the data lives, how it comes back, what replacing the service would require, and which internal skills have atrophied because the vendor now supplies them. Revisiting those answers periodically, rather than treating last year’s renewal as the rationale for next year’s, is a step that’s easy to skip when nobody’s asking the questions in the first place.</p>



<h2 class="wp-block-heading"><strong>From hobbyists to city hall</strong></h2>



<p class="wp-block-paragraph">Dan suspects the search for alternatives will happen outside big organizations first. He compared it to the early personal computer movement. Hobbyists and enthusiasts experiment first, long before organizations decide the ideas are practical. His hunch is executives will eventually look at how much of their budgets go to recurring subscriptions and ask a simpler question: How much are programmers? Again, I don’t think the answer will be “hire programmers and cancel SaaS.” But more organizations will ask the question behind the question. What are they paying for convenience? What are they paying for capability? And what are they paying because leaving has become too difficult?</p>



<p class="wp-block-paragraph">Software can grow to define an organization rather than serve it, especially when the cost of paying for or maintaining a tool outgrows the tool’s value. That shift becomes a problem when systems are so deeply embedded that replacing them feels impossible or when years of subscriptions leave an organization unable to perform a basic function on its own.</p>



<p class="wp-block-paragraph">Dan’s new job has given that shift a different scale. He told me the biggest adjustment from council member to mayor was realizing that the job is administration and management. He’s less interested in ceremonial appearances than in answering email, returning calls, setting meetings, and, in his words, getting stuff done. Software is part of getting stuff done. So is knowing when to buy it and when to build it. The latest addition to that list may be knowing that the tool with the most impressive new capability is not as valuable as the one you can still leave.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Will TypeSafe’s Jev Change How We Build AI Applications?</title>
		<link>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/</link>
				<comments>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 16:28:07 +0000</pubDate>
					<dc:creator><![CDATA[Laurie Voss]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19782</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Arize’s blog and is being reposted here with the author’s permission. This week the AI community was in uproar about Jev from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on <a href="https://arize.com/blog/typesafe-jev-llm-judge/" target="_blank" rel="noopener">Arize’s blog</a> and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">This week the AI community was in uproar about <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener">Jev</a> from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the fuss is simple. It’s radically faster and cheaper than using an LLM to perform the same task (up to 200x faster and 400x cheaper if TypeSafe’s numbers are to be trusted). In one small independent test, a general-purpose model spent about 910 output tokens reasoning its way to each yes-or-no answer while Jev spent 85, and it doesn’t even bill for them.</p>



<p class="wp-block-paragraph">That’s potentially a really big deal. An enormous share of LLM-powered components in AI applications today are being asked to make decisions: pass or fail, route A or route B, which of five labels to pick. In particular, that’s something that LLM-as-a-judge evaluations are doing all the time, so it really made our ears perk up at Arize AI. This post is about how we got here, what this new kind of model buys you, what you lose, and what choices you should be making about your application’s architecture as a result.</p>



<h2 class="wp-block-heading">TypeSafe shipped a model that can’t write, only decide</h2>



<p class="wp-block-paragraph">Here’s how Jev works. You send it some data that represents a state (a support ticket or an agent trace or a JSON blob) plus a list of typed questions (Choose one of these options; Score this on a scale; Is this statement true?). It doesn’t generate a token stream. It returns <a href="https://docs.typesafe.ai/concepts/system-one.md" target="_blank" rel="noopener">typed answers with probability distributions</a> in a single parallel pass, in 70 ms to 500 ms, at $0.042 per million input tokens. No free-form text comes back.</p>



<p class="wp-block-paragraph">The training method used to create Jev is what TypeSafe calls Reinforcement Learning for Calibrated Decisions or RLCD, described in their <a href="https://docs.typesafe.ai/introduction/machine-learning-primer.md" target="_blank" rel="noopener">primer</a> as training the model so that a higher stated probability means a higher chance the answer is right. (You’d think that’s always what a higher stated probability should mean, but read on for surprising facts about how LLMs work.)</p>



<p class="wp-block-paragraph">TypeSafe also claims Jev “can’t hallucinate,” but that really feels like an overreach. Jev can’t return an answer outside the schema you gave it. Within that schema, it could still be giving the wrong answer, although its probability score should give you a clue if it’s not confident.</p>



<p class="wp-block-paragraph">And the whole thing is incredibly fast and incredibly cheap: 40x to 200x faster and 40x to 400x cheaper, depending on the task, are TypeSafe’s numbers from TypeSafe’s evals. Of course, we know better than to take a vendor’s word for these things, so Arize will be running our own benchmarks just as soon as we can. But other people have already started doing that.</p>



<h2 class="wp-block-heading">On the early data, Jev is mid-tier intelligence at a two-orders-of-magnitude discount</h2>



<p class="wp-block-paragraph">TypeSafe’s <a href="https://evals.typesafe.ai/" target="_blank" rel="noopener">published evals</a> run four decision workflows, one of which is reviewing a finished agent trace to decide whether a human needs to look at it. Averaged across the four workflows, Jev lands at 68% accuracy at $0.0004 and 0.4 seconds per case. GPT-5.6 Terra is at 68% for $0.03 and 10 seconds. Opus 5 is at 73% for $0.18 and 38 seconds. That’s five points behind Opus 5, but on the other hand it’s 440x cheaper. That’s a very interesting cost-benefit trade-off, and such a radical one that it may change how we architect our applications.</p>



<p class="wp-block-paragraph">The independent data so far is small but it points the same way. <a href="https://every.to/" target="_blank" rel="noopener">Every</a>’s head of evals ran <a href="https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds" target="_blank" rel="noopener">777 judgments in under 0.7 seconds</a> for about a quarter of a cent. A UK events site, NearHere, tested listing moderation and got <a href="https://nearhere.events/blog/typesafe-jev-mistral-gemini-event-validation" target="_blank" rel="noopener">96% from Jev against 86% from Gemini Flash-Lite</a>, 58x cheaper per decision. That’s where the 910-versus-85 token count I mentioned earlier came from. And a developer ran Jev zero-shot over <a href="https://github.com/bitnovus/jev-spam-eval" target="_blank" rel="noopener">18,514 spam emails</a>, getting a result that was a statistical tie versus a classifier trained on the labels.</p>



<p class="wp-block-paragraph">These are small samples and early data but hey, the thing was released a few days ago.</p>



<h2 class="wp-block-heading">We used LLM judges because nothing else worked without labels</h2>



<p class="wp-block-paragraph">Why are we using LLMs to make decisions in the first place? The reason is simple: They are able to do it without huge, expensive training sets, which is what most ML solutions prior to LLMs required. Here’s the options on the table now:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac018eb81147&quot;}" data-wp-interactive="core/image" data-wp-key="6ac018eb81147" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="2034" height="740" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png" alt="ML solutions table" class="wp-image-19783" style="aspect-ratio:2.748898678414097" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png 2034w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-300x109.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-767x279.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1600x582.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1536x559.png 1536w" sizes="auto, (max-width: 2034px) 100vw, 2034px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The first two rows are the old-school options, which need a training dataset: hundreds to thousands of labeled examples before you get a single prediction, and then a training run, and then someone to maintain it. Nobody building a first version of an AI product has that data or that kind of time. The LLM as a judge, on the other hand, just asks for a paragraph-long prompt. The decision was easy.</p>



<p class="wp-block-paragraph">But the results weren’t without trade-offs. On a <a href="https://www.latent.space/p/benchmarks-201" target="_blank" rel="noopener"><em>Latent Space</em> episode in July 2024</a>, Clémentine Fourrier of Hugging Face laid out what LLM judges are bad at: They prefer their own model family, and they can’t score on a continuous scale. Asked what benchmark she wished existed, Fourrier said, “Nobody’s evaluating model calibration at the moment.” With the release of Jev, the need for that benchmark is even greater, because real progress seems to have been made.</p>



<h2 class="wp-block-heading">Jev gives you zero-shot probabilities without training and without a generator</h2>



<p class="wp-block-paragraph">Jev takes the same plain-English criteria you’d put in a judge prompt, needs no labels, and returns a probability. Zero-shot and autoregressive text generation are no longer tied together. We were paying for the second to get the first, and it turns out you don’t have to.</p>



<p class="wp-block-paragraph">The spam evaluation I mentioned earlier is an impressive demonstration of how attractive this new offering is. With zero labeled examples and a simply well-written definition of spam, Jev hit 98.3% accuracy. A TF-IDF logistic regression trained on about 14,800 labeled emails hit 98.4%. The two disagreed on 466 emails and split them almost evenly, with no statistically meaningful difference between the two. So a classifier from 2003, trained on a dataset, only ties a decision model trained on nothing. It’s early data that’s yet to be reproduced, but if it holds up, that’s an amazing new capability unlocked.</p>



<p class="wp-block-paragraph">But there are still some trade-offs you’re making.</p>



<h2 class="wp-block-heading">A radically cheaper decision loses you some things</h2>



<p class="wp-block-paragraph">The biggest loss is the explanation. TypeSafe’s docs say plainly that System One models don’t generate explanations of their reasoning, and NearHere’s test noted the same thing: a category and probabilities came back, nothing else.</p>



<p class="wp-block-paragraph">Depending on your use case, that could matter a lot. LLM judge explanations are an incredibly valuable tool that tells you not just what was wrong, but why. That provides real signal that can be fed en masse back to a coding agent and used to automatically improve your software. Jev on the other hand just gives you a probability, which leaves you with much less directional signal of how to improve.</p>



<p class="wp-block-paragraph">Of course, at these prices, you can do both: run Jev on every single trace for broad, comparably accurate measurement and monitoring, and then take samples of failures and rerun them through an LLM judge to get your directional signal. That involves changing how you work, which is why I say that this may require rearchitecting your systems.</p>



<h2 class="wp-block-heading">To automate a decision, you need to know which 5% to hand to a human</h2>



<p class="wp-block-paragraph">TypeSafe’s launch post makes the point that a model that’s right 95% of the time but can’t tell you when it’s in the other 5% can’t automate anything, because a person still has to review all of it. That’s an important point because it highlights a problem with LLM judges.</p>



<p class="wp-block-paragraph">We evaluate LLM judges by accuracy against a gold set. Accuracy tells you how many errors to expect, but not where they will be. If your LLM application is making decisions for you, it feeds three things: a threshold that decides when to act, an escalation path that decides when to ask a human, and a drift monitor that decides when the world has changed under it. All three need a probability score, but LLM judges don’t provide reliable probabilities. A <a href="https://arxiv.org/abs/2508.06225" target="_blank" rel="noopener">2025 study of 14 models on JudgeBench</a> found judges clustering their predictions at 90% to 100% confidence while landing well below that in accuracy, and argued for exactly this shift from accuracy-centric to confidence-driven evaluation.</p>



<p class="wp-block-paragraph">The same small spam evaluation test shows what a usable probability looks like. Of the emails Jev scored under 0.1, 0.1% were spam. Of those scored 0.9 or above, 99.9% were. In the 0.5 to 0.6 band, only 38% were. That curve tells you where to set your threshold and how much human review you’re buying: Sending the 4.6% of emails scored between 0.3 and 0.7 to a person left the rest at 99.5% accuracy. Your overconfident LLM judge can’t get you there.</p>



<p class="wp-block-paragraph">Another metric to consider is tokens per decision. A component that spends thousands of output tokens to emit one of five labels is telling you it’s the wrong tool for the job. UkisAI’s <a href="https://huggingface.co/ukisai/Swift-Qwen3.8-27b" target="_blank" rel="noopener">Swift-Qwen3.8-27B</a> cut 58% of its reasoning tokens on GPQA-Diamond and lost 0.1 points of accuracy. A lot of these tokens aren’t making a critical difference to accuracy.</p>



<p class="wp-block-paragraph">In <a href="https://arize.com/products/ax/?utm_source=lvoss&amp;utm_medium=linkedin&amp;utm_campaign=devrel&amp;utm_content=Have%20we%20been%20using%20the%20wrong%20kind%20of%20model%20to%20make%20decisions%3F" target="_blank" rel="noopener">Arize AX</a>, eval labels, the judge’s explanation, and the token count and cost of the judge call sit on the same trace, so tokens per decision is a column you can sort by rather than a number you have to figure out.</p>



<h2 class="wp-block-heading">Cheap decisions change the math of how you build and measure AI applications</h2>



<p class="wp-block-paragraph">As I mentioned earlier, at $0.0004 and 0.4 seconds a decision, you can stop sampling. You can check every output, every tool call, and every agent step as it happens. For some use cases that’s a total game changer.</p>



<p class="wp-block-paragraph">But it might require that you rearchitect how your application works to make the most of it. Take the work and decompose into many small typed questions; only call the expensive LLM generator when text actually needs to be written. That’s a stack where the decision layer is something you can version, measure, and swap independently of the model that writes the words, and it’s the first time decision-making has been cheap and fast enough to make that practical without requiring training data.</p>



<p class="wp-block-paragraph">So go count how many of your LLM calls end in one of five labels. Then work out what you’d check, and how often, if each of those calls cost a fraction of a cent and came back with a probability you could trust.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>MCP Is Not Just Another API Standard</title>
		<link>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/</link>
				<comments>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 10:17:25 +0000</pubDate>
					<dc:creator><![CDATA[Balaji Venkatasubramaniyar]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Operations]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19772</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What Model Context Protocol actually changes for agentic systems, and what it doesn’t]]></custom:subtitle>
		
				<description><![CDATA[Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water through pipes you already designed. MCP changes who’s holding the wrench. Once you’ve felt that shift in a real production system, the “just another API standard” framing stops making sense.</p>



<p class="wp-block-paragraph">This piece is the long version of that argument. It walks through what MCP’s primitives actually are and why they’re the right primitives, where the standard genuinely collapses integration work that used to be duplicated per framework, where the abstraction leaks in ways that only show up once you’re past the demo, and what the protocol’s own 2026 evolution tells you about where the real pain has been. Nearly everything worth knowing about building on MCP falls out of understanding these pieces and how they interact.</p>



<h2 class="wp-block-heading"><strong>What MCP actually standardizes</strong></h2>



<p class="wp-block-paragraph">Strip away the framing and MCP is a JSON-RPC-based protocol that lets a client (the thing driving an LLM) talk to a server that exposes capabilities, over a small, fixed set of primitives:</p>



<p class="wp-block-paragraph"><strong>Tools</strong> are callable functions. Each one has a name, a description, and a JSON Schema describing its inputs. This is the primitive most people mean when they say “MCP,” and it’s the one doing the heavy lifting in most production deployments: “look up an order,” “run a query,” “create a ticket,” etc.</p>



<p class="wp-block-paragraph"><strong>Resources</strong> are readable context, addressed by URI, that a client can pull in without the model having to call a function to get it: a file, a record, or a document, for instance. Think of this as the read side of the interface, separate from the “do something” side that tools represent.</p>



<p class="wp-block-paragraph"><strong>Prompts</strong> are reusable templates a server offers to the client, so common workflows don’t have to be respecified from scratch every time.</p>



<p class="wp-block-paragraph">On top of those three, the spec defines capabilities that flow the other direction, from server back to client: Sampling lets a server ask the client’s model to generate text on its behalf; elicitation, added in the 2025-06-18 revision, lets a server pause and ask the human for more input mid-task; and roots let a server learn which directories or URIs it’s actually allowed to touch.</p>



<p class="wp-block-paragraph">None of these primitives are individually novel. What’s novel is that they’re the same five primitives regardless of which model, which framework, or which vendor is on the client side. That’s the entire value proposition in one sentence, and it’s also the source of everything that goes right and everything that goes wrong when you build on top of it.</p>



<h2 class="wp-block-heading"><strong>Where the old model breaks down</strong></h2>



<p class="wp-block-paragraph">Before MCP, wiring an LLM into an enterprise system meant writing tool-calling code for that specific model, that specific framework, that specific integration. Every agent framework had its own function-calling convention: its own way of describing a schema, its own way of parsing a model’s intent to call something, and its own error-handling contract. Every system you wanted to expose needed its own adapter written to whichever dialect that framework spoke. Add a second framework to your stack and you don’t get twice the work. You get a second, incompatible copy of the same logic, maintained by whoever drew the short straw.</p>



<p class="wp-block-paragraph">MCP replaces that with one contract, written once, usable by any compliant client regardless of which model sits behind it. That’s the part every MCP explainer gets right, and it’s a real, measurable win. I’ve watched it collapse from a maintenance burden that used to scale with the number of frameworks a team happened to be supporting that quarter down to something that scales with the number of systems, full stop.</p>



<p class="wp-block-paragraph">But the more consequential change is where the integration decision gets made. A traditional API integration is an agreement two systems make in advance. You negotiate a contract: endpoints, payloads, auth, versioning, and both sides build to it, because a project plan said this integration should exist. The plan predates the code.</p>



<p class="wp-block-paragraph">An MCP server doesn’t get that luxury. It has no idea which agent will call it, in what sequence, alongside which other servers, in service of what goal a human typed into a chat box 30 seconds ago. The plan doesn’t exist as a concrete thing until the agent composes one, at runtime, out of whatever tools happen to be available to it. That’s not a stylistic difference from the old model. It’s a different category of integration problem, because the party doing the composing isn’t your code anymore. It’s a model, reasoning over natural-language descriptions you wrote weeks or months earlier, with no idea what context it would eventually be reasoning inside of.</p>



<h2 class="wp-block-heading"><strong>The description is the interface now</strong></h2>



<p class="wp-block-paragraph">A tool’s JSON Schema tells the agent what parameters it takes and what shape they need to be. That part is mechanical, and MCP handles it well. The tool’s name and description tell the agent when to use it at all, and whether to prefer it over some other tool that does something adjacent. Those are two different jobs, and only one of them is solved by a well-formed schema.</p>



<p class="wp-block-paragraph">Picture two versions of the same tool description. The first is technically correct and nothing more:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current status of a record given its ID."
}</code></pre>



<p class="wp-block-paragraph">An agent reading that has no idea when this is the right tool versus three other tools that also return some kind of status, no idea what “record” means in this system, and no idea whether IDs are case-sensitive, numeric, or prefixed. The second version spells out the domain the tool operates in, gives the ID format explicitly, states what the returned status values mean, and flags the one adjacent tool this one is commonly confused with and why they’re different:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current fulfillment status for an order record. 
IDs are numeric order numbers (e.g. 48213), not SKUs or customer IDs. Status 
values are one of: pending, processing, shipped, delivered, cancelled. Use this 
instead of get_shipment_status, which returns carrier tracking events rather 
than the order's internal state."
}
</code></pre>



<p class="wp-block-paragraph">That’s a longer description, and it will feel like overexplaining to the engineer writing it, because the engineer already knows all of this. The agent doesn’t. It’s encountering the tool for the first time, with a handful of tokens to decide whether it’s the right call, and no colleague to ask.</p>



<p class="wp-block-paragraph">I’ve watched teams ship a technically correct MCP server that agents used badly, or avoided entirely in favor of a worse but better-described alternative, purely because of this gap. The failure mode isn’t a stack trace. It’s an agent confidently calling the wrong tool, or the right tool with an assumption baked in that happened to be wrong for this case, and nobody notices until the output looks slightly off downstream. Writing tool descriptions well is closer to technical writing and product design than it is to backend engineering, and it’s not a skill most integration teams (mine included, early on) walked in the door with.</p>



<h2 class="wp-block-heading"><strong>Composition is emergent, and that cuts both ways</strong></h2>



<p class="wp-block-paragraph">The entire appeal of MCP is that an agent can combine tools from servers that never agreed to work together, in combinations their respective authors never planned for. That’s also the risk, and it’s structural, not a bug you fix with better testing.</p>



<p class="wp-block-paragraph">In a traditional integration, the sequencing logic (call A, then check its result, then decide whether to call B or C) lives in a script that a human wrote and a reviewer read. You can unit test it. In an MCP-based agent, that same sequencing logic lives in the model’s runtime reasoning, generated fresh for each task based on the goal it was given and whatever tools happen to be available in that session. You can’t unit test a decision that doesn’t exist until the moment it’s made.</p>



<p class="wp-block-paragraph">A tool that behaves correctly in isolation, with the exact inputs its author tested against, can still produce a bad outcome the first time an agent calls it third instead of first in a chain, or passes it a value that came from a different server’s output rather than a human’s direct input. This is qualitatively different from a normal integration bug, because it doesn’t show up in code review, and it won’t show up in testing unless your test suite happens to exercise that specific, unplanned chain of calls. It shows up in production, once, when a particular combination finally occurs. That’s exactly the kind of failure mode that’s cheap to dismiss as an edge case until it happens to the wrong customer.</p>



<h2 class="wp-block-heading"><strong>The protocol is catching up to its own success, in specific and telling ways</strong></h2>



<p class="wp-block-paragraph">To be fair to MCP, it isn’t standing still, and the shape of its evolution tells you a lot about where the real production pain has been. The July 28, 2026 specification is the largest revision since the protocol’s November 2024 launch, and every major change in it traces back to something that broke, or nearly broke, at scale.</p>



<p class="wp-block-paragraph">The protocol core is now stateless. The original design tracked sessions with an MCP-Session-Id header, workable for a single server instance but painful the moment you’re running behind a normal horizontally scaled fleet and discover that “any instance can answer any request” and “sticky session state” don’t coexist. Removing protocol-level sessions means the same request can be served by any instance behind ordinary load-balancing infrastructure, which sounds unglamorous right up until you’re the one who has to explain in an incident review why a routine deploy dropped a chunk of in-flight sessions.</p>



<p class="wp-block-paragraph">Tasks formalize long-running work. A lot of real enterprise work (document processing, multistep approvals, anything involving a human in the loop) doesn’t complete inside a single request/response cycle. Before this extension existed, teams hand-rolled this with polling loops and webhook callbacks, each implementation slightly different, each one a source of its own edge cases. Tasks turn that into a first-class protocol concept.</p>



<p class="wp-block-paragraph">MCP Apps let a server return interactive UI, not just structured data. That matters the moment a “tool” is something a human needs to actually look at and approve before it fires, which in any environment with real consequences attached is often.</p>



<p class="wp-block-paragraph">Authorization was hardened to align with OAuth 2.1 and OpenID Connect. This one isn’t novel so much as overdue, and the gap it closes was a real one; see the governance section below.</p>



<p class="wp-block-paragraph">A formal deprecation policy now governs the legacy HTTP+SSE transport, with a 12-month offramp. That’s the kind of unglamorous governance maturity a protocol only earns after it’s been run in production long enough for someone to need it.</p>



<p class="wp-block-paragraph">None of this is exciting reading. All of it is the sound of a two-year-old protocol absorbing genuine operational scar tissue, which is a far better signal about its trajectory than raw adoption numbers. And the adoption numbers are themselves striking: The official registry tracks close to 10,000 distinct servers, Tier 1 SDK downloads run into the tens of millions monthly, and both the TypeScript and Python SDKs have individually crossed a billion total downloads. Competitors of the protocol’s original author adopted it within months. That combination, real scale plus a spec that keeps changing in response to real production failure modes, is a much stronger signal of durability than either fact alone.</p>



<h2 class="wp-block-heading"><strong>Governance hasn’t caught up</strong></h2>



<p class="wp-block-paragraph">Here’s what I’d want any team to weigh before connecting MCP to anything that matters. Independent security research through 2026 paints a specific, and specifically uncomfortable, picture of the current ecosystem.</p>



<p class="wp-block-paragraph">Scans across thousands of publicly registered servers have found the large majority carrying file-operation patterns prone to path traversal. A meaningful share of tested servers are vulnerable to command injection or server-side request forgery, and there are documented, disclosed cases of tool description poisoning, where the attack lives in the text a model reads to decide what to do rather than in the code the tool actually executes. A closely related failure mode, configuration poisoning, targets the server’s operational baseline directly: stealthy permission changes or altered defaults that persist across sessions and are hard to catch in a normal code review because the malicious logic lives in configuration state, not application code. Multiple high-severity vulnerabilities, including at least one missing-authentication flaw in a major vendor’s own production package, have already been disclosed and patched.</p>



<p class="wp-block-paragraph">None of that is a reason to avoid MCP. It’s a reason to treat it the way you’d treat any protocol that hands an autonomous caller real privileges inside your systems: skeptically, and with the controls in place before the agent gets access rather than after an incident teaches you why you needed them. In practice that means an explicit, enforced allowlist of vetted tools per agent rather than open discovery of whatever happens to be reachable; authentication on every remote endpoint with no quiet exception carved out for “internal” traffic; centralized, immutable audit logging of every tool call an agent makes; and secrets pulled dynamically from a real secrets manager rather than sitting in a server’s local config where a configuration-poisoning attack can find them. None of this is exotic. It’s the same discipline any experienced integration team already applies to systems with real privileges, applied here to a caller that can now improvise its own sequence of actions.</p>



<p class="wp-block-paragraph">Here’s what I’d tell a team starting today.</p>



<p class="wp-block-paragraph">Treat the tool description as reviewed engineering output, not documentation you write last and skim once. Test it against how an agent actually behaves when given it, not just against whether a human reviewer nods along.</p>



<p class="wp-block-paragraph">Assume composition you didn’t plan for will eventually happen, and design tools to fail safely and legibly when it does, rather than assuming a chain of calls you never tested simply won’t occur.</p>



<p class="wp-block-paragraph">Put governance in front of capability, not after it. The allowlist, the auth, the audit log, and the secrets manager are the entry price given where the current vulnerability data sits, not optional hardening for later.</p>



<p class="wp-block-paragraph">And build against the current specification baseline, not whichever example repository you copied six months ago. The stateless core and the authorization changes in the July 2026 spec aren’t cosmetic; targeting an older baseline today is technical debt you’re taking on knowingly, on day one.</p>



<p class="wp-block-paragraph">MCP earned the “not just another API standard” framing honestly. It didn’t get there by being a cleaner REST, or a nicer SDK, or a better-documented function-calling convention: all real, all incremental. It got there by changing who, or what, is actually doing the integration work at runtime. The parts of that job the protocol doesn’t standardize (how well you describe a capability, how safely your tools behave when composed in ways you never anticipated, and how seriously you take governance before you grant an agent real privileges) are exactly the parts worth taking seriously before you bet production traffic on it.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 117/118 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-10-02 20:49:47 by W3 Total Cache
-->