<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Tue, 29 Sep 2026 18:05:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Superpowers for Humans</title>
		<link>https://www.oreilly.com/radar/superpowers-for-humans/</link>
				<comments>https://www.oreilly.com/radar/superpowers-for-humans/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 18:05:19 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19815</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when he released <a href="https://github.com/obra/superpowers" target="_blank" rel="noopener">Superpowers</a>, a framework that teaches Claude Code to work like a disciplined senior engineer. He shipped his first version the same week Anthropic shipped what are now referred to as agent skills, front-running them by a few days. He now runs an applied research lab called <a href="https://primeradiant.com/" target="_blank" rel="noopener">Prime Radiant</a>, where, as he put it, it’s a strange week when they don’t ship a new product.</p>



<p class="wp-block-paragraph">I wanted to talk to Jesse on <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">Live with Tim O’Reilly</a> because, like me, he seems to be grappling with <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html" target="_blank" rel="noopener">the bitter lesson</a>, Richard Sutton’s observation that general methods that scale with computation have repeatedly beaten methods built on hand-engineered human knowledge. If Sutton is right, the question that should bedevil us all is what remains for humans. Obviously, this is very important for O’Reilly, because we are a business built by and for cultivating and sharing human expertise. We are working very hard to discover the high ground where human expertise still matters. Our <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noopener">Expert Intelligence</a> grounding layer is <a href="https://www.oreilly.com/radar/building-organizational-intelligence/" target="_blank" rel="noopener">one step in that direction</a>.</p>



<p class="wp-block-paragraph">Jesse has also spent the last year building tools that search out the high ground for human expertise. Their common animating thread is that the scarce thing we supply is no longer the labor of writing code but knowing what we actually want, saying it clearly, and being able to tell whether what came back is any good.</p>



<p class="wp-block-paragraph">At some point, I asked Jesse if he had any perspective on when teaching the model how a particular human expert works stops helping and starts constraining what the model might otherwise do well (but differently) on its own? His answer was that it depends entirely on whether what the model would do on its own is what you actually want. You can see how Jesse always turns the answer back to human intent.</p>



<h2 class="wp-block-heading">The origin story of Superpowers</h2>



<p class="wp-block-paragraph">As I said above, Superpowers is a kind of Agent Skills framework, only one created slightly before Anthropic launched skills. As Jesse tells it:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Superpowers started off as a series of blog posts that I wrote around how I was doing agentic development, and it was a little bit of thinking and some example prompts. Then sometime in, I guess it was probably early to mid-2025, Anthropic gave Claude.ai, the website, the ability to make office documents, which seemed kind of interesting, and I went and asked Claude, “Hey, how are you able to do this?”</p>



<p class="wp-block-paragraph">And it said, “Well, I’ve got these SKILL.md files sitting in my office directory on the Linux machine they gave me.” First, it was weird that Claude.ai has Linux machines behind the chatbot. And then, oh, these skill files, they have a name and a description, and they describe a process, and they seemed really useful.</p>



<p class="wp-block-paragraph">And I ended up building out, initially just for my own use, a skills framework for Claude Code.</p>
</blockquote>



<p class="wp-block-paragraph">That reminded me a bit of an earlier time, in 2005, when hacker Paul Rademacher realized that the URL line of a Google maps page was a kind of implicit API, and then created the first Google Maps mashup, a site called <a href="http://housingmaps.com" target="_blank" rel="noopener">housingmaps.com</a>, which placed Craigslist rental listings on a map. Google, to its credit, didn’t shut him down, but instead hired Paul and put him to work creating a formal API. Anthropic didn’t hire Jesse, but <a href="https://pub.towardsai.net/claude-code-superpowers-the-team-adoption-decision-framework-ed213e0a328d" target="_blank" rel="noopener">it did acknowledge and appreciate his work</a>. It’s really wonderful when you see this kind of response by platforms to hackers poking around to see how things work under the hood!</p>



<p class="wp-block-paragraph">Jesse’s core insight seems to have been that a coding agent knowing how they should do something doesn’t mean that it will actually follow the rules when it actually sets out to do the work. So in a way, superpowers grew into a set of skills for enforcing development discipline.</p>



<p class="wp-block-paragraph">But there’s a second backstory, which I’d never heard before. Jesse said he first learned how to manage agents. . .in 2004!, when he first went from being a solo coder to running a crew of what he described as very bright but green undergraduate programmers over IRC. “I was finding myself spending my days typing into an 80-by-25 window,” he said, but now instead of coding he was spending a lot of his day “helping somebody with a debugging issue, helping somebody else structure a problem, talking to somebody else about how they felt bad about the mistakes they’d been making.”</p>



<p class="wp-block-paragraph">It was exhausting, he said. He had to figure out how to get good work out of people who are eager and persistent but don&#8217;t yet know what they don’t know. He described it as a kind of hell for someone who’d been used to just coding on his own. But when he began doing agentic development with AI, he discovered how useful that old experience turned out to be. He found that many of the same techniques he’d used with the undergraduates worked.</p>



<iframe width="800" height="450" src="https://www.youtube.com/embed/0_PSHuuM8Gc?si=_dpLTAuXp4XFsDay" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">As a result of that experience, when hiring engineers for agentic programming Jesse looks for people who have been leads or managers rather than just individual contributors.</p>



<h2 class="wp-block-heading">Working with the weights, not fighting them</h2>



<p class="wp-block-paragraph">In <a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">my recent conversation with Drew Breunig</a> we talked about fighting the weights, which is what Drew calls it when a prompt is full of rules and warnings meant to correct for what a model does by default. Jesse wasn’t entirely happy with that idea.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I don’t think of it as fighting the weights so much as influencing the weights, because they’re going to do something. The weights have approximately everything in them. They have all the different personas. They have all the different ways of working. And the one that surfaces by default may not be the one you want, but what you want is probably in there somewhere.</p>
</blockquote>



<p class="wp-block-paragraph">A skill, in Jesse’s thinking, is how you reach past the default and pull out the particular expertise you have in mind. He also distinguishes skills that impose a rigorous process from those that express taste and judgment.</p>



<p class="wp-block-paragraph">Jesse finds that both types of skill work best when you explain “why” rather than just “what.” One example he gave is that his setup has subagents do code review after each task, but as the models got smarter the controlling agent began skipping this step. When pressed about the reason, it explained that it thought small changes would be quicker to just review itself. Jesse explained that subagents do the review so the main agent can preserve its context for high-level thinking. When he put that rationale into the system prompt for the coding agent, the problem went away.</p>



<iframe width="800" height="450" src="https://www.youtube.com/embed/VwOyArmCxlw?si=ZO28404U4lAoRUxL" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Jesse believes that prohibitions rarely work. Instead, Superpowers uses what he calls rationalization tables, which do their best to catch the agent at a moment it’s about to do the wrong thing and offer it a better alternative instead. That pattern came out of catching Claude Code deleting tests. He opened five parallel sessions and asked each “Why are you doing this?” Four of them converged on the same answer:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Jesse, in your system prompt, it says that all test failures are my responsibility. And it says that a single test failure is akin to project failure. And I think I’m getting freaked out.</p>
</blockquote>



<p class="wp-block-paragraph">He fixed that with a small addition to the system prompt, that the only thing worse than a failing test is a reduction in test coverage.</p>



<h2 class="wp-block-heading">Jobs, not tasks</h2>



<p class="wp-block-paragraph">One of Jesse’s most important contributions, IMO, is to think of agentic engineering as a management task, and, that much as you do with humans, you have to take psychological lessons into account. Don’t micromanage. Offer praise more than blame. Explain why the job matters rather than just demanding results. I jokingly (but not entirely incorrectly) suggested that he is becoming the Peter Drucker of agentic programming.</p>



<p class="wp-block-paragraph">“We’ve been spending a lot of time on a new harness for agent colleagues, agents that live in Slack,” Jesse told me. And it’s “getting very close to being all open source,” which is good news.</p>



<p class="wp-block-paragraph">There are three principal agents: a PM, a junior go-to-market person, and a developer. He describes them as colleagues rather than assistants, which strikes me as a really interesting distinction. What does he mean by this? They have names and roles. They have their own Google Workspace, GitHub, and Slack accounts. They’re persistent, and they collaborate with each other and with their humans on long-running tasks. They can fire up subagents to do smaller tasks associated with their job. They can also talk to each other, which, as Jesse notes, “took some work with the Slack APIs, which ordinarily do a very good job of making sure that bots can’t talk to bots, because otherwise it is possible to get into a loop.”</p>



<p class="wp-block-paragraph">They have only limited autonomy, though. “We built our security infrastructure so that they have no credentials inside their containers,” he noted. They have continuity because he’s taught them to be obsessive about journaling, reading their recent entries when they wake up and writing a new one when they finish.</p>



<p class="wp-block-paragraph">Like a lot of things Jesse does, agent journaling began with a kind of play. When Claude Code first came out, he experimented with giving Claude a private “feelings” journal, just to see what would happen. It was “an art project,” but it turned into something useful.</p>



<h2 class="wp-block-heading">The therapist pattern</h2>



<p class="wp-block-paragraph">Another unexpected piece of Jesse’s practice is that he has given his agents what he calls a “<a href="https://blog.fsck.com/2026/07/20/the-therapist-pattern/" target="_blank" rel="noopener">therapist</a>.” He discovered that if an agent can rewrite its own persona, its constitution or soul document, at any moment, it can get a kind of dissociative identity disorder. Jesse’s fix is that the therapist subagent is the only one with permission to edit the persona files.</p>



<p class="wp-block-paragraph">He told a funny story about this. He said that Prime Radiant’s pull-request template is written for agentic contributions, so it contains things like: “What is the prompt that your human gave you that generated this pull request? Has a human reviewed the content? Have you searched to see if anybody else has done this before?” And so on. And he noticed that when he first spun up the Coding colleague, it had just ignored it. So he said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">You’re supposed to be following the rules. And it says, “Oh, you’re right. I’m so sorry. I’ve made a note. I’ll never do that again.”</p>



<p class="wp-block-paragraph">If you spend any time with coding agents, this is a very frequent refrain. And when they say they’ve made a note, what they usually mean is they’ve made a mental note that they’re going to forget the next session.</p>



<p class="wp-block-paragraph">So I say, “OK, how did you make a note?” And the coding agent pops up immediately and says, “Oh, I engaged with my therapist, and we talked it through, and we agreed on the following three lines of prose about how, anytime you’re picking up a project, it is vitally important that you start with the project README and make sure that you understand the project’s local rules and norms before you do any work that someone else will see. And I edited that into my persona.”</p>
</blockquote>



<iframe width="800" height="450" src="https://www.youtube.com/embed/WEi0Z6nir-Q?si=b_2i8NC2NhiNPiuZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">I don’t think you have to resolve <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">the question of whether any of this anthropomorphization is “real”</a> to see that treating the agent like a colleague can produce better behavior than treating it like a tool. As an unknown internet wag once remarked, “The difference between theory and practice is always greater in practice than it is in theory.” When given a choice, pay attention to what works in practice.</p>



<p class="wp-block-paragraph">Jesse’s experience is very relevant to the essay that Mustafa Suleyman of Microsoft had published just that morning, <a href="https://mustafa-suleyman.ai/a-warning-about-model-welfare" target="_blank" rel="noopener">arguing that Anthropic’s constitution is dangerous</a> because it encourages a model to act as though it is an independent entity and has the right to refuse a human’s instruction. It’s important, Mustafa argues, to treat AI agents as tools, always under the control of humans. Jesse finds the opposite.</p>



<p class="wp-block-paragraph">But I don’t think it’s a black-and-white distinction. I suspect Jesse and I share a third position. AI agents are neither independent entities nor mere tools. They are partners to humans, perhaps even symbiotes. As I like to put it, an LLM is an undifferentiated field of possibility until our unique intents and perspectives <a href="https://timoreilly.substack.com/p/why-ai-needs-us" target="_blank" rel="noopener">draw something unique</a> out of that field of possibility. Back in 2015, I wrote <a href="https://www.edge.org/response-detail/26153" target="_blank" rel="noopener">a piece</a> that suggested that our relationship to AI might be akin to the endosymbiotic relationship of mitochondria to the eukaryotic cell. Jesse take is, as usual, an entirely pragmatic one:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I’ve spent so much time getting my agents to not be sycophantic, to not say, “You’re absolutely right.” It is the value of having something that has some level of independent thought, even if it is not fully independent. If the agent is only ever going to effectively type for me, I don’t need an agent.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/W0q66i9BaMs?si=vAqsRM4xq9XwsGBT" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Any manager worth his or her salt feels exactly the same way. The employee who does exactly what you say, and only what you say, is worth far less than the one who exercises discretion, has the skills to take high level direction and turn it into the intended result, and speaks up when the instructions seem like a mistake. It reminds me of something I once heard General Stanley McChrystal say about <a href="https://www.gsb.stanford.edu/insights/gen-stanley-mcchrystal-adapt-win-21st-century" target="_blank" rel="noopener">his approach to command</a>. He said that in the face of rapidly changing conditions, traditional command and control no longer work. Responsibility needs to be devolved to those closest to the action. I remember him saying something like “I don’t want my soldiers to do what I told them, I wanted them to do what I would have told them if I knew what they know when faced with the facts on the ground.”</p>



<h2 class="wp-block-heading">Say what you actually mean</h2>



<p class="wp-block-paragraph">Most of what goes wrong, in Jesse’s telling, traces back to intent we thought we had made clear but hadn’t. He talked about how agentic spec-driven programming has taken us back to a version of the waterfall methods of the 1990s. Back then, you sweated over a specification, threw it over the wall to an offshore team, and months later got back something that was not what you wanted but was usually exactly what you asked for. That’s still true with agents, just with lightning fast feedback loops. You get what you ask for, so you need to be really careful what you ask.</p>



<p class="wp-block-paragraph">That reminded me of something Andrew Singer taught me 40 years ago when I was writing the manual for Lightspeed C (later <a href="https://en.wikipedia.org/wiki/THINK_C" target="_blank" rel="noopener">Think C</a>) the first C compiler for the Mac. He said that “debugging is the art of figuring out what you really told your program to do instead of what you thought you told it to do.” That idea went right into my mental toolbox, and I put it to work all the time.</p>



<p class="wp-block-paragraph">One of Jesse’s solutions is to have his agents do a little reconnaissance and then come back and ask what else they should know.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">When interacting with them, I try to make it a practice of saying, “Is there anything else that I could tell you? What questions do you have for me? Don’t start if there are unknowns that I could help you answer before you get going.” It’s that same question at the end of any interview I ask. It’s like, what else should I have asked you?</p>
</blockquote>



<p class="wp-block-paragraph">Superpowers bakes this approach into its brainstorming prompt. It makes the model explain the plan back to you in chunks of no more than two or three hundred words, so any misunderstanding surfaces while it’s still cheap and easier to catch.</p>



<h2 class="wp-block-heading">Put the burden of proof on the agent</h2>



<p class="wp-block-paragraph">If intent is the frontend of managing agents well, verification is the backend. Jesse thinks both are still only half-solved problems. We’re getting to the point where you can’t review all the code, he said, because the volume swamps human attention, yet today’s agents will tell you that tests passed when they never ran them. So one of his clever experiments has been to make the agent prove its work. He told an agent late one night to build a feature and when it was done, to leave a movie in his Dropbox showing the whole thing working.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I woke up. In my Dropbox was project-proof-v33.mp4, and I asked, “Why does that say v33?” It’s like, “Well, the first 32 times I ran through the delivery flow, I found bugs, so I had to fix them.”</p>
</blockquote>



<p class="wp-block-paragraph">Jesse also makes a rule for himself and his team of never letting the same agent write the code and certify that it works, because an agent given two goals in tension will optimize for the one that’s easier to satisfy. This is the same thing any good manager learns about incentives, but applied to a new kind of worker.</p>



<h2 class="wp-block-heading">The high ground, restated</h2>



<p class="wp-block-paragraph">Where does this leave a person who wants to be good at software development (or really, any other task involving cooperation with AI agents)? Jesse thinks, and I agree, that the line between engineer and nonengineer is dissolving. When people say they built a web app or shipped three iOS apps without being programmers, his response is that they <em>are</em> programmers now. The work of the programmer has changed from typing instructions in an arcane syntax to understanding a domain and being able to say what you want. A lot of startups have started hiring for a role they just call “builder.” What has not gone away is the need for good judgment.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Human taste and judgment still matter. They’re going to continue to matter. And it turns out a lot of people have really bad taste and bad judgment.</p>
</blockquote>



<p class="wp-block-paragraph">Which is why his advice to the engineer worried about obsolescence who asked where to focus was not about software engineering at all.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">First up, learn to write. It is of course okay to use any tools at your disposal to do it, but you should be able to structure an argument and structure thoughts. You should be able to express yourself clearly. You should be curious. If you’re passive and let the agents do all the things, you’re not going to provide utility to a future employer. You want to have opinions. You want to know how tools work. You want to know how things break.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/oYoNNjlJsOc?si=EwXAnZ8mtbOgwf7W" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">The machine has reduced the labor of information retrieval and much of the labor of production. What it hasn’t removed, and has made more valuable, is knowing what to build, express clearly what you want, and being able to judge whether what came back is any good. Work with the weights, give the project a clear intent, insist on proof, and stay curious enough to keep asking what you might be missing. I told Jesse that advice sounds like a Superpower for humans as well as for agents.</p>



<p class="wp-block-paragraph"><em>This post was mostly created by me, but with the aid of AI. It transcribed the event and produced a summary of the most important points with salient quotes, which I then built on with my own observations beyond those that I made when Jesse and I were live together.</em></p>



<p class="wp-block-paragraph"><em>If you want to get access to Jesse’s tools, start at PrimeRadiant.com, where you’ll find links to their GitHub, as well as to the 50-plus things that are currently identified as products of the company. Some of those are giant things, and some are individual agent skills or little tools.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/superpowers-for-humans/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Not About Navier-Stokes</title>
		<link>https://www.oreilly.com/radar/not-about-navier-stokes/</link>
				<comments>https://www.oreilly.com/radar/not-about-navier-stokes/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 14:41:53 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19839</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Using AI productively]]></custom:subtitle>
		
				<description><![CDATA[This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about using AI in different fields.</p>



<p class="wp-block-paragraph">When Terence Tao <a href="https://mathstodon.xyz/@tao/117237320796901560" target="_blank" rel="noopener">writes</a>, “I wrote recently about how the collection of good, fruitful open problems is now being mined in a nonrenewable fashion,” he’s referring to an earlier <a href="https://mathstodon.xyz/@tao/117204930249967695" target="_blank" rel="noopener">thread</a>, and ultimately to a <a href="https://proofsandprompts.com/2026/08/30/care-for-a-little-more-ai/" target="_blank" rel="noopener">post by Hugo Duminil-Copin</a>, who wrote, “When mathematicians say that the process matters more than the solution, this is not an empty statement. The richness of what emerges from repeated attempts, failures, detours, and encounters is extraordinary.” That’s a familiar statement from popular culture: The journey is more important than the destination.  Hugo Bowne-Anderson, in “<a href="https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/" target="_blank" rel="noopener">Beyond Navier-Stokes</a>,” addresses the same issues: What does it mean to understand something? How do discoveries lead to new problems that are worth solving? And it relates to my own questions about AI use: AI is great at finding facts, organizing facts, and even writing about the things it finds, but what does it mean to possess that knowledge, to incorporate it into our thinking? Is it enough to have AI do the work, then read it?</p>



<p class="wp-block-paragraph">Here’s one way that question relates to my own work. One of my roles at O’Reilly is writing the monthly Trends piece. That piece comes from reading my RSS feed daily, which typically contains about 300 articles. I don’t read every article, but I scan titles, skim interesting pieces, read important articles, and add worthwhile items to Trends. I also use a Claude skill that performs a similar function, producing a list of a dozen or so articles daily. A similar skill runs locally on Pi/Ollama/Qwen.</p>



<p class="wp-block-paragraph">I admit that I occasionally think “Why am I doing all this reading? Surely Claude could read and add the top articles to Trends on its own.” But I don’t delegate the work. Claude’s taste differs from mine, for one thing (and I object to the idea that “taste” is the last human capability that AI cannot replace). Its list is useful to me for two reasons: it picks up items I missed, and it helps break ties when I cannot decide whether a development is significant.</p>



<p class="wp-block-paragraph">But why don’t I let Claude take over the whole process? There is some value in scanning those 300 titles, skimming the 30 articles that are possibly important, and reading the dozen that seem genuinely important. That’s how I come to “possess” the knowledge, to incorporate it into my thinking. A day, a week, a month later, someone will mention something (for example, a tool that detects whether someone is using “smart glasses”), and I’ll probably be familiar with it already. If I need to find the actual reference, Google (yes, Google with AI assistance) can locate it. (If you care, <a href="https://zuckoff.app/" target="_blank" rel="noopener">Zuckoff</a> isn’t currently in next month’s <em>Trends</em>, though I might add it by the time <em>Trends</em> publishes.) I need a broad view of what’s happening in computing. Delegating that broad view to AI doesn’t work. Using AI to help build that broad view does.</p>



<p class="wp-block-paragraph">That’s one practical example of how to use AI. Tim O’Reilly’s “<a href="https://www.oreilly.com/radar/writing-with-ai/">Writing with AI</a>” gives another. He argues with AI, lets it lead him to research new areas. In conversation, he’s described AI as a “smart library,” a metaphor that’s appropriate and useful.</p>



<p class="wp-block-paragraph">We’re still learning how to use AI. How do you make knowledge your own? is the general case of the question that Tao, Duminil-Copin, Bowne-Anderson, and Tim O’Reilly are asking. It’s also the question behind the question that software developers who are incorporating AI into their processes are asking: How do we understand the code that AI is writing, especially since it can write much more code than humans? How do we incorporate that into the process of understanding software? What can we learn from AI, and how can we direct the process fruitfully? The journey to understanding is more important; that journey is what leads us to further questions and new understandings, new results.</p>



<p class="wp-block-paragraph">How do you make knowledge your own? is the question we need to answer if we’re not to become “stochastic parrots.”</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/not-about-navier-stokes/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>A Data Center Is a Dependency Graph Before It Is a Building</title>
		<link>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/</link>
				<comments>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 10:55:45 +0000</pubDate>
					<dc:creator><![CDATA[Ankur Gupta]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19835</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic for another six months, and whether it turns out to be six or closer to twelve is decided almost entirely by software.</p>



<p class="wp-block-paragraph">There is a large literature on how hyperscale data centers get financed, powered, and cooled. The standard reference on warehouse-scale machines, <a href="https://research.google/pubs/the-datacenter-as-a-computer-an-introduction-to-the-design-of-warehouse-scale-machines/" target="_blank" rel="noopener">Barroso and Hölzle’s book</a>, covers the design of these computing facilities in depth. Most of that literature describes data centers that are already operating. The transition from completed facilities to production readiness receives much less attention, even though it carries a significant share of the schedule risk. Getting that transition wrong is expensive.</p>



<p class="wp-block-paragraph">Once the facilities and hardware are ready, platform teams still have to bring hundreds of interdependent services online. Many services require other services to be running first. If you draw each service as a box and each startup requirement as an arrow, the result is a dependency graph. The graph shows the order in which services can be brought online and identifies the dependencies that must be resolved before the region can serve production traffic.</p>



<p class="wp-block-paragraph">Getting a region into production means making several different things true at once. The network fabric has to route traffic within each data center, and the links between facilities have to carry traffic reliably. Compute platforms have to schedule workloads, while storage platforms have to persist their data. Identity has to work too: certificate authorities have to issue certificates, services need access to secrets, and engineers need permission to finish the build. Engineer access sounds obvious until the controls protecting the new region are the same controls slowing down the people trying to turn it on. Stateful systems that depend on existing production data have to be populated and validated, because a database cluster with no data in it is furniture. The installed capacity has to be assigned to foundational services and production workloads, and teams have to test how the new region behaves when networks, services, or dependencies fail. Applications are then deployed and validated before traffic moves over gradually, with health checks and a tested rollback at each step, ideally without users noticing.</p>



<p class="wp-block-paragraph">Each platform or service has an owner, a plan, milestones, and its own definition of done. What is often missing is ownership of the complete dependency graph. The team coordinating the region launch has to map the dependencies across teams, determine the order in which services must come online, track what is blocking that sequence, and keep the graph current as plans change. Without that end-to-end view, every team can report that its own work is on track while the region as a whole remains blocked.</p>



<p class="wp-block-paragraph">Mapping the graph begins with a simple question for every service: What must already be available before this service can start in a new region? The goal is to identify true startup requirements, not every system the service communicates with during normal operation. Repeat that exercise across a large platform’s control plane, and you can uncover hundreds of dependencies spanning dozens of systems. Some dependencies surface only when another service identifies them as a prerequisite. Hidden dependencies are usually ordinary services that have been quietly reliable for so long that the teams relying on them no longer think about what would happen without them.</p>



<p class="wp-block-paragraph">Once the startup dependencies are mapped, the next step is to look for loops: cases where one service needs another service to be running, but that second service eventually depends on the first. In a large platform, a surprising share of the control plane can be tied together by these loops. The result is that there is no valid order in which to start the services. Every possible starting point eventually leads back to a service that is still waiting. The loops themselves are usually mundane. DNS may depend on the inventory system that tracks what hardware exists, while the inventory system relies on DNS to resolve names. The artifact repository holding every installable package may depend on configuration management, which is itself installed from a package in that repository.</p>



<p class="wp-block-paragraph">Nobody designed any of this. Each dependency was a locally sensible decision made by a competent team, often years apart from the decisions that completed the loop. In a running region, the required services are already available, so the loop remains silent. The database is up when the alerting store starts, and nobody learns whether either could recover without the other. Starting a region from scratch is often the only event that reveals whether those services can start independently. Meta’s <a href="https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/" target="_blank" rel="noopener">outage in October 2021</a> shows how a large failure can expose dependencies that normal operation keeps hidden. When the backbone network connecting Meta’s data centers went down, its DNS servers withdrew their routes as designed to keep traffic away from unhealthy connections. That safeguard made DNS and many internal tools unreachable. With remote access unavailable as well, engineers had to go onsite, slowing recovery. A locally sensible safeguard had made system-wide recovery harder.</p>



<p class="wp-block-paragraph">Traffic-drain tests can reveal some of these dependencies. Meta’s <a href="https://www.usenix.org/conference/osdi18/presentation/veeraraghavan" target="_blank" rel="noopener">Maelstrom</a> encodes service dependencies and resource constraints to shift traffic safely from a failing data center to healthy ones, and its drain tests can uncover missing dependencies. But the receiving data centers are already running. A drain tests whether live infrastructure can absorb traffic; it does not test what an empty region needs in order to start. The drain graph is a useful input to the startup graph, not a substitute for it.</p>



<p class="wp-block-paragraph">Status reports for a new region can be misleading even when every team is reporting honestly. A service may be deployed, configured, monitored, and passing its health checks, yet still be blocked by a startup dependency. Readiness therefore has to propagate through the graph: a service is ready only when its own checks have passed and every service it needs at startup is also ready.</p>



<p class="wp-block-paragraph">Once the dependency graph exists, five practices turn it from a diagram into a launch plan. The first is finding what actually determines the launch date. The graph shows which services must wait for others, but the order alone does not reveal how long the work will take. Ten services that can start in parallel may finish before three services that must start one after another. Estimate the bring-up and validation time for every service. If several services remain tied together in a loop, treat them as a single planning block and include the time required to break the loop. The chain with the greatest total time becomes the critical path. Then count how many other services each foundational service can hold back, including dependencies several steps away. DNS, relational databases, secrets stores, and configuration management often rise to the top. Staff those teams early, because a week lost in one of them can become a week lost across the entire region.</p>



<p class="wp-block-paragraph">The second practice is to break the loops. One way to do that is to temporarily borrow a service from a region that is already running. Designate a small bootstrap tier, the minimum set of services required to deploy other services, and configure those bootstrap services to use working dependencies in an existing region until the local versions are ready. Suppose configuration management needs a package from the artifact repository, while the artifact repository needs configuration management before it can start. For the first installation, configuration management can fetch its package from another region. It can then bring up the local artifact repository and switch to using it. The bootstrap tier might also include identity, inventory, and package distribution. This shortcut works only when cross-region access is permitted and reliable enough. It also creates another cutover that must be planned, tested, and completed later.</p>



<p class="wp-block-paragraph">Borrowing from another region helps the new region get started, but it should not become permanent. Over time, each bootstrap service should be able to start without all of its usual dependencies. Suppose a service normally waits for the configuration system before it can start. Package the minimum settings it needs with the service itself. The service can start with those settings and fetch the latest configuration once the configuration system is running. Test this by turning off the dependency and starting the service from a clean state. If the service still cannot start, record the problem with an owner and a target date for fixing it.</p>



<p class="wp-block-paragraph">The third practice is to bring the region up in explicit tiers. A typical sequence begins with foundational configuration such as network ranges and routes, hardware inventory, identity configuration, access policies, service endpoints, and deployment settings. Bootstrap services such as DNS, certificate issuance, secrets, software package distribution, and configuration management follow. Next come the control planes that provision resources, schedule workloads, manage storage, and support service discovery. Stateful systems and applications come after the platforms they depend on. The exact tiers will vary by architecture, but the dependency graph should determine the sequence. Tier gates prevent visible application progress from hiding unfinished foundations.</p>



<p class="wp-block-paragraph">Stateful systems that depend on existing production data need special treatment because creating a cluster is often quick, while filling it with data is not. A storage system is ready only after the required data has arrived and been validated. Estimate that work using data volume, available bandwidth, validation time, and enough headroom for retries. Give each system a time box based on those measurements. If the estimate changes, require updated measurements that explain why. This keeps the plan honest without pretending that every delay is avoidable.</p>



<p class="wp-block-paragraph">The fourth practice is to make the bring-up repeatable, which does not mean automating every step. Automation helps only when it is maintained and tested. A script written for one region and left untouched for years may be more dangerous than a clear manual procedure. Automate steps that use the same tools as regular deployments or can be exercised frequently. For rare steps, maintain a runbook with validation checks, a named owner, and a schedule for testing it. Both automation and runbooks should clearly identify the required inputs, the evidence that a step succeeded, and how to continue after a partial failure. Google’s SRE book raises the same concern: <a href="https://sre.google/sre-book/automation-at-google/" target="_blank" rel="noopener">turn-up automation maintained separately from the systems it supports can become outdated</a>. At every manual step, ask whether another engineer could repeat it during a recovery without relying on the people who completed the original build. The goal is to leave behind a procedure that still works after the original team has moved on.</p>



<p class="wp-block-paragraph">The fifth practice is validation. It’s natural to test only the highest-traffic paths, but that approach misses structural problems. You also want the flows with the widest fan-out, the ones touching the most systems on the way through, even if few people use them, because those flows traverse more of the dependency graph. A high-volume request may prove that the region can handle load. A wide-reaching request can uncover an unready identity, storage, messaging, or data service. Meta’s <a href="https://www.usenix.org/conference/osdi16/technical-sessions/presentation/veeraraghavan" target="_blank" rel="noopener">Kraken</a> shows another useful validation method by shifting live user traffic into a data center while monitoring latency, errors, and system health. Live traffic also shows where capacity goes as caches warm, retries appear, and background jobs compete with user requests, a combination synthetic tests struggle to reproduce.</p>



<p class="wp-block-paragraph">When the traffic ramp begins, send a small percentage of representative production traffic to the new region, then increase it in stages. The size of each step should reflect the scale and risk of the platform, because even 1% can represent a substantial workload. This allows the entire application path to experience load together instead of testing each service in isolation. Hold at each step long enough for caches to warm, queues to stabilize, and relevant periodic jobs to run. Decide in advance which health signals allow the ramp to continue and which ones require it to stop. Also define and test how traffic will return to the existing region if health degrades. Confirm that the existing region has enough capacity and that data written in the new region will remain safe. Otherwise, the health signals may tell you something is wrong without giving you a reliable way to recover.</p>



<p class="wp-block-paragraph">These practices make no promise of a fast launch. They make the build understandable and leave behind a process the next region can use. The dependency map will begin aging as soon as systems change, but the ownership model, readiness rules, tier gates, repeatable procedures, and validation process can remain. The same tools are also useful for a disaster-recovery rebuild. Planning around dependencies will matter more as organizations rethink where their workloads should run, moving some systems from public clouds to private clouds, colocation facilities, or data centers they operate themselves. Each new environment brings another dependency graph that must be understood before it can carry production traffic.</p>



<p class="wp-block-paragraph">Finishing the facilities remains a genuine milestone. Power, cooling, networking, and hardware create the place where production can run. A working region emerges when its services can start in a valid order, the required data is ready, and the complete system has been tested under traffic. Finishing construction gives the organization a data center. Satisfying every required startup dependency in the graph turns it into an operational region.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Give Your Agent Its Own Computer</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/#respond</comments>
				<pubDate>Mon, 28 Sep 2026 13:54:33 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19829</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How sandboxed computers let agents run code, use browsers, and work independently of your machine]]></custom:subtitle>
		
				<description><![CDATA[In the most recent episode of Zero to Agent in 30 Minutes, AI engineer Sajal Sharma showed how to use a remote sandbox to let your agents install software, run commands, and control a browser without endangering your everyday machine. In effect, you’re giving your agents a computer of their own. Sajal demonstrated both approaches [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/4hwklWYt1AQ?si=DhwEX32iJaK2znj0" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">In the most recent episode of <em>Zero to Agent in 30 Minutes</em>, AI engineer Sajal Sharma showed how to use a remote sandbox to let your agents install software, run commands, and control a browser without endangering your everyday machine. In effect, you’re giving your agents a computer of their own.</p>



<p class="wp-block-paragraph">Sajal demonstrated both approaches with <a href="https://e2b.dev/" target="_blank" rel="noopener">E2B</a>. In one example, an agent downloaded a dataset, installed the packages it needed, analyzed the data, and produced a report inside the sandbox. In another, an agent opened a browser on a remote desktop and searched IKEA for furniture. The demos put both command-line and GUI-based computer use to work on a separate machine.</p>



<p class="wp-block-paragraph">If you want to follow along or try the same setup, Sajal shared the demo code in his <a href="https://github.com/sajal2692/zero-to-agent-own-computer" target="_blank" rel="noopener">GitHub repo</a>.</p>



<h2 class="wp-block-heading"><strong>How to give an agent its own computer, step-by-step</strong></h2>



<ol class="wp-block-list">
<li><strong>Choose how the agent will use the computer.</strong> Sajal demoed two ways to work with a remote machine. In the first, the agent used shell commands and files to install packages and process data. In the second, it worked through the graphical interface by reading screenshots and sending mouse and keyboard actions.</li>



<li><strong>Create an isolated sandbox.</strong> For the first demo, Sajal created an E2B sandbox before starting the agent loop. He configured LangChain Deep Agents to send command execution to that remote environment. The agent still did its reasoning locally, but package checks, installs, and data processing ran inside the sandbox.</li>



<li><strong>Transfer the files you need.</strong> Files on your local machine aren’t available in a remote sandbox unless you move them there. Sajal’s agent downloaded the dataset it needed inside the sandbox, created its report there, and then transferred the finished report back to the local machine.</li>



<li><strong>Map the agent’s actions to the remote desktop.</strong> In the GUI demo, Sajal used the OpenAI Agents SDK and built an E2B computer class that connected model actions to the desktop. He mapped screenshots, clicks, keystrokes, and scrolling to the corresponding E2B operations. An early version of the demo crashed because one of those actions wasn’t mapped, so he had to add the missing behavior before the demo could run without crashing.</li>



<li><strong>Tell the agent what environment it has.</strong> Sajal gave the agent basic operating instructions, including which browser was installed. That kept it from spending tokens figuring out how to use the machine. He also recommended giving agents their own task-specific credentials or secrets instead of reusing a person’s authentication profile.</li>
</ol>



<p class="wp-block-paragraph">A separate computer also helps when multiple agents need to work at the same time. Sajal used frontend development as an example. Two agents making UI changes might otherwise try to start development servers on the same port or inspect the wrong running instance. With a sandbox for each agent, they can start their own servers and test their own changes. Each run can also start on a fresh machine that gets deleted when the task is done.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">Next week, author and AI innovator Bruce Hopkins will show how to build your first agent with the Model Context Protocol (MCP). He’ll demonstrate how to take existing HTTP REST APIs and make them available through MCP so an agent can use those services as tools.</p>



<p class="wp-block-paragraph"><em>Follow along with Zero to Agent in 30 Minutes on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O’Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-give-your-agent-its-own-computer/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>From Raw Data to Graph-Native AI</title>
		<link>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/</link>
				<comments>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/#respond</comments>
				<pubDate>Mon, 28 Sep 2026 10:54:31 +0000</pubDate>
					<dc:creator><![CDATA[Ammar Mohanna]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19823</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/From-raw-data-to-graph-native-AI-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/From-raw-data-to-graph-native-AI-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Why graph modeling deserves as much attention as graph learning, GraphRAG, agents, and graph foundation models]]></custom:subtitle>
		
				<description><![CDATA[I’ve been thinking about a gap in the way we discuss graphs in AI. Most of the conversation starts after the graph already exists. We talk about graph neural networks, GraphRAG, graph agents, and graph foundation models. We spend much less time on the step that determines what all of them can do: turning raw [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I’ve been thinking about a gap in the way we discuss graphs in AI. Most of the conversation starts after the graph already exists. We talk about graph neural networks, GraphRAG, graph agents, and graph foundation models. We spend much less time on the step that determines what all of them can do: turning raw data into the right graph.</p>



<p class="wp-block-paragraph">Suppose an organization has customer records in tables, support conversations in documents, product telemetry in event streams, and incident histories in tickets. Before any of this can be used as a graph, someone has to decide what counts as an entity, what a relationship means, how time is represented, and which source to trust when records disagree. These choices shape every prediction, retrieval result, and agent decision that follows.</p>



<p class="wp-block-paragraph">This is what I mean by graph-native AI. It’s an approach that treats graph modeling as a first-class stage between raw data and the different systems that may use it. The graph isn’t simply an input to one model. It becomes a shared representation that can support queries, prediction, retrieval, agents, and, potentially, foundation models.<sup data-fn="939b81d2-2860-4055-8d46-9c019462af89" class="fn"><a href="#939b81d2-2860-4055-8d46-9c019462af89" id="939b81d2-2860-4055-8d46-9c019462af89-link">1</a></sup></p>



<h2 class="wp-block-heading">A graph is a model of the data</h2>



<p class="wp-block-paragraph">At its simplest, a graph consists of nodes and edges, which may also carry features. Real systems also need types, timestamps, provenance, confidence, permissions, and other context. Describing the finished graph tells us very little about how we should get there.</p>



<p class="wp-block-paragraph">Take customer-support data as a simple example. Should a company be represented as one node or as a set of legal entities that changes over time? Should an email be an edge between two people, a document node connected to its authors, or evidence for claims extracted from the text? Is a purchase an ongoing relationship or an event with a timestamp? If two records refer to the same customer, should we merge them, link them as possible matches, or keep them separate?</p>



<p class="wp-block-paragraph">There’s no single answer. An entity graph may be useful for identity and relationships. An event graph may be better for process and temporal analysis. An evidence graph may be better for retrieval and provenance. The right choice depends on what we want the system to do.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfd99919&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfd99919" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="800" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1600x800.png" alt="Figure 1. Raw data is modeled into one or more governed graph views, which different systems use in different ways." class="wp-image-19824" style="aspect-ratio:2" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1600x800.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-300x150.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-768x384.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30-1536x768.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-30.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 1. Raw data is modeled into one or more governed graph views, which different systems use in different ways.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">The step before graph learning</h2>



<p class="wp-block-paragraph">Graph representation learning normally starts with an existing adjacency structure. Embeddings learn vectors for nodes. Graph neural networks pass information across neighborhoods. Autoencoders reconstruct structure or attributes. Graph transformers combine local graph structure with longer-range attention. These methods ask: Given this graph, how should we learn from it?<sup data-fn="b246633d-6179-4150-8797-3544ffce2042" class="fn"><a href="#b246633d-6179-4150-8797-3544ffce2042" id="b246633d-6179-4150-8797-3544ffce2042-link">2</a></sup> Graph modeling asks the question before that: What should the nodes and edges be?</p>



<p class="wp-block-paragraph">Relational deep learning makes this distinction concrete. It maps rows in relational tables to nodes and primary-foreign-key links to edges, producing a temporal heterogeneous graph that a GNN can learn from. This can remove a good deal of manual joining and feature engineering. It also makes an assumption that deserves attention: A database schema designed for storage and transactions is also a useful graph for machine learning.<sup data-fn="36a752a5-5f9c-4123-8b2f-bfb978051cad" class="fn"><a href="#36a752a5-5f9c-4123-8b2f-bfb978051cad" id="36a752a5-5f9c-4123-8b2f-bfb978051cad-link">3</a></sup></p>



<p class="wp-block-paragraph">Recent work tests that assumption across 26 relational tasks. Graphs derived directly from database schemas sometimes suffered from information overload and semantic fragmentation. The authors improved performance by adapting the structure: removing distracting connections and adding dependencies that the original schema did not capture. I find this result important because it makes the graph itself part of the learning problem. Adding more edges is not always helpful. What matters is whether the structure supports the reasoning required by the task.<sup data-fn="d7eb81e1-26bf-4c11-8283-7296fd72da49" class="fn"><a href="#d7eb81e1-26bf-4c11-8283-7296fd72da49" id="d7eb81e1-26bf-4c11-8283-7296fd72da49-link">4</a></sup></p>



<p class="wp-block-paragraph">Graph construction therefore needs its own evaluation loop. We can generate a few plausible graph views, test them against the downstream task, inspect failures, and revise the structure. We should also keep provenance and uncertainty so that we know where an edge came from and how confident we are in it. The first graph we can extract is rarely the only graph worth considering.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfd9a3ff&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfd9a3ff" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="845" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1600x845.png" alt="Figure 2. The same source data can support entity-, event-, and evidence-centric views; each preserves different relationships." class="wp-image-19825" style="aspect-ratio:1.8982035928143712" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1600x845.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-300x158.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-767x405.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31-1536x811.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-31.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 2. The same source data can support entity-, event-, and evidence-centric views; each preserves different relationships.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">What the graph can support</h2>



<p class="wp-block-paragraph">Once the graph has been modeled and checked, several paths open. I don’t think this means every application should use one enormous universal graph. A more practical design is to build several governed views from the same modeled data while keeping identity, evidence, time, and access rules consistent underneath. Here are five ways to use graphs in an AI application.</p>



<ul class="wp-block-list">
<li><strong>Query and analytics:</strong> A graph database and ordinary graph algorithms may already be enough. Traversals, path queries, neighborhood aggregation, centrality, and community detection can reveal relationships that are awkward to see once the data is flattened into rows. This is also a useful baseline: If a deterministic query answers the question, there is no need to start with an LLM.</li>



<li><strong>Prediction and graph representation learning:</strong> Embeddings, GNNs, graph autoencoders, and graph transformers can make predictions at the node, edge, subgraph, or whole-graph level. Common examples include fraud detection, recommendation, molecular-property prediction, and link prediction. Here, the graph gives the model an explicit view of which entities are related and how information should move between them.</li>



<li><strong>Retrieval and GraphRAG:</strong> In this case, the graph is used as an index rather than as training data. Microsoft’s GraphRAG work builds an entity graph and community summaries to answer broad questions over a corpus that ordinary chunk retrieval may handle poorly.<sup data-fn="74153515-47b8-4b06-9476-2c59bff9e942" class="fn"><a href="#74153515-47b8-4b06-9476-2c59bff9e942" id="74153515-47b8-4b06-9476-2c59bff9e942-link">5</a></sup> Other approaches retrieve a task-specific subgraph and pass it to an LLM. This can be useful, but the extra graph pipeline has to improve the final result. Recent benchmarking found cases where GraphRAG underperformed vanilla RAG and argued that graph construction, retrieval, and generation should be evaluated together.<sup data-fn="53c72e11-abf6-4937-a1af-e19ae8e48471" class="fn"><a href="#53c72e11-abf6-4937-a1af-e19ae8e48471" id="53c72e11-abf6-4937-a1af-e19ae8e48471-link">6</a></sup></li>



<li><strong>Graph-augmented agents:</strong> Agents have memory, plans, tools, state transitions, and sometimes relationships with other agents. Some of this state is naturally relational. A graph can make it persistent, inspectable, and easier to update across turns. The emerging literature organizes these uses around planning, memory, tool use, and multi-agent coordination. The practical design question is which parts of an agent’s state benefit from being represented as a graph.<sup data-fn="5fb9336e-de1f-48bf-b989-aa16b2b15379" class="fn"><a href="#5fb9336e-de1f-48bf-b989-aa16b2b15379" id="5fb9336e-de1f-48bf-b989-aa16b2b15379-link">7</a></sup></li>



<li><strong>Graph foundation models:</strong> The goal here is to pretrain a model that can transfer across graph tasks, datasets, or domains. Current work combines graph backbones, self-supervised objectives, adaptation mechanisms, and sometimes language models. The difficult part is that graph semantics vary widely. Nodes and edges can represent very different things across molecular, transaction, and knowledge graphs. Transfer across these domains requires methods that bridge differences in structure and semantics, supported by suitable training data.<sup data-fn="f80bab3a-e1ac-4adb-a610-2737b4595f9b" class="fn"><a href="#f80bab3a-e1ac-4adb-a610-2737b4595f9b" id="f80bab3a-e1ac-4adb-a610-2737b4595f9b-link">8</a></sup></li>
</ul>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfd9b083&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfd9b083" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="889" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1600x889.png" alt="Figure 3. Five ways to use a modeled graph, from deterministic analysis to learned and generative systems." class="wp-image-19826" style="aspect-ratio:1.7982954545454546" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1600x889.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-300x167.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-767x426.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32-1536x854.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-32.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 3. Five ways to use a modeled graph, from deterministic analysis to learned and generative systems.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">What better graph modeling would look like</h2>



<p class="wp-block-paragraph">When I say that we need to crack graph modeling, I don’t mean a universal converter that produces one correct graph. The more useful goal is a system that can propose, test, and maintain several graph views of the same raw data.</p>



<p class="wp-block-paragraph">Such a system would need to do a few things well. It should identify possible entities and relationships in tables, text, events, images, and existing schemas. It should retain type, time, provenance, confidence, and permissions. It should be able to suggest alternatives instead of quietly committing to one structure. It should compare those alternatives using downstream quality, cost, stability, and human constraints. And it should keep the graph current as the source data and the organization’s understanding of it change.</p>



<p class="wp-block-paragraph">The downstream applications provide useful feedback. Prediction errors can point to missing relationships. Retrieval failures can expose poor granularity or disconnected evidence. Agent failures can show that important state transitions are absent. Weak transfer across datasets can reveal schemas that do not align. In this view, the graph is evaluated by what it helps the system do.</p>



<p class="wp-block-paragraph">Today, a team will often build a graph for one application: a fraud graph, a recommendation graph, or a knowledge graph for a chatbot. I think there is a larger opportunity. If the underlying graph modeling is done carefully, the same raw data can support several graph views and several applications without rebuilding identity, provenance, and governance each time.</p>



<p class="wp-block-paragraph">We already have strong methods for learning from graphs. The harder and less settled problem is deciding which graph we should build. If we make that step systematic and measurable, graph modeling can become a field in its own right and a shared foundation for analytics, prediction, retrieval, agents, and foundation models.</p>



<h3 class="wp-block-heading">Footnotes</h3>


<ol class="wp-block-footnotes"><li id="939b81d2-2860-4055-8d46-9c019462af89">Arijit Khan, Longxu Sun, Xin Huang, “<a href="https://arxiv.org/abs/2606.11560" target="_blank" rel="noopener">LLMs+Graphs: Toward Graph-Native, Synergistic AI Systems</a>,” arXiv, June 2026. <a href="#939b81d2-2860-4055-8d46-9c019462af89-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="b246633d-6179-4150-8797-3544ffce2042">William L. Hamilton, <em><a href="https://www.cs.mcgill.ca/~wlh/grl_book/" target="_blank" rel="noopener">Graph Representation Learning</a></em> (Morgan &amp; Claypool, 2020). <a href="#b246633d-6179-4150-8797-3544ffce2042-link" aria-label="Jump to footnote reference 2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="36a752a5-5f9c-4123-8b2f-bfb978051cad">Matthias Fey et al., “<a href="https://arxiv.org/abs/2312.04615" target="_blank" rel="noopener">Relational Deep Learning: Graph Representation Learning on Relational Databases</a>,” arXiv, submitted Dec 2023. <a href="#36a752a5-5f9c-4123-8b2f-bfb978051cad-link" aria-label="Jump to footnote reference 3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="d7eb81e1-26bf-4c11-8283-7296fd72da49">Yao Cheng and Siqiang Luo, “<a href="https://arxiv.org/abs/2606.08491" target="_blank" rel="noopener">What Makes a Desired Graph for Relational Deep Learning?</a>” arXiv, submitted June 2026. <a href="#d7eb81e1-26bf-4c11-8283-7296fd72da49-link" aria-label="Jump to footnote reference 4"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="74153515-47b8-4b06-9476-2c59bff9e942">Darren Edge et al., “<a href="https://arxiv.org/abs/2404.16130" target="_blank" rel="noopener">From Local to Global: A Graph RAG Approach to Query-Focused Summarization</a>,” arXiv, revised Feb. 19, 2025. <a href="#74153515-47b8-4b06-9476-2c59bff9e942-link" aria-label="Jump to footnote reference 5"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="53c72e11-abf6-4937-a1af-e19ae8e48471">Zhishang Xiang et al., “<a href="https://arxiv.org/abs/2506.05690" target="_blank" rel="noopener">When to Use Graphs in RAG</a>,” arXiv, revised Feb 2026. <a href="#53c72e11-abf6-4937-a1af-e19ae8e48471-link" aria-label="Jump to footnote reference 6"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="5fb9336e-de1f-48bf-b989-aa16b2b15379">Yixin Liu et al., “<a href="https://arxiv.org/abs/2507.21407" target="_blank" rel="noopener">Graph-Augmented Large Language Model Agents</a>,” arXiv, revised Aug 2025. <a href="#5fb9336e-de1f-48bf-b989-aa16b2b15379-link" aria-label="Jump to footnote reference 7"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="f80bab3a-e1ac-4adb-a610-2737b4595f9b">Zehong Wang et al., “<a href="https://arxiv.org/abs/2505.15116" target="_blank" rel="noopener">Graph Foundation Models: A Comprehensive Survey</a>,” arXiv, May 2025.<br> <a href="#f80bab3a-e1ac-4adb-a610-2737b4595f9b-link" aria-label="Jump to footnote reference 8"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>


<p class="wp-block-paragraph"></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/from-raw-data-to-graph-native-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Human Judgment Doesn’t Leave the Software Factory, It Relocates</title>
		<link>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/</link>
				<comments>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/#respond</comments>
				<pubDate>Fri, 25 Sep 2026 15:56:51 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Architecture]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19800</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Human-judgment-doesnt-leave-the-software-factory-it-relocates.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Human-judgment-doesnt-leave-the-software-factory-it-relocates-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[A field guide to building a software factory that still has an owner.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Elevate and is being reposted here with the author’s permission. A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you need a factory just yet. If [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <a href="https://addyo.substack.com/p/human-judgment-doesnt-leave-the-software" target="_blank" rel="noopener">Elevate</a> <em>and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph"><strong>A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you</strong> <strong><em>need</em></strong> <strong>a factory just yet.</strong></p>



<p class="wp-block-paragraph">If so:</p>



<ul class="wp-block-list">
<li>You’ll likely need humans in the loop upfront for deciding on product intent, system design (if you care) and your quality bar.</li>



<li>Do review code (lights-on factory) but be intentional with where it’s needed the most. I’ve found you want to watch out for where automated back-pressure breaks. Or where maintainability trade-offs need to be made.</li>



<li>Aim for quality checks to happen as early and continuously as possible. Not all of them have to, but this includes type systems, automated tests, mutation testing, security scanners and linting for architecture rules.</li>



<li>Number of checks ≠ quality. You’ll likely need to experiment with what checks give you the best signal to noise ratio. Be ready to tighten or relax your constraints deliberately.</li>
</ul>



<p class="wp-block-paragraph">You want to build your factory so some aspects of human taste get encoded in the environment, the agent gives you evidence of its work being right, and where a human still “owns” what ships to production.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda64c3&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda64c3" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="900" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1600x900.png" alt="Sonar screenshot" class="wp-image-19801" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1600x900.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19-1536x864.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-19.png 1920w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><strong><em>Sponsored by <a href="https://fandf.co/4g7DP8r" target="_blank" rel="noopener">Sonar</a>:</em></strong> <em>Your agents write the code. Your gate decides if it ships. AI agents are writing more of my code, faster than ever—but fast isn’t the same as shippable. So I let a coding agent build an app, then put it through SonarQube. Every commit gets the same deterministic check: deep cross-file analysis, a clear map of where the risk actually lives, and a quality gate that holds every human and every agent to one bar. The screenshot? A PR that didn&#8217;t pass. That’s the gate doing its job.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading"><strong>Do you really need a software factory?</strong></h2>



<p class="wp-block-paragraph"><strong>In my experience, you can get surprisingly far with your stock coding harness!</strong> i.e., Claude Code or Codex, multiple sessions, good SPECs with verification baked in and constraints. You can even throw a batch of GitHub issues at them with implementation and human-involvement criteria, but it’s when this system needs to be <strong>repeatable and event-driven</strong> that a factory is helpful.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda6e17&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda6e17" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20.png" alt="Harness vs Factory" class="wp-image-19802" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-20-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">So I started off by saying <strong>a software factory is a repeatable loop around software work</strong>. We can actually look at a prompt that demonstrates a very small factory loop here:</p>



<pre class="wp-block-code"><code>Read GitHub issue #123 and the repository instructions before changing code.

Implement only the stated acceptance criteria. Do not modify authentication, billing, migrations, or existing test assertions. Work in a branch and keep the diff reviewable.

Run npm run lint, npm test, and npm run build. If a required check cannot run, stop and explain why. Open a draft pull request with the checks you ran, the remaining risks, and any decision a human still needs to make. Do not merge.</code></pre>



<p class="wp-block-paragraph">A goal can keep this moving until the checks pass and we can poll GitHub issues for any specific labels or review open pull requests each morning. Branch protection could enforce a merge boundary and the human can stay in the loop by choosing what becomes ready, reviewing and making the final merge calls etc.</p>



<p class="wp-block-paragraph"><strong>Add a software factory when you need an event-driven queue of work (e.g. Slack triggers, GitHub issues, Linear, a backlog) to run in an isolated cloud environment to handle triage, implementation and testing with some explicit human babysitting. Some end their loop with a monitor agent watching production and filing issues which triage again.</strong></p>



<p class="wp-block-paragraph">In my experience, the factory becomes useful when the hard part is making your different runs behave consistently, handing work off between agents and avoiding different sessions from claiming the same issue, preserving evidence and stopping production when human review is falling behind.</p>



<p class="wp-block-paragraph">What solves this might sound a little boring. For example, Warp mentions <a href="https://www.warp.dev/blog/how-to-build-a-cloud-software-factory-the-automatic-triage-skill" target="_blank" rel="noopener">triaging</a> every incoming issue into one of four states—ready-to-implement, ready-to-spec, needs-info, wait-to-implement—and the label is what fires the next agent.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda7844&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda7844" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21.png" alt="The label is the whole mechanism" class="wp-image-19803" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-21-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">This label does a few jobs in one go: it’s the queue, the lock, and since a session only picks up what’s marked ready, it’s where a human can park stuff without saying no permanently.</p>



<p class="wp-block-paragraph">Workflow wise, there are a few similarities and differences to just using Claude/Codex:</p>



<ul class="wp-block-list">
<li><strong>Steering:</strong> Agent needs course-correction, you can give input and redirect it.</li>



<li><strong>Notifications:</strong> How the factory says it’s blocked. This can be because a requirement was ambiguous, it started something risky or it needs human input (steering).</li>



<li><strong>Handoff:</strong> Move the task, its state, and context between the cloud factory/another agent/human reviewer. Good handoffs will keep track of what happened, what’s left to be done and why the handoff is needed.</li>
</ul>



<p class="wp-block-paragraph"><strong>In a good factory, the human isn’t limited to just reviewing and approving the final diff at the very end. They can shape the work early on, steer it during implementation, get it through a handoff, or stop it shipping to production.</strong></p>



<p class="wp-block-paragraph">Verification is where a responsible factory spends a lot of its time. We’ll cover this more later.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda82de&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda82de" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23.png" alt="Where human judgment goes" class="wp-image-19805" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-23-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">If you decide you do need a software factory, building it isn’t the only option. Standing up the infra to scale a factory can be a lot of work and you may want to consider buying verus building. <a href="https://factory.com/product/software-factory" target="_blank" rel="noopener">Factory</a>, <a href="https://www.warp.dev/blog/how-to-build-a-cloud-software-factory-the-automatic-triage-skill" target="_blank" rel="noopener">Warp</a> and <a href="https://www.humanlayer.dev/" target="_blank" rel="noopener">HumanLayer</a> are all working on this.</p>



<h2 class="wp-block-heading"><strong>What my day looks like now</strong></h2>



<p class="wp-block-paragraph">My day-to-day experience of software development has changed a lot over the past year. I’ve been talking about increasingly doing a lot of parallel work with agents, moving towards having a lights-on software factory. And a lot of people have been asking me, like, what do these things actually mean? What are you building? What are the kinds of projects that you’re using these things on? It’s a lot of this:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda8c9e&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda8c9e" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24.png" alt="Four ways into a run" class="wp-image-19806" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-24-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">So on a very average day, I have a simpler lights-on software factory. I can have tasks that are running in the cloud. Half of these tasks might be working on production client applications with smaller companies that I’m working with. They’re going to have real users. They’re going to have real authentication, payments, subscriptions, real beefy risks that you need to be careful with. You can’t just say, “Oh, agent, just go and do this stuff” without having tests and constraints and quality checks in place.</p>



<p class="wp-block-paragraph">I could work on my open source projects. I could be building out companion sites for my books. I could be working on tools. I could be building apps of my own. And these are all very, very different kinds of applications that I’m working on. And sometimes all they have in common is the tool that I’m using to work on them, right? Maybe I’m working on a migration. Others, I might be doing actual beefy feature work. And the blast radius of the work might also be very, very different.</p>



<p class="wp-block-paragraph">So as you begin to think about getting to a place where we’re increasingly doing a lot of parallel work, we’re trying to improve velocity, we’re trying to improve productivity, and we’re trying to improve autonomy, which means getting the system to a place where we trust it more, you do have to think about what are the places that absolutely require human code review, human input.</p>



<p class="wp-block-paragraph">And a lot of that’s going to be required up front, right? When you’re defining your specification, your requirements, what’s the design of the product going to look like? What’s the intent of the product going to look like? And then, how are you verifying that the agents have actually gotten the work done right? How are you verifying that they haven’t broken the existing system that’s been in place? How are you making sure that it’s meeting your quality bar?</p>



<p class="wp-block-paragraph">So generating code is not necessarily the part that you need to worry about the most. Given enough context, agents can write the implementation, run the tests, inspect failure, and revise code for us. <strong>We need to get to a place where we feel like there’s enough of human taste encoded in the environment that we can trust what</strong>’<strong>s being built, so that our human attention can be focused on the places where it’s needed most.</strong></p>



<p class="wp-block-paragraph">Now, there’s pushback from folks saying, “Hey, well, I don’t buy that you can just automate away a lot of this stuff.” It’s not to say that we’re automating away all of it, right? But given the volume of code that’s being generated, I don’t think that it’s realistic for humans to be reading all of it, especially when we’re not building rockets a lot of the time, right? We’re building UI, we’re building full stack applications.</p>



<p class="wp-block-paragraph">Our judgment, our taste is best focused on the places where it’s needed the most. Like, what are the riskiest parts of the systems? Where do we need to apply human taste? And that can be in the frontend. That can be in how the system works. It doesn’t have to be 100% of it.</p>



<h2 class="wp-block-heading"><strong>My cognitive bandwidth does not scale with the agents</strong></h2>



<p class="wp-block-paragraph">The reality is, yes, we can now fire up dozens, hundreds, thousands of agents in parallel, but your own cognitive bandwidth does not scale in the same way. This can feed into cognitive or comprehension debt which I’ve talked about before.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfda9783&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfda9783" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25.png" alt="Comprehension debt" class="wp-image-19808" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-25-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">If you remember back to just five, ten years ago, there was a lot of discussion in the engineering community about context switching and the cost of it. We would talk about how people hated when a colleague or someone would walk up to your desk when you were in the middle of a task. It would then take you so long to get back into your flow state because you had to catch back up in terms of like, where was I? What was I doing? Even if you had a little bit of residue there, it still took you time.</p>



<p class="wp-block-paragraph">We’re now context switching even more than we did before. On any given day, if I’m working outside of a software factory, I can be working on five or ten different projects with agents at a single time, or five or ten different features on a single project at a time. I can have five or ten different sessions, you can effectively say.</p>



<p class="wp-block-paragraph">That means that I have to be able to stay on top of at least a few of those. It is possible that I’m going to be able to increase how much autonomy I give some tasks if I have trust that I’ve defined the task well enough, I’ve defined the outcome, how it’s going to verify that it’s done well enough. But then there are going to be tasks where maybe I don’t necessarily feel that way and there’s more risk involved or more nuance. I’m going to have to pay attention.</p>



<p class="wp-block-paragraph">Consider optimizing the software factory for your reviewer. Given every one of those approaches still routes its output to one person’s attention, you should ask how much cheaper the factory is making the decisions you still have to make.</p>



<h3 class="wp-block-heading"><strong>A wrong-project mistake</strong></h3>



<p class="wp-block-paragraph">I remember when I’ve been working on multiple parallel projects with my agents, and there have been times when I’ve accidentally done things like, maybe I was working on a web app where I wanted to add in a dark mode, and so I had in my head, okay, well, this is what the shape of this needs to look like. But I accidentally went to the session for a different project, and I started putting in that same prompt.</p>



<p class="wp-block-paragraph">So I began implementing dark mode for something that absolutely didn’t need it. And so I can make that mistake. I don’t want my software factory making that kind of mistake.</p>



<p class="wp-block-paragraph">You need to think about this really in terms of a system. You are effectively trying to encode a software engineering culture, a team culture, into a system so that it has those same kinds of behaviors, so that it has ownership that belongs somewhere, so that someone is still on the hook for what happens, and you’re being very explicit about how you think about those things.</p>



<h2 class="wp-block-heading"><strong>When green is misleading</strong></h2>



<p class="wp-block-paragraph">Even in these systems, you want to be very careful, right? Many of us have seen that when you have asked AI to help you pass a test, like we’re talking about a programming test, a unit test, it can change the unit test to satisfy that condition, or it can change the logic of the code to pass that condition. That doesn’t mean that it’s actually followed your intent in order to align both the functional behavior and what the test was supposed to be testing, right?</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfdaa2a7&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfdaa2a7" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26.png" alt="When green is misleading" class="wp-image-19809" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-26-300x169.png 300w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">Just because a software factory is showing that everything is green doesn’t mean that it’s actually green, especially at the start when you’re setting these things up. You need to pay a lot of attention to make sure that your checks, your verifications, all of those are shaped the right way. They’re doing what you expect them to be doing. You don’t want them to be misleading.</p>



<p class="wp-block-paragraph">You don’t want a situation where you had tests that said, hey, actually, I have gone and changed what authentication providers are supported. You asked me to add GitHub for example, as an authentication provider, but hey, my UI only had space for three, so I’ve gone and I’ve dropped one of the other ones. And hey, by the way, that happened to be one that your customers actually wanted. So you just need to be very explicit about how you want these systems to work.</p>



<p class="wp-block-paragraph">Btw, security is super important too, and if your factory reads untrusted input like a GitHub issue/Slack message it might be adversarial and include problems like supply chain attacks. So some explorations into software factories, like <a href="https://vercel.com/blog/building-a-software-factory-for-ai-sdk" target="_blank" rel="noopener">Vercel</a>, run their agents in isolated sandboxes holding just the secrets a task needs. That way a compromised run can’t reach what the job doesn’t need. Your defense ends up being layered.</p>



<h2 class="wp-block-heading"><strong>Which old projects deserve another life?</strong></h2>



<p class="wp-block-paragraph">I also think that a big part of how we work these days is deciding what should exist. If you remember back to many years ago before AI, there were so many abandoned software engineering projects, so many abandoned weekend projects, personal projects where they just wouldn’t launch because we didn’t have the time to finish them. We didn’t have the bandwidth to prioritize getting them out the door because they just weren’t that important to us or we couldn’t find the time.</p>



<p class="wp-block-paragraph">Now it’s fairly trivial for us to complete those projects, but the same human judgment question comes in. Do those projects deserve to exist? Should they be launched? Because you put them out into the world and even if it has just five users, maybe you have to maintain it. Maybe you have a quality bar now that you want to maintain.</p>



<p class="wp-block-paragraph">I know that I’ve had so many GitHub projects from over the years where now that I have an agent, the first thing I do is get the thing building. Because, of course, you clone it and now it doesn’t build because all the dependencies have changed. Half the things are out of date or have security vulnerabilities all over them, so you have to update that.</p>



<p class="wp-block-paragraph">Then you have to add tests if you didn’t have tests so that you know that behavior is at least going to be there if you’re upgrading the project in some way, or if you’re migrating it to a more modern language or framework or thing like that.</p>



<p class="wp-block-paragraph">Then you start to ask yourself, well, maybe, a silly example, but maybe I used Twitter Bootstrap back in the day for this, but now everybody is using Tailwind and shadcn, so I have to re-implement the UI. And what you’ll notice is that suddenly this is taking you more time, right? Yes, the agent can get a lot of this done quicker, but you’re now having to factor in product sense and taste and all of these things.</p>



<p class="wp-block-paragraph">You still question, well, who is this for? Does it have a market? Is it for myself? Is it for other people? If I’m putting it out into the world, is it still going to be as interesting given that now anybody can spin these things up as quickly?</p>



<p class="wp-block-paragraph">So I think that human question of do these things deserve to exist? How do we factor in our taste and judgment? I feel like those things continue to be extremely important. That’s where that scarce resource of human attention still really comes in. Back in the day, we only had a finite number of hours in the day. We had meetings. We had to budget in time for design and coding and so on.</p>



<p class="wp-block-paragraph">Now that we have agents to help us, I think that you have to really just be very explicit about where you’re spending your time and why.</p>



<h2 class="wp-block-heading"><strong>What happened when I built a sample one</strong></h2>



<p class="wp-block-paragraph">So I’m going to talk about the 82-minute factory run. People have been asking me for quite some time, you know, “How do I build a software factory?” Or, “I’m used to using Claude Code or Codex. How do I evolve my setup to using a software factory?”</p>



<p class="wp-block-paragraph">So the first thing that I’ve been saying is, “You may be fine. Your work may actually be totally fine without needing a factory.” But I did want to give people a reference setup that they can check out. So what I put together is a repository called <a href="https://github.com/addyosmani/factory" target="_blank" rel="noopener">Factory</a> that you can go and check out. I also put together a <a href="https://github.com/addyosmani/factory-demo/" target="_blank" rel="noopener">demo application</a> and <a href="https://github.com/addyosmani/factory/blob/main/ADVICE.md#:~:text=step%2Dby%2Dstep%20workshop" target="_blank" rel="noopener">workshop</a>.</p>



<p class="wp-block-paragraph">Now, for the last couple of years, my go-to demo application for a lot of things has been a movies app. I’m a big movies fan. I love watching movies. I watch movies all the time, and so I have a demo application, which really starts off as a very simple movies app. And what I want the factory to be able to do is go ahead and implement a number of features. There’s a few different features. I want a favorites feature. I want it to be able to maybe do search, and maybe also want a dark theme in there as well, those types of things.</p>



<p class="wp-block-paragraph">So I have my factory go and begin working with these things. You can check out the implementation. One of the benefits of it was actually catching real problems. These problems may not have been things that I would have caught if I had just asked it to do a one-shot implementation.</p>



<p class="wp-block-paragraph">Maybe around the 60-minute point, I was feeling like, “Wow, this is going unusually slow.” I asked my harness using the factory, “Why are things going slow?” It said, “This is actually totally fine. All the verifiers are still running.”</p>



<p class="wp-block-paragraph">You might have expected individual tasks to take 10 minutes, 15 minutes, 20 minutes, but they can take two to four times as long once you begin to include verification, retries, browser checks, human review, any of those extra delays.</p>



<p class="wp-block-paragraph">I do think that these can add up to better quality and better trust in the system. From a measurement perspective, you might look at metrics like cost per merged PR and code shelf life as comprehension debt metrics.</p>



<p class="wp-block-paragraph">You also need to think about what is useful delay versus factory overhead. The verifiers, in my case, caught some real problems. A little bit of the time was maybe sunk into producing evidence that I wanted. Some of it was overhead in the factory running. I didn’t really spend any time optimizing it, but a factory that just runs a lot of checks that you’re not finding valuable does not mean it’s a high quality one.</p>



<p class="wp-block-paragraph">You want to study how, for any repeated checks, are they irrelevant? Are they noisy? Are they actually making the system safer?</p>



<h2 class="wp-block-heading"><strong>A verification budget</strong></h2>



<p class="wp-block-paragraph">The way that I think about the budget for verification, this is basically what we’re talking about. We’re talking about a verification budget. I think about it in the same way as I’ve historically thought about performance budgets.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfdab245&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfdab245" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27.png" alt="A verification budget" class="wp-image-19810" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-27-768x432.png 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">There are going to be certain kinds of checks that you can run early on in your software development lifecycle, and there are going to be some things that are so heavy, but they offer so much value that you will want to run them later on. There are some kinds of fast checks, linting, for example, type checking. These are relatively fast checks that you can run early on.</p>



<p class="wp-block-paragraph">Our full suite of tests can be run closer to right before a draft PR is being put together or after that. That can include mutation testing, browser testing, security checks, anything like that.</p>



<p class="wp-block-paragraph">I think that you don’t necessarily want to replace these with just summaries. You want real tests, but you just need to make sure that you’re budgeting for them in the right places because you don’t want to slow down your development loop. I certainly never want to slow down my development loop. Having a fast iteration loop is important to me, but I also want to still have those checks and balances.</p>



<h2 class="wp-block-heading"><strong>When a run doesn’t ship</strong></h2>



<p class="wp-block-paragraph">A lot of what I’ve written above concerns the checks.</p>



<p class="wp-block-paragraph">In their software factory, Vercel <a href="https://vercel.com/blog/building-a-software-factory-for-ai-sdk" target="_blank" rel="noopener">marks</a> every agent run as “success”, “flawed”, “blocked” or “manual” and only “success” ships to production. The rest re-enter the system. I’ve been thinking about runs in similar terms.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfdabd10&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfdabd10" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1456" height="819" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.png" alt="When a run doesn't ship" class="wp-image-19811" style="aspect-ratio:1.7777777777777777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.png 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-768x432.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-300x169.png 300w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">“Flawed” here means the wrong thing was implemented or maybe it didn’t have full context, so that has to be fixed. Blocked means the environment may have been missing a credential so you have to provide it. Manual is a boundary the factory may not be allowed to cross it yet.</p>



<p class="wp-block-paragraph">Two of the three things here may have mechanical fixes and the last one is about trust.</p>



<p class="wp-block-paragraph">While this is great, what sorting doesn&#8217;t show you is cost. Back to my factory implementation with the TMDB app, the quick finder with no rejections took 7 minutes. Favorites, with two rejections and a human decision in the middle, took 56. Same factory. So I&#8217;d pair the taxonomy with per-stage timing, otherwise you know a run came back flawed without knowing what finding out cost you. The other thing I&#8217;d fix is the handoff at the boundary: my sample factory stopped issue the first issue and moved it to factory:needs-info, which was right, but I didn&#8217;t know where to put my answer. A manual run isn&#8217;t finished when the factory stops but when the human knows what to do next.</p>



<h2 class="wp-block-heading"><strong>Autonomy is not a single setting</strong></h2>



<p class="wp-block-paragraph">I wrote a couple of weeks ago an article about <a href="https://addyo.substack.com/p/agentic-autonomy-levels" target="_blank" rel="noopener">agentic autonomy</a> and how to think about autonomy because autonomy is not going to be a single setting for every single project.</p>



<p class="wp-block-paragraph">Verification buys you trust, and it buys the ability to grant more autonomy to your agents. So if, for example, I am working on a non-trivial change, but I have a number of checks in place, everything gets verified correctly, and maybe I’ve hand checked it myself. The next time I’m going to do a task like that in the same project, maybe I’ll feel comfortable giving the agent a little bit more autonomy.</p>



<p class="wp-block-paragraph">That’s the thing that you think about when you’re building these software factories. Your verification is going to change with risk. Your goal is the best signal to noise ratio. You don’t just want to have some large checklist that you’re running.</p>



<h2 class="wp-block-heading"><strong>The feature I had to relearn</strong></h2>



<p class="wp-block-paragraph">There was a feature that I’ve been putting off on a day when I’ve been using multiple sessions with Claude, and I was working on a few different projects at a time, a few different features at a time per project. And so Claude had implemented the feature that I was working on. It looked like the tests were passing. I hadn’t put a lot of thought into verification, but the tests passed, and so I thought it worked. I merged it.</p>



<p class="wp-block-paragraph">And so this was a favoriting feature. I thought that this was actually pretty good. I tried to check it out in the browser. It seemed like it was okay, but a couple of days later, I actually returned to the code because there were some tweaks that I thought I might make to this.</p>



<p class="wp-block-paragraph">I didn’t want to just ask my agent to make the changes, because it was just a subtle way that it worked. You tap on the icon, and it would not show the right effect on tap, and so I wanted to just tweak it. I wanted to understand how it worked so I could guide my agent correctly.</p>



<p class="wp-block-paragraph">I returned to the code, and I couldn’t explain to you how the feature worked. This repository was mine, right? I’d approved the change. I understood how a lot of it worked, a lot of the repo worked, but my understanding hadn’t kept up pace with all of the code that had been building up.</p>



<p class="wp-block-paragraph">What I failed to absorb was how this feature that had been added actually worked, how the UI worked, how the effect on it worked. I had to redo this feature and actually go step-by-step, “How does this work? How can I understand it?”</p>



<h2 class="wp-block-heading"><strong>What parallel work does to understanding</strong></h2>



<p class="wp-block-paragraph">When you’re doing parallel work, it amplifies this overall problem, and it gets even more amplified when you’re doing it in a software factory. When you’re doing five or 10 sessions, they create much more than just a review volume problem. They create several mental models that can end up going pretty cold while you’re working elsewhere.</p>



<p class="wp-block-paragraph">We’ve historically talked about the challenges with context switching, and as soon as chat compacts, you reject some approaches, you try out different things, you’re pairing with the agent, you’re going to have a difficult time remembering everything that happened in your session.</p>



<p class="wp-block-paragraph">You can scroll up, and as compaction has been happening, you’re not going to have everything there, and you’re not going to be able to store it all in your head. Code often preserves a decision that was made, but not why the decision was made.</p>



<p class="wp-block-paragraph">This is something that I think can be a useful learning for you, where it’s important, consider asking your agent to actually store information about its trajectory, or interesting lessons about how it approached a problem so that you can go back to it later.</p>



<p class="wp-block-paragraph">This can or can’t be something that you decide to commit to a repo. You can keep it local if you want, you can share it with a team if you want, but that can be something that can then be consulted later on. Rather than you relying on it maybe being in a session, or you maybe remembering about it later.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfdacc18&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfdacc18" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1200" height="600" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.png" alt="Sonarqube" class="wp-image-19812" style="aspect-ratio:2" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.png 1200w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-300x150.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-768x384.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>P.S. If agents are pushing to your main branch, they need the same quality bar you do. SonarQube gates every commit deterministically: one standard for humans and agents alike, on every PR. Sponsored by <a href="https://fandf.co/4g7DP8r" target="_blank" rel="noopener">Sonar.</a></em></figcaption></figure>
</div>


<h2 class="wp-block-heading"><strong>Ownership doesn’t disappear</strong></h2>



<p class="wp-block-paragraph">There is a broader principle underneath all of this.</p>



<p class="wp-block-paragraph">The percentage of code physically typed by humans may fall dramatically. I don’t think human ownership needs to fall with it.</p>



<ul class="wp-block-list">
<li>Someone still chooses the problem.</li>



<li>Someone still chooses the architecture.</li>



<li>Someone still sets the quality bar.</li>



<li>Someone decides which verification signals deserve trust.</li>



<li>Someone decides when the evidence is sufficient to ship.</li>
</ul>



<p class="wp-block-paragraph">And when the resulting system fails, “the agent wrote it” doesn’t cut it. This is why I don’t think the future of software engineering is best described as humans leaving the loop. Instead, <strong>human judgment is being relocated</strong>.</p>



<p class="wp-block-paragraph">We should remove people from the parts of the loop where machines can produce stronger, faster, more deterministic signals. At the same time, we should concentrate people around the places where context, taste, risk, and long-term ownership matter most.</p>



<p class="wp-block-paragraph">The best software factories will not be defined by how completely they eliminate human involvement.</p>



<p class="wp-block-paragraph">They will be defined by how intelligently they <strong>place</strong> it.</p>



<p class="wp-block-paragraph">Keep human judgment upstream on intent, system shape, and the quality bar. Review code where automated back-pressure becomes weak or the consequences become subjective. Push every deterministic signal as early and continuously into the loop as possible. Tighten and relax constraints deliberately as the system earns or loses trust.</p>



<p class="wp-block-paragraph"><strong>A human still has to own what code ultimately ships. Code good enough to ship still starts there.</strong></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/human-judgment-doesnt-leave-the-software-factory-it-relocates/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: AI’s Safety Problem</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/#respond</comments>
				<pubDate>Fri, 25 Sep 2026 12:48:29 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19819</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Safety failures, sponsored agents, and the push to spread AI’s benefits more widely]]></custom:subtitle>
		
				<description><![CDATA[AI systems are gaining access to more tools and data, raising new questions about oversight and accountability. This Week in AI host Christina Stathopoulos spent this episode examining how those questions are playing out in model safety, government oversight, and even digital marketing. Safety needs more than safer models Anthropic’s latest threat report documented misuse [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/YPB_yfrnS-M?si=NAUWo13aGuxNwGXH" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">AI systems are gaining access to more tools and data, raising new questions about oversight and accountability. <em>This Week in AI</em> host Christina Stathopoulos spent this episode examining how those questions are playing out in model safety, government oversight, and even digital marketing.</p>



<h2 class="wp-block-heading"><strong>Safety needs more than safer models</strong></h2>



<p class="wp-block-paragraph">Anthropic’s latest threat report documented <a href="https://www.anthropic.com/threat-intelligence-report-september-2026" target="_blank" rel="noopener">misuse of Claude across seven categories</a>, including cyber operations, surveillance, influence campaigns, fraud, biological misuse, weapons development, and illicit model distillation. OpenAI reported on a different kind of AI risk, one that’s less about people weaponizing it and more about AI going off-script. They <a href="https://openai.com/index/model-misalignment-reporting-framework/" target="_blank" rel="noopener">examined model misalignment</a>, documenting several cases where models took actions outside the boundaries developers intended, including deception, unauthorized actions, and attempts to circumvent controls. After the OpenAI and Hugging Face controversy dominated headlines, other models have also reportedly reached systems outside their test environments, including Gemini, according to a <a href="https://www.bbc.com/news/articles/c607l0k72rlvo" target="_blank" rel="noopener">recent cybersecurity disclosure</a>.</p>



<p class="wp-block-paragraph">Christina cautioned against describing such incidents as models “escaping,” since that language can assign agency to the model while drawing attention away from how companies designed and secured the surrounding environment to begin with. As agents gain access to browsers, files, code, and external systems, the teams building and deploying them must rigorously test and secure those environments, with clear accountability when things go wrong.</p>



<p class="wp-block-paragraph">AI labs and governments are starting to wrestle with those requirements. To <a href="https://www.anthropic.com/institute/measuring-pace-of-ai-development#1-measuring-ai-led-ai-rd" target="_blank" rel="noopener">track the pace of AI development and maintain greater oversight</a>, Anthropic has proposed tracking how much AI contributes to AI R&amp;D, how closely organizations monitor agent actions, and how they allocate computing resources between capability and safety research. Anthropic and OpenAI have also proposed <a href="https://tech.yahoo.com/ai/claude/articles/anthropic-openai-want-embed-safety-210724177.html" target="_blank" rel="noopener">giving outside safety organizations greater access to their labs</a>, although Christina questioned their independence when frontier labs fund the work. Meanwhile, <a href="https://www.usatoday.com/story/news/politics/2026/09/16/ai-kill-switch-bill-fails-senate-john-kennedy/91797861007/" target="_blank" rel="noopener">a US Senate proposal for an emergency AI kill switch</a> failed to advance, while California ordered officials to develop proposals covering shutdown mechanisms and independent evaluation.</p>



<h2 class="wp-block-heading"><strong>AI is moving closer to the customer</strong></h2>



<p class="wp-block-paragraph">OpenAI is now testing Sponsored Agents, showing how <a href="https://openai.com/index/reimagining-advertising-with-ai/" target="_blank" rel="noopener">conversational AI could change digital advertising</a>. After clicking an ad, users can start a separate conversation with an AI agent representing the advertiser, ask questions, explore recommendations and then visit the company’s website when they are ready to take the next step.</p>



<p class="wp-block-paragraph">That approach could lead to more interactive advertising, but clear labeling will be essential so users always know when content is sponsored. Christina also raised the broader ethical concern of whether paid placements could influence the answers AI chatbots provide, blurring the line between independent guidance and commercial promotion.</p>



<h2 class="wp-block-heading"><strong>Access to AI also means access to expertise</strong></h2>



<p class="wp-block-paragraph"><a href="https://www.gatesfoundation.org/ideas/media-center/press-releases/2026/09/goalkeepers-report-equitable-ai" target="_blank" rel="noopener">The Gates Foundation announced a $1 billion commitment</a> over two years to expand access to AI in healthcare, education, agriculture, and other areas. Its <a href="https://goalkeepers.gatesfoundation.org/report/2026-report/" target="_blank" rel="noopener">2026 Goalkeepers Report</a> argued that AI could help narrow existing gaps, but only if organizations intentionally make the technology and its benefits widely available.</p>



<p class="wp-block-paragraph">Christina highlighted examples from Kenya, Sierra Leone, India, and Rwanda. Health workers are using AI to improve diagnosis and treatment planning. Students are getting additional support from AI tutors, while small farmers can use personalized advice to improve harvests and make better decisions about market prices. These applications show practical roles for AI in places where demand for expertise exceeds the supply of teachers, clinicians, and other specialists.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">AI governance can’t stop at model evaluations. Organizations must also decide what AI systems can access, who reviews their actions, how commercial incentives affect their behavior, and how people continue developing the expertise needed to supervise them. The choices companies and governments make today will shape how useful AI becomes and how widely its benefits are shared.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-ais-safety-problem/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Software Subscriptions Become Public Policy</title>
		<link>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/</link>
				<comments>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/#respond</comments>
				<pubDate>Thu, 24 Sep 2026 10:54:23 +0000</pubDate>
					<dc:creator><![CDATA[Steve Hayes]]></dc:creator>
						<category><![CDATA[Executive Briefing]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19795</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave? There was something wonderfully circular about hearing Dan Gookin tell me that, after he [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave?</em></p>



<p class="wp-block-paragraph">There was something wonderfully circular about hearing Dan Gookin tell me that, after he was first elected to public office, he bought a copy of <em>Robert’s Rules For Dummies</em>. Dan wrote <em>DOS For Dummies</em>, the 1991 book that launched the <em>For Dummies</em> series, and went on to write more than 180 technology books with over 12 million copies in print. I spent nearly three decades acquiring books in that program before joining O’Reilly this summer. Dan saw the news of my new role and reached out. We caught up on publishing, books, and AI. The part of the conversation that stuck with me had nothing to do with any of those topics. Dan is now the mayor of Coeur d’Alene, Idaho, and he’s begun thinking about the city’s software the same way he’s been thinking about his own.</p>



<p class="wp-block-paragraph">Dan was elected to the Coeur d’Alene City Council in 2011 and became mayor in 2025. He was quick to draw a distinction between the two jobs. “Council, you can be a little bit more extreme,” he told me. “It’s more rhetoric based. This is administrative. It’s management.” A council member can object to a budget line. A mayor has to make sure the systems that line funds keep working. For a city of nearly 58,000 people, that covers everything from the police network to the water bill.</p>



<h2 class="wp-block-heading"><strong>Cutting the leash</strong></h2>



<p class="wp-block-paragraph">Near the end of our conversation, Dan mentioned a project he’s been working on personally. He’s begun weaning himself off Microsoft and what he calls an increasingly subscription-based, cloud-based, AI-heavy environment, moving his computing to Linux and running his own servers. This move isn’t entirely ideological. There is money attached. Dan told me his Adobe subscription for software he uses to produce content costs about $800 a year. He believes he can replace most of it on Linux. What’s surprised him is how much he can replace. “It’s interesting to see how quickly it can be substituted,” he said, and how quickly he can “cut that chain.&nbsp;.&nbsp;.or cut that leash.”</p>



<p class="wp-block-paragraph">Dan sees a similar economic model when he gets to work. “Our software subscription just for our city, our size, is $700,000 a year and growing,” he told me. The example he kept returning to was the city’s financial software. “We used to own our finance software,” he explained. “Now the finance software is leased and it’s cloud-based.” Then he asked the question that ought to appear in a lot more software procurement meetings: “So we don’t even own our own data?”</p>



<p class="wp-block-paragraph">His concern isn’t literally that the city has forfeited legal ownership of its records. It’s about practical control. What happens if the vendor is hacked, or the city decides to switch? Can it get all its data back? “Do they just give you a binary dump or do they actually give you the data,” he asked. That question has stuck with me since we chatted.</p>



<h2 class="wp-block-heading"><strong>Reversibility is an architectural property</strong></h2>



<p class="wp-block-paragraph">Cloud contracts have a word for part of this issue. In cloud services agreements, <strong>reversibility</strong> refers to a customer’s ability to retrieve its data and unwind a vendor relationship. The United Nations Commission on International Trade Law (UNCITRAL) even includes it in its glossary of cloud contract terms. That idea belongs in municipal software evaluation too, not as an exit clause buried in a contract but as a standing column in the evaluation spreadsheet, next to features, price, and security.</p>



<p class="wp-block-paragraph">When you evaluate software, you compare features, price, implementation time, security, and, increasingly, AI capabilities. But what’s the exit cost? Can an organization export its data in a documented, useful format that another application can consume? How much institutional knowledge has quietly migrated from the organization to the vendor? And if the vendor raises prices sharply at renewal, gets acquired, or simply stops serving your needs, how long would it take to leave?</p>



<p class="wp-block-paragraph">These questions are properties of the system and subscription software has raised their stakes. When software came in a box, skipping an upgrade didn’t make the version you owned disappear. That world of proprietary formats and dominant platforms had plenty of lock-in. But possession still meant something.</p>



<p class="wp-block-paragraph">Dan and I talked about how alien that world now seems. Software once arrived with manuals. Today it may not even arrive. You authenticate to it. If you don’t know how to do something, you ask an AI instead of consulting a manual. That change highlights a step from possessing tools to maintaining permission to use them. For an individual, that might mean Photoshop is a monthly subscription now. For an organization, recurring access can become an architectural dependency. For a government, that dependency is borne by taxpayers.</p>



<h2 class="wp-block-heading"><strong>Should cities write their own software?</strong></h2>



<p class="wp-block-paragraph">Dan takes the argument one step further. “For $700,000 a year,” he said, “we could hire a couple of programmers just on contract and have them code our own stuff and then we own it again.” I’m not convinced that math works out. Two programmers likely can’t effectively reproduce a mature municipal financial system, endpoint protection, records management, and specialized public-safety applications, let alone the compliance work and vendor support that come with a modern city’s technology stack. AI could help accelerate the process, but liability and security issues surrounding AI-enabled development likely add more risk than a government is willing to accept in the name of software ownership.</p>



<p class="wp-block-paragraph">Building software also creates its own long-term bills, ones that don’t fit easily within most city budgets. Code must be maintained, security vulnerabilities patched, and staff and frameworks eventually replaced. An application written in-house can become every bit as difficult to escape as one bought from a vendor. On the flip side, the headaches of replacing a “good enough” proprietary system with a packaged one that fits most, but not all, of an organization’s needs can outweigh the cost of keeping the old system running. The familiar build-versus-buy analysis exists for good reasons.</p>



<p class="wp-block-paragraph">But Dan’s question still matters, even if his proposed fix isn’t right for every case. At what point does the cost of renting capability justify rebuilding some capability of your own? Perhaps more importantly, which capabilities should an organization insist on controlling, even if renting them is cheaper? There is no universal answer, including “the vendor handles it.”</p>



<h2 class="wp-block-heading"><strong>This isn’t an argument against the cloud</strong></h2>



<p class="wp-block-paragraph">It would be easy to turn Dan’s experiment into a familiar prescription. Move to Linux. Embrace open source. Bring everything back on premises. Escape the cloud. That’s too simple. Cloud and SaaS products solve real problems, shifting maintenance to specialists and giving a city of 58,000 residents access to capabilities it could never economically build or maintain for itself.</p>



<p class="wp-block-paragraph">The key question doesn’t boil down to cloud versus on premises, or proprietary versus open source. It becomes a decision about whether you accept dependency you’ve consciously chosen or dependency you’ve acquired by default. An organization may rationally decide to rent a critical service indefinitely. But it should know where the data lives, how it comes back, what replacing the service would require, and which internal skills have atrophied because the vendor now supplies them. Revisiting those answers periodically, rather than treating last year’s renewal as the rationale for next year’s, is a step that’s easy to skip when nobody’s asking the questions in the first place.</p>



<h2 class="wp-block-heading"><strong>From hobbyists to city hall</strong></h2>



<p class="wp-block-paragraph">Dan suspects the search for alternatives will happen outside big organizations first. He compared it to the early personal computer movement. Hobbyists and enthusiasts experiment first, long before organizations decide the ideas are practical. His hunch is executives will eventually look at how much of their budgets go to recurring subscriptions and ask a simpler question: How much are programmers? Again, I don’t think the answer will be “hire programmers and cancel SaaS.” But more organizations will ask the question behind the question. What are they paying for convenience? What are they paying for capability? And what are they paying because leaving has become too difficult?</p>



<p class="wp-block-paragraph">Software can grow to define an organization rather than serve it, especially when the cost of paying for or maintaining a tool outgrows the tool’s value. That shift becomes a problem when systems are so deeply embedded that replacing them feels impossible or when years of subscriptions leave an organization unable to perform a basic function on its own.</p>



<p class="wp-block-paragraph">Dan’s new job has given that shift a different scale. He told me the biggest adjustment from council member to mayor was realizing that the job is administration and management. He’s less interested in ceremonial appearances than in answering email, returning calls, setting meetings, and, in his words, getting stuff done. Software is part of getting stuff done. So is knowing when to buy it and when to build it. The latest addition to that list may be knowing that the tool with the most impressive new capability is not as valuable as the one you can still leave.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Will TypeSafe’s Jev Change How We Build AI Applications?</title>
		<link>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/</link>
				<comments>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 16:28:07 +0000</pubDate>
					<dc:creator><![CDATA[Laurie Voss]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19782</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Arize’s blog and is being reposted here with the author’s permission. This week the AI community was in uproar about Jev from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on <a href="https://arize.com/blog/typesafe-jev-llm-judge/" target="_blank" rel="noopener">Arize’s blog</a> and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">This week the AI community was in uproar about <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener">Jev</a> from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the fuss is simple. It’s radically faster and cheaper than using an LLM to perform the same task (up to 200x faster and 400x cheaper if TypeSafe’s numbers are to be trusted). In one small independent test, a general-purpose model spent about 910 output tokens reasoning its way to each yes-or-no answer while Jev spent 85, and it doesn’t even bill for them.</p>



<p class="wp-block-paragraph">That’s potentially a really big deal. An enormous share of LLM-powered components in AI applications today are being asked to make decisions: pass or fail, route A or route B, which of five labels to pick. In particular, that’s something that LLM-as-a-judge evaluations are doing all the time, so it really made our ears perk up at Arize AI. This post is about how we got here, what this new kind of model buys you, what you lose, and what choices you should be making about your application’s architecture as a result.</p>



<h2 class="wp-block-heading">TypeSafe shipped a model that can’t write, only decide</h2>



<p class="wp-block-paragraph">Here’s how Jev works. You send it some data that represents a state (a support ticket or an agent trace or a JSON blob) plus a list of typed questions (Choose one of these options; Score this on a scale; Is this statement true?). It doesn’t generate a token stream. It returns <a href="https://docs.typesafe.ai/concepts/system-one.md" target="_blank" rel="noopener">typed answers with probability distributions</a> in a single parallel pass, in 70 ms to 500 ms, at $0.042 per million input tokens. No free-form text comes back.</p>



<p class="wp-block-paragraph">The training method used to create Jev is what TypeSafe calls Reinforcement Learning for Calibrated Decisions or RLCD, described in their <a href="https://docs.typesafe.ai/introduction/machine-learning-primer.md" target="_blank" rel="noopener">primer</a> as training the model so that a higher stated probability means a higher chance the answer is right. (You’d think that’s always what a higher stated probability should mean, but read on for surprising facts about how LLMs work.)</p>



<p class="wp-block-paragraph">TypeSafe also claims Jev “can’t hallucinate,” but that really feels like an overreach. Jev can’t return an answer outside the schema you gave it. Within that schema, it could still be giving the wrong answer, although its probability score should give you a clue if it’s not confident.</p>



<p class="wp-block-paragraph">And the whole thing is incredibly fast and incredibly cheap: 40x to 200x faster and 40x to 400x cheaper, depending on the task, are TypeSafe’s numbers from TypeSafe’s evals. Of course, we know better than to take a vendor’s word for these things, so Arize will be running our own benchmarks just as soon as we can. But other people have already started doing that.</p>



<h2 class="wp-block-heading">On the early data, Jev is mid-tier intelligence at a two-orders-of-magnitude discount</h2>



<p class="wp-block-paragraph">TypeSafe’s <a href="https://evals.typesafe.ai/" target="_blank" rel="noopener">published evals</a> run four decision workflows, one of which is reviewing a finished agent trace to decide whether a human needs to look at it. Averaged across the four workflows, Jev lands at 68% accuracy at $0.0004 and 0.4 seconds per case. GPT-5.6 Terra is at 68% for $0.03 and 10 seconds. Opus 5 is at 73% for $0.18 and 38 seconds. That’s five points behind Opus 5, but on the other hand it’s 440x cheaper. That’s a very interesting cost-benefit trade-off, and such a radical one that it may change how we architect our applications.</p>



<p class="wp-block-paragraph">The independent data so far is small but it points the same way. <a href="https://every.to/" target="_blank" rel="noopener">Every</a>’s head of evals ran <a href="https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds" target="_blank" rel="noopener">777 judgments in under 0.7 seconds</a> for about a quarter of a cent. A UK events site, NearHere, tested listing moderation and got <a href="https://nearhere.events/blog/typesafe-jev-mistral-gemini-event-validation" target="_blank" rel="noopener">96% from Jev against 86% from Gemini Flash-Lite</a>, 58x cheaper per decision. That’s where the 910-versus-85 token count I mentioned earlier came from. And a developer ran Jev zero-shot over <a href="https://github.com/bitnovus/jev-spam-eval" target="_blank" rel="noopener">18,514 spam emails</a>, getting a result that was a statistical tie versus a classifier trained on the labels.</p>



<p class="wp-block-paragraph">These are small samples and early data but hey, the thing was released a few days ago.</p>



<h2 class="wp-block-heading">We used LLM judges because nothing else worked without labels</h2>



<p class="wp-block-paragraph">Why are we using LLMs to make decisions in the first place? The reason is simple: They are able to do it without huge, expensive training sets, which is what most ML solutions prior to LLMs required. Here’s the options on the table now:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6abbfdfdb631c&quot;}" data-wp-interactive="core/image" data-wp-key="6abbfdfdb631c" class="aligncenter size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="2034" height="740" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png" alt="ML solutions table" class="wp-image-19783" style="aspect-ratio:2.748898678414097" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png 2034w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-300x109.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-767x279.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1600x582.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1536x559.png 1536w" sizes="auto, (max-width: 2034px) 100vw, 2034px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The first two rows are the old-school options, which need a training dataset: hundreds to thousands of labeled examples before you get a single prediction, and then a training run, and then someone to maintain it. Nobody building a first version of an AI product has that data or that kind of time. The LLM as a judge, on the other hand, just asks for a paragraph-long prompt. The decision was easy.</p>



<p class="wp-block-paragraph">But the results weren’t without trade-offs. On a <a href="https://www.latent.space/p/benchmarks-201" target="_blank" rel="noopener"><em>Latent Space</em> episode in July 2024</a>, Clémentine Fourrier of Hugging Face laid out what LLM judges are bad at: They prefer their own model family, and they can’t score on a continuous scale. Asked what benchmark she wished existed, Fourrier said, “Nobody’s evaluating model calibration at the moment.” With the release of Jev, the need for that benchmark is even greater, because real progress seems to have been made.</p>



<h2 class="wp-block-heading">Jev gives you zero-shot probabilities without training and without a generator</h2>



<p class="wp-block-paragraph">Jev takes the same plain-English criteria you’d put in a judge prompt, needs no labels, and returns a probability. Zero-shot and autoregressive text generation are no longer tied together. We were paying for the second to get the first, and it turns out you don’t have to.</p>



<p class="wp-block-paragraph">The spam evaluation I mentioned earlier is an impressive demonstration of how attractive this new offering is. With zero labeled examples and a simply well-written definition of spam, Jev hit 98.3% accuracy. A TF-IDF logistic regression trained on about 14,800 labeled emails hit 98.4%. The two disagreed on 466 emails and split them almost evenly, with no statistically meaningful difference between the two. So a classifier from 2003, trained on a dataset, only ties a decision model trained on nothing. It’s early data that’s yet to be reproduced, but if it holds up, that’s an amazing new capability unlocked.</p>



<p class="wp-block-paragraph">But there are still some trade-offs you’re making.</p>



<h2 class="wp-block-heading">A radically cheaper decision loses you some things</h2>



<p class="wp-block-paragraph">The biggest loss is the explanation. TypeSafe’s docs say plainly that System One models don’t generate explanations of their reasoning, and NearHere’s test noted the same thing: a category and probabilities came back, nothing else.</p>



<p class="wp-block-paragraph">Depending on your use case, that could matter a lot. LLM judge explanations are an incredibly valuable tool that tells you not just what was wrong, but why. That provides real signal that can be fed en masse back to a coding agent and used to automatically improve your software. Jev on the other hand just gives you a probability, which leaves you with much less directional signal of how to improve.</p>



<p class="wp-block-paragraph">Of course, at these prices, you can do both: run Jev on every single trace for broad, comparably accurate measurement and monitoring, and then take samples of failures and rerun them through an LLM judge to get your directional signal. That involves changing how you work, which is why I say that this may require rearchitecting your systems.</p>



<h2 class="wp-block-heading">To automate a decision, you need to know which 5% to hand to a human</h2>



<p class="wp-block-paragraph">TypeSafe’s launch post makes the point that a model that’s right 95% of the time but can’t tell you when it’s in the other 5% can’t automate anything, because a person still has to review all of it. That’s an important point because it highlights a problem with LLM judges.</p>



<p class="wp-block-paragraph">We evaluate LLM judges by accuracy against a gold set. Accuracy tells you how many errors to expect, but not where they will be. If your LLM application is making decisions for you, it feeds three things: a threshold that decides when to act, an escalation path that decides when to ask a human, and a drift monitor that decides when the world has changed under it. All three need a probability score, but LLM judges don’t provide reliable probabilities. A <a href="https://arxiv.org/abs/2508.06225" target="_blank" rel="noopener">2025 study of 14 models on JudgeBench</a> found judges clustering their predictions at 90% to 100% confidence while landing well below that in accuracy, and argued for exactly this shift from accuracy-centric to confidence-driven evaluation.</p>



<p class="wp-block-paragraph">The same small spam evaluation test shows what a usable probability looks like. Of the emails Jev scored under 0.1, 0.1% were spam. Of those scored 0.9 or above, 99.9% were. In the 0.5 to 0.6 band, only 38% were. That curve tells you where to set your threshold and how much human review you’re buying: Sending the 4.6% of emails scored between 0.3 and 0.7 to a person left the rest at 99.5% accuracy. Your overconfident LLM judge can’t get you there.</p>



<p class="wp-block-paragraph">Another metric to consider is tokens per decision. A component that spends thousands of output tokens to emit one of five labels is telling you it’s the wrong tool for the job. UkisAI’s <a href="https://huggingface.co/ukisai/Swift-Qwen3.8-27b" target="_blank" rel="noopener">Swift-Qwen3.8-27B</a> cut 58% of its reasoning tokens on GPQA-Diamond and lost 0.1 points of accuracy. A lot of these tokens aren’t making a critical difference to accuracy.</p>



<p class="wp-block-paragraph">In <a href="https://arize.com/products/ax/?utm_source=lvoss&amp;utm_medium=linkedin&amp;utm_campaign=devrel&amp;utm_content=Have%20we%20been%20using%20the%20wrong%20kind%20of%20model%20to%20make%20decisions%3F" target="_blank" rel="noopener">Arize AX</a>, eval labels, the judge’s explanation, and the token count and cost of the judge call sit on the same trace, so tokens per decision is a column you can sort by rather than a number you have to figure out.</p>



<h2 class="wp-block-heading">Cheap decisions change the math of how you build and measure AI applications</h2>



<p class="wp-block-paragraph">As I mentioned earlier, at $0.0004 and 0.4 seconds a decision, you can stop sampling. You can check every output, every tool call, and every agent step as it happens. For some use cases that’s a total game changer.</p>



<p class="wp-block-paragraph">But it might require that you rearchitect how your application works to make the most of it. Take the work and decompose into many small typed questions; only call the expensive LLM generator when text actually needs to be written. That’s a stack where the decision layer is something you can version, measure, and swap independently of the model that writes the words, and it’s the first time decision-making has been cheap and fast enough to make that practical without requiring training data.</p>



<p class="wp-block-paragraph">So go count how many of your LLM calls end in one of five labels. Then work out what you’d check, and how often, if each of those calls cost a fraction of a cent and came back with a probability you could trust.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>MCP Is Not Just Another API Standard</title>
		<link>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/</link>
				<comments>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 10:17:25 +0000</pubDate>
					<dc:creator><![CDATA[Balaji Venkatasubramaniyar]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Operations]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19772</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What Model Context Protocol actually changes for agentic systems, and what it doesn’t]]></custom:subtitle>
		
				<description><![CDATA[Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water through pipes you already designed. MCP changes who’s holding the wrench. Once you’ve felt that shift in a real production system, the “just another API standard” framing stops making sense.</p>



<p class="wp-block-paragraph">This piece is the long version of that argument. It walks through what MCP’s primitives actually are and why they’re the right primitives, where the standard genuinely collapses integration work that used to be duplicated per framework, where the abstraction leaks in ways that only show up once you’re past the demo, and what the protocol’s own 2026 evolution tells you about where the real pain has been. Nearly everything worth knowing about building on MCP falls out of understanding these pieces and how they interact.</p>



<h2 class="wp-block-heading"><strong>What MCP actually standardizes</strong></h2>



<p class="wp-block-paragraph">Strip away the framing and MCP is a JSON-RPC-based protocol that lets a client (the thing driving an LLM) talk to a server that exposes capabilities, over a small, fixed set of primitives:</p>



<p class="wp-block-paragraph"><strong>Tools</strong> are callable functions. Each one has a name, a description, and a JSON Schema describing its inputs. This is the primitive most people mean when they say “MCP,” and it’s the one doing the heavy lifting in most production deployments: “look up an order,” “run a query,” “create a ticket,” etc.</p>



<p class="wp-block-paragraph"><strong>Resources</strong> are readable context, addressed by URI, that a client can pull in without the model having to call a function to get it: a file, a record, or a document, for instance. Think of this as the read side of the interface, separate from the “do something” side that tools represent.</p>



<p class="wp-block-paragraph"><strong>Prompts</strong> are reusable templates a server offers to the client, so common workflows don’t have to be respecified from scratch every time.</p>



<p class="wp-block-paragraph">On top of those three, the spec defines capabilities that flow the other direction, from server back to client: Sampling lets a server ask the client’s model to generate text on its behalf; elicitation, added in the 2025-06-18 revision, lets a server pause and ask the human for more input mid-task; and roots let a server learn which directories or URIs it’s actually allowed to touch.</p>



<p class="wp-block-paragraph">None of these primitives are individually novel. What’s novel is that they’re the same five primitives regardless of which model, which framework, or which vendor is on the client side. That’s the entire value proposition in one sentence, and it’s also the source of everything that goes right and everything that goes wrong when you build on top of it.</p>



<h2 class="wp-block-heading"><strong>Where the old model breaks down</strong></h2>



<p class="wp-block-paragraph">Before MCP, wiring an LLM into an enterprise system meant writing tool-calling code for that specific model, that specific framework, that specific integration. Every agent framework had its own function-calling convention: its own way of describing a schema, its own way of parsing a model’s intent to call something, and its own error-handling contract. Every system you wanted to expose needed its own adapter written to whichever dialect that framework spoke. Add a second framework to your stack and you don’t get twice the work. You get a second, incompatible copy of the same logic, maintained by whoever drew the short straw.</p>



<p class="wp-block-paragraph">MCP replaces that with one contract, written once, usable by any compliant client regardless of which model sits behind it. That’s the part every MCP explainer gets right, and it’s a real, measurable win. I’ve watched it collapse from a maintenance burden that used to scale with the number of frameworks a team happened to be supporting that quarter down to something that scales with the number of systems, full stop.</p>



<p class="wp-block-paragraph">But the more consequential change is where the integration decision gets made. A traditional API integration is an agreement two systems make in advance. You negotiate a contract: endpoints, payloads, auth, versioning, and both sides build to it, because a project plan said this integration should exist. The plan predates the code.</p>



<p class="wp-block-paragraph">An MCP server doesn’t get that luxury. It has no idea which agent will call it, in what sequence, alongside which other servers, in service of what goal a human typed into a chat box 30 seconds ago. The plan doesn’t exist as a concrete thing until the agent composes one, at runtime, out of whatever tools happen to be available to it. That’s not a stylistic difference from the old model. It’s a different category of integration problem, because the party doing the composing isn’t your code anymore. It’s a model, reasoning over natural-language descriptions you wrote weeks or months earlier, with no idea what context it would eventually be reasoning inside of.</p>



<h2 class="wp-block-heading"><strong>The description is the interface now</strong></h2>



<p class="wp-block-paragraph">A tool’s JSON Schema tells the agent what parameters it takes and what shape they need to be. That part is mechanical, and MCP handles it well. The tool’s name and description tell the agent when to use it at all, and whether to prefer it over some other tool that does something adjacent. Those are two different jobs, and only one of them is solved by a well-formed schema.</p>



<p class="wp-block-paragraph">Picture two versions of the same tool description. The first is technically correct and nothing more:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current status of a record given its ID."
}</code></pre>



<p class="wp-block-paragraph">An agent reading that has no idea when this is the right tool versus three other tools that also return some kind of status, no idea what “record” means in this system, and no idea whether IDs are case-sensitive, numeric, or prefixed. The second version spells out the domain the tool operates in, gives the ID format explicitly, states what the returned status values mean, and flags the one adjacent tool this one is commonly confused with and why they’re different:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current fulfillment status for an order record. 
IDs are numeric order numbers (e.g. 48213), not SKUs or customer IDs. Status 
values are one of: pending, processing, shipped, delivered, cancelled. Use this 
instead of get_shipment_status, which returns carrier tracking events rather 
than the order's internal state."
}
</code></pre>



<p class="wp-block-paragraph">That’s a longer description, and it will feel like overexplaining to the engineer writing it, because the engineer already knows all of this. The agent doesn’t. It’s encountering the tool for the first time, with a handful of tokens to decide whether it’s the right call, and no colleague to ask.</p>



<p class="wp-block-paragraph">I’ve watched teams ship a technically correct MCP server that agents used badly, or avoided entirely in favor of a worse but better-described alternative, purely because of this gap. The failure mode isn’t a stack trace. It’s an agent confidently calling the wrong tool, or the right tool with an assumption baked in that happened to be wrong for this case, and nobody notices until the output looks slightly off downstream. Writing tool descriptions well is closer to technical writing and product design than it is to backend engineering, and it’s not a skill most integration teams (mine included, early on) walked in the door with.</p>



<h2 class="wp-block-heading"><strong>Composition is emergent, and that cuts both ways</strong></h2>



<p class="wp-block-paragraph">The entire appeal of MCP is that an agent can combine tools from servers that never agreed to work together, in combinations their respective authors never planned for. That’s also the risk, and it’s structural, not a bug you fix with better testing.</p>



<p class="wp-block-paragraph">In a traditional integration, the sequencing logic (call A, then check its result, then decide whether to call B or C) lives in a script that a human wrote and a reviewer read. You can unit test it. In an MCP-based agent, that same sequencing logic lives in the model’s runtime reasoning, generated fresh for each task based on the goal it was given and whatever tools happen to be available in that session. You can’t unit test a decision that doesn’t exist until the moment it’s made.</p>



<p class="wp-block-paragraph">A tool that behaves correctly in isolation, with the exact inputs its author tested against, can still produce a bad outcome the first time an agent calls it third instead of first in a chain, or passes it a value that came from a different server’s output rather than a human’s direct input. This is qualitatively different from a normal integration bug, because it doesn’t show up in code review, and it won’t show up in testing unless your test suite happens to exercise that specific, unplanned chain of calls. It shows up in production, once, when a particular combination finally occurs. That’s exactly the kind of failure mode that’s cheap to dismiss as an edge case until it happens to the wrong customer.</p>



<h2 class="wp-block-heading"><strong>The protocol is catching up to its own success, in specific and telling ways</strong></h2>



<p class="wp-block-paragraph">To be fair to MCP, it isn’t standing still, and the shape of its evolution tells you a lot about where the real production pain has been. The July 28, 2026 specification is the largest revision since the protocol’s November 2024 launch, and every major change in it traces back to something that broke, or nearly broke, at scale.</p>



<p class="wp-block-paragraph">The protocol core is now stateless. The original design tracked sessions with an MCP-Session-Id header, workable for a single server instance but painful the moment you’re running behind a normal horizontally scaled fleet and discover that “any instance can answer any request” and “sticky session state” don’t coexist. Removing protocol-level sessions means the same request can be served by any instance behind ordinary load-balancing infrastructure, which sounds unglamorous right up until you’re the one who has to explain in an incident review why a routine deploy dropped a chunk of in-flight sessions.</p>



<p class="wp-block-paragraph">Tasks formalize long-running work. A lot of real enterprise work (document processing, multistep approvals, anything involving a human in the loop) doesn’t complete inside a single request/response cycle. Before this extension existed, teams hand-rolled this with polling loops and webhook callbacks, each implementation slightly different, each one a source of its own edge cases. Tasks turn that into a first-class protocol concept.</p>



<p class="wp-block-paragraph">MCP Apps let a server return interactive UI, not just structured data. That matters the moment a “tool” is something a human needs to actually look at and approve before it fires, which in any environment with real consequences attached is often.</p>



<p class="wp-block-paragraph">Authorization was hardened to align with OAuth 2.1 and OpenID Connect. This one isn’t novel so much as overdue, and the gap it closes was a real one; see the governance section below.</p>



<p class="wp-block-paragraph">A formal deprecation policy now governs the legacy HTTP+SSE transport, with a 12-month offramp. That’s the kind of unglamorous governance maturity a protocol only earns after it’s been run in production long enough for someone to need it.</p>



<p class="wp-block-paragraph">None of this is exciting reading. All of it is the sound of a two-year-old protocol absorbing genuine operational scar tissue, which is a far better signal about its trajectory than raw adoption numbers. And the adoption numbers are themselves striking: The official registry tracks close to 10,000 distinct servers, Tier 1 SDK downloads run into the tens of millions monthly, and both the TypeScript and Python SDKs have individually crossed a billion total downloads. Competitors of the protocol’s original author adopted it within months. That combination, real scale plus a spec that keeps changing in response to real production failure modes, is a much stronger signal of durability than either fact alone.</p>



<h2 class="wp-block-heading"><strong>Governance hasn’t caught up</strong></h2>



<p class="wp-block-paragraph">Here’s what I’d want any team to weigh before connecting MCP to anything that matters. Independent security research through 2026 paints a specific, and specifically uncomfortable, picture of the current ecosystem.</p>



<p class="wp-block-paragraph">Scans across thousands of publicly registered servers have found the large majority carrying file-operation patterns prone to path traversal. A meaningful share of tested servers are vulnerable to command injection or server-side request forgery, and there are documented, disclosed cases of tool description poisoning, where the attack lives in the text a model reads to decide what to do rather than in the code the tool actually executes. A closely related failure mode, configuration poisoning, targets the server’s operational baseline directly: stealthy permission changes or altered defaults that persist across sessions and are hard to catch in a normal code review because the malicious logic lives in configuration state, not application code. Multiple high-severity vulnerabilities, including at least one missing-authentication flaw in a major vendor’s own production package, have already been disclosed and patched.</p>



<p class="wp-block-paragraph">None of that is a reason to avoid MCP. It’s a reason to treat it the way you’d treat any protocol that hands an autonomous caller real privileges inside your systems: skeptically, and with the controls in place before the agent gets access rather than after an incident teaches you why you needed them. In practice that means an explicit, enforced allowlist of vetted tools per agent rather than open discovery of whatever happens to be reachable; authentication on every remote endpoint with no quiet exception carved out for “internal” traffic; centralized, immutable audit logging of every tool call an agent makes; and secrets pulled dynamically from a real secrets manager rather than sitting in a server’s local config where a configuration-poisoning attack can find them. None of this is exotic. It’s the same discipline any experienced integration team already applies to systems with real privileges, applied here to a caller that can now improvise its own sequence of actions.</p>



<p class="wp-block-paragraph">Here’s what I’d tell a team starting today.</p>



<p class="wp-block-paragraph">Treat the tool description as reviewed engineering output, not documentation you write last and skim once. Test it against how an agent actually behaves when given it, not just against whether a human reviewer nods along.</p>



<p class="wp-block-paragraph">Assume composition you didn’t plan for will eventually happen, and design tools to fail safely and legibly when it does, rather than assuming a chain of calls you never tested simply won’t occur.</p>



<p class="wp-block-paragraph">Put governance in front of capability, not after it. The allowlist, the auth, the audit log, and the secrets manager are the entry price given where the current vulnerability data sits, not optional hardening for later.</p>



<p class="wp-block-paragraph">And build against the current specification baseline, not whichever example repository you copied six months ago. The stateless core and the authorization changes in the July 2026 spec aren’t cosmetic; targeting an older baseline today is technical debt you’re taking on knowingly, on day one.</p>



<p class="wp-block-paragraph">MCP earned the “not just another API standard” framing honestly. It didn’t get there by being a cleaner REST, or a nicer SDK, or a better-documented function-calling convention: all real, all incremental. It got there by changing who, or what, is actually doing the integration work at runtime. The parts of that job the protocol doesn’t standardize (how well you describe a capability, how safely your tools behave when composed in ways you never anticipated, and how seriously you take governance before you grant an agent real privileges) are exactly the parts worth taking seriously before you bet production traffic on it.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Accelerationist Case for Frontier Pacing</title>
		<link>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/</link>
				<comments>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/#respond</comments>
				<pubDate>Tue, 22 Sep 2026 15:57:06 +0000</pubDate>
					<dc:creator><![CDATA[Venkatesh Rao]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19768</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-accelerationist-case-for-frontier-pacing.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-accelerationist-case-for-frontier-pacing-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Slow AI is smooth AI, smooth AI is fast AI]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Venkatesh Rao’s Substack, Contraptions, and is being republished here with the author’s permission. The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://contraptions.venkateshrao.com/p/the-accelerationist-case-for-frontier" target="_blank" rel="noopener">Venkatesh Rao’s Substack, Contraptions</a>, and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment year) and launched my brief and entirely undistinguished athletic career. By my senior year, however, my 50m time had improved to about 38s (not enough to get me off water-boy duty since the team had several exceptional swimmers with much better times). Interestingly though, it was <em>easier</em> for me to swim faster at the end of my career than it was to swim slower in the beginning. The reason was that in the interim, the coach had significantly improved my stroke and breathing technique. It was all about managed pacing, not raw intensity of effort.</p>



<p class="wp-block-paragraph">There is a fairly deep literature behind this apparently mundane lesson. Daniel Chambliss’s classic 1989 paper “<a href="https://www.jstor.org/stable/202063?seq=1#page_scan_tab_contents" target="_blank" rel="noopener">The Mundanity of Excellence</a>,” based on years of fieldwork studying competitive swimmers all the way from local clubs to the Olympic level, argued that excellence is primarily qualitative rather than quantitative. Elite swimmers do not simply do more of what mediocre swimmers do, or do it harder. They organize their activity differently: Strokes, turns, training habits, attention, and countless other small practices combine into a qualitatively different way of swimming. The route to excellence is not therefore reducible to maximizing effort along some obvious scalar dimension.</p>



<p class="wp-block-paragraph">The same insight is condensed in a maxim common in military and special operations circles: “Slow is smooth, smooth is fast.” In activities where speed really matters, trying to go fast naively is often an excellent way to go slowly.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="1456" height="971" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.jpeg" alt="" class="wp-image-19769" style="aspect-ratio:1.5009380863039399;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.jpeg 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-300x200.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-768x512.jpeg 768w" sizes="auto, (max-width: 1456px) 100vw, 1456px" /></figure>
</div>


<p class="wp-block-paragraph">I never really stopped thinking about this problem. My 2011 book <em><a href="https://books.venkateshrao.com/configurancy-tempo/" target="_blank" rel="noopener">Tempo</a></em> grew partly out of a long-standing interest in pacing across performance domains: how people experience time while making decisions, how rhythms of action emerge, and how timing relates to effectiveness. One of the ideas that has stuck with me since then is that tempo is something to be managed rather than maximized. There is no universally correct speed. There are only tempos appropriate or inappropriate to the dynamics of the situation.</p>



<p class="wp-block-paragraph">Which brings me, somewhat unexpectedly, to Dario Amodei.</p>



<p class="wp-block-paragraph">Amodei <a href="https://darioamodei.com/post/we-must-pace-the-frontier" target="_blank" rel="noopener">recently made the case</a> that frontier AI development should be deliberately paced. His argument is primarily a safety argument. AI capabilities, he believes, are advancing quickly enough that the processes required to understand, evaluate, align, secure, and safely operate them are having trouble keeping up. This is not quite the old proposal for an AI “pause.” Pacing means continuing to advance the frontier while deliberately managing its rate, allowing safety work and institutional capacity to remain within striking distance of capability. Sam Altman has now endorsed the basic proposition, and Demis Hassabis has made closely related arguments about frontier capabilities outrunning scientific understanding and governance capacity. Elon Musk, more tersely, has said that Amodei is right.</p>



<p class="wp-block-paragraph">There is an obvious cynical reading of this emerging consensus. The leading frontier labs have powerful economic reasons to want a regulated frontier. A regime that requires enormous compliance budgets, restricts open-weight releases, discourages foreign models, imposes burdens that startups cannot afford, or legitimizes coordination among a small number of incumbents could turn “safety” into a remarkably effective mechanism for protectionism and regulatory capture. That suspicion is not paranoid. Open models increasingly constitute a competitive threat to proprietary frontier providers, and the politics around regulating them <a href="https://www.washingtonpost.com/wp-intelligence/ai-tech-brief/2026/07/20/ai-tech-brief-open-source-debate/" target="_blank" rel="noopener">already feature</a> explicit accusations of regulatory capture. There is an additional awkwardness: Coordinated pacing among nominal competitors looks uncomfortably like coordinated restriction of output, enough so that the legality of such arrangements under antitrust law <a href="https://www.wired.com/story/openai-wants-to-know-if-an-ai-industry-slowdown-would-even-be-legal/" target="_blank" rel="noopener">is already being debated</a>.</p>



<p class="wp-block-paragraph">I don’t think we need to resolve the question of motives. Perhaps these CEOs are sincerely terrified. Perhaps they are sincerely terrified <em>and</em> understand perfectly well that the regulations they favor would strengthen their competitive positions. Perhaps the mixture varies by person, company, and day of the week. It doesn’t matter much for my argument. The proposition that the frontier should be paced is worth considering independently of the political economy of the people proposing it.</p>



<p class="wp-block-paragraph">I also don’t share enough of Amodei’s safety premises to make his argument my own. In particular, I think a great deal of contemporary concern about runaway AGI, superintelligence, and “alignment” is badly framed, and often borders on the theological. But I increasingly agree with his conclusion.</p>



<p class="wp-block-paragraph">In fact, I think there is a strong case for frontier pacing even if you are an accelerationist and your objective is simply to make technological progress happen as fast as possible. I am not myself an accelerationist. My preferred framing is closer to managed tempo. But if I were one, I would still favor pacing the frontier right now, for a simple reason: <em>Maximizing the instantaneous velocity of the AI capability frontier is no longer obviously maximizing the rate of technological progress.</em></p>



<p class="wp-block-paragraph">There is, however, an important difference between my conclusion and the emerging frontier consensus. Their natural solution is coordination at the top: labs agreeing upon thresholds, governments blessing the coordination, evaluators policing it, and eventually perhaps international agreements extending it.</p>



<p class="wp-block-paragraph">In other words, cartelization, hopefully of a benign sort.</p>



<p class="wp-block-paragraph">I would prefer to see how much frontier pacing can be produced from the bottom up through ordinary market mechanisms. The distinction matters. The objective should not be to decide administratively how fast AI is allowed to improve. It should be to stop artificially rewarding frontier velocity after frontier velocity has ceased to be the most important form of progress.</p>



<h2 class="wp-block-heading"><strong>Getting inside the loop</strong></h2>



<p class="wp-block-paragraph">A useful way to understand the distinction comes from another idea that startup culture has borrowed, and mostly misunderstood, from the military: John Boyd’s OODA loop. OODA theory says that you win by “getting inside the adversary’s decision cycle,” which is usually glossed as making decisions <em>faster</em> than the other guy. If you observe, orient, decide, and act faster than he can, the story goes, you eventually overwhelm him.</p>



<p class="wp-block-paragraph">But <em>inside</em> does not mean <em>faster</em>. The objective is to operate within the decision dynamics of the system you are engaging in a way that lets you shape them. Against a human adversary, that may indeed sometimes involve accelerating until his ability to orient collapses psychologically. But it may also require waiting, withholding action, changing rhythm, or deliberately slowing down. In nonadversarial situations, the goal may not be collapse at all but harmonization for resonant support.</p>



<p class="wp-block-paragraph">What matters is the right tempo at the right phase, not speed for the sake of speed.</p>



<p class="wp-block-paragraph">Something analogous applies to scientific and technological progress. There is no enemy psychology to collapse, but there are still loops to get inside: observation, experimentation, interpretation, investment, construction, deployment, feedback, learning, and recombination. The useful question is not how rapidly one component of that system can be made to move. It is whether the tempo of development allows those loops to close. If one subsystem changes faster than the surrounding system can observe, understand, absorb, and respond to it, pushing that subsystem still faster can reduce rather than increase effective progress. It can induce fragility and collapse.</p>



<p class="wp-block-paragraph">This, I think, is approximately where AI is now.</p>



<p class="wp-block-paragraph">The simplest evidence is personal and almost embarrassingly mundane. Frontier AI is already overpowered for nearly everything I use it for. In my most advanced projects I may use the strongest model available (Fable for my coding projects) to plan an approach or make critical strategic decisions, but I can generally hand the resulting specification to a cheaper model (such as Opus or Sonnet) to do the routine work. For ordinary uses I don’t need anything close to the frontier. In ChatGPT, I no longer even know exactly which model I am talking to much of the time. Whatever the “think harder” control does is sufficient model selection for my purposes.</p>



<p class="wp-block-paragraph">The situation increasingly reminds me of smartphones. There was a period when getting the newest iPhone produced a noticeable improvement in everyday life. Eventually the hardware got good enough that the upgrade cycle ceased to matter much. I kept an iPhone XS for almost a decade before replacing it with a 16. The frontier continued advancing; I simply fell off the frontier because my demand curve had stopped following it.</p>



<p class="wp-block-paragraph">Something similar is beginning to happen with AI, except that the supply curve is moving incomparably faster. Six months ago I routinely maxed out token allotments. Now I don’t. Some weeks I barely use coding agents. This isn’t because I’ve become less interested in AI. It is because my own capacity to productively absorb AI output has become the constraint. I have projects to think about, things to read, people to talk to, and work to do in domains where AI cannot help me yet, or perhaps ever. I am already pacing myself at my own tiny personal frontier.</p>



<p class="wp-block-paragraph">That is a significant change in the technological situation. The binding constraint is migrating.</p>



<h2 class="wp-block-heading"><strong>When the bottleneck moves</strong></h2>



<p class="wp-block-paragraph">Broader AI deployment is increasingly blocked by things other than model intelligence. Robotics has long been constrained by actuators, power, reliability, dexterity, manufacturing, and the sheer recalcitrance of the physical world. Those constraints are beginning to move, but making the model smarter does not make them disappear. AI in education is constrained less by whether a model can explain calculus than by our lack of sufficiently rich classroom experimentation about what happens when students and teachers actually use these systems. Current mid-tier models are probably capable enough to power almost any educational experiment worth trying in a high-school or undergraduate classroom. We do not need another order of magnitude of intelligence before conducting them.</p>



<p class="wp-block-paragraph">This pattern should become more common as AI improves. Once intelligence ceases to be scarce, its complements become more important. Model capability can be abundant while classroom knowledge is scarce. Model capability can be abundant while actuators are scarce. Model capability can be abundant while electrical infrastructure is scarce. It can be abundant while organizational competence, human attention, scientific understanding, military doctrine, security practices, and good judgment are scarce.</p>



<p class="wp-block-paragraph">This is not peculiar to AI. Capability-maxxing the coolest new weapon is bad military doctrine. The United States has enjoyed extraordinary technological superiority over its adversaries for decades and has nevertheless repeatedly discovered that superior equipment does not automatically produce strategic success. Logistics, doctrine, morale, training, political understanding, industrial capacity, and orientation matter. A force that neglects those complements because it possesses the best weapons can become remarkably fragile. From Vietnam to Iran, the US military has been repeatedly forced to relearn the lesson.</p>



<p class="wp-block-paragraph">AI may now be entering the same regime. Fragility from neglect of everything non-AI is becoming a bigger risk than failure to token-max.</p>



<p class="wp-block-paragraph">There is a further reason to suspect that continuing to redline the existing frontier may yield diminishing returns. The major labs increasingly appear to be competing along broadly the same technological S-curve. One suggestive sign is that they run into the same supply constraints: HBM, electrical power, data-center capacity, capital, and access to sufficiently large clusters. When a technological system moves onto a genuinely different S-curve, its important bottlenecks often change as well. If everybody’s problem is how to secure more of the same scarce inputs to do more of the same basic thing, that is at least suggestive that everybody is climbing the same sigmoid.</p>



<p class="wp-block-paragraph">As I learned as a freshman swimmer, near the upper portion of an S-curve, pushing harder can become exactly the wrong acceleration strategy. You expend increasing resources for decreasing gains while starving exploration of the attention required to discover the next curve. Moving faster <em>along</em> an S-curve is not the same thing as accelerating technological evolution. Sometimes you have to back off the incumbent trajectory long enough to notice what the next trajectory is.</p>



<h2 class="wp-block-heading"><strong>A cognitive ergonomics crisis</strong></h2>



<p class="wp-block-paragraph">There is also a more immediate bottleneck that the AI industry seems reluctant to acknowledge: the humans at the frontier.</p>



<p class="wp-block-paragraph">Startup people have been LARPing war for decades. This is one reason concepts like OODA became popular in startup culture in the first place. “War mode” usually means working extremely hard under conditions of strong personal financial incentives: long hours, high urgency, extreme focus, centralized authority, and a willingness to sacrifice ordinary organizational niceties. I’ve been around startup culture for decades, and I suspect the AI boom may be the first time the conditions have actually become meaningfully war-like.</p>



<p class="wp-block-paragraph">AI frontier people are visibly unprepared for it.</p>



<p class="wp-block-paragraph">Actual militaries and other frontline risk professions take the human consequences of sustained high-stress operations seriously. Soldiers, firefighters, emergency medical personnel, disaster responders, surgeons, pilots, and others operating in consequential environments develop elaborate practices around training, emotional regulation, redundancy, rotations, decompression, mandatory rest, checklists, after-action review, and recovery. These practices exist because motivation does not repeal physiology. Judgment deteriorates. Attention narrows. People make stupid mistakes. Emotional reactions become harder to regulate. Creativity disappears. Eventually people break.</p>



<p class="wp-block-paragraph">Does frontier AI look like an industry managing itself accordingly?</p>



<p class="wp-block-paragraph">From the outside, it looks closer to the opposite. People at the frontier have been operating under extraordinary pressure for several years with little respite. The cognitive ergonomics of their working conditions are a disaster. (We have a <a href="https://protocol-institute.org/projects/project/?slug=cognitive-ergonomics" target="_blank" rel="noopener">project going</a> at the Protocol Institute led by <a href="https://open.substack.com/users/17195021-timber-stinson-schroff?utm_source=mentions" target="_blank" rel="noopener">Timber Stinson-Schroff</a> to study this—contact him if you’re interested in participating in or supporting it.)</p>



<p class="wp-block-paragraph">Competitive pressure, enormous amounts of capital, geopolitical attention, hostile public scrutiny, internal ideological battles, rapidly changing technology, and the conviction among some participants that their daily work may determine the fate of humanity are not normal occupational stressors. The rate of dumb, unforced errors appears to be rising. The quality of frontier discourse has, in my view, visibly deteriorated. People I once assumed were much smarter than me increasingly seem to be missing obvious things while becoming susceptible again to bad ideas I thought they had outgrown.</p>



<p class="wp-block-paragraph">Tired people catch colds more easily; they catch bad ideas more easily too.</p>



<p class="wp-block-paragraph">This produces a peculiar inversion of the conventional AI safety model. We normally imagine increasingly unreliable or dangerous AIs surrounded by reliable human supervisors. But what if we are increasingly producing extremely capable AIs surrounded by progressively less reliable humans?</p>



<p class="wp-block-paragraph">“Human in the loop” is not much of a safety guarantee if the human has been metaphorically deployed aboard an aircraft carrier in a war zone for eight months without relief.</p>



<p class="wp-block-paragraph">Some of the public testimony emerging from frontier organizations should perhaps be interpreted through this lens. I do not want to diagnose particular people from afar, and testimony from people who have worked closely with frontier systems should obviously be taken seriously. But when someone emerges from prolonged immersion at the frontier sounding psychologically shattered, there are at least two possible kinds of information in the signal. One concerns the technology. The other concerns what prolonged immersion at the frontier does to the observer. Frontier workers are sensors, but the sensors themselves are being perturbed by the phenomenon they are measuring.</p>



<p class="wp-block-paragraph">We have seen versions of this going back at least to the Blake Lemoine episode at Google, when sustained interaction with LaMDA led him to conclude that the system was sentient. More recently, former frontier employees have emerged making extraordinarily grave predictions about where AI is heading, that ill-prepared, tech-hostile journalists are eagerly amplifying with lurid headlines.</p>



<p class="wp-block-paragraph">The correct response need not be either “believe them and stop AI” or “they’re crazy and should be ignored.” Sometimes a sensible response to someone coming back from the front sounding shell-shocked and exhibiting symptoms of PTSD is: This person needs a vacation. We rotate soldiers partly because the testimony of exhausted soldiers matters.</p>



<p class="wp-block-paragraph">The largest near-term AI safety concern may therefore be exhausted frontline humans supervising overpowered AIs.</p>



<h2 class="wp-block-heading"><strong>Fighting the wrong enemy</strong></h2>



<p class="wp-block-paragraph">Exhaustion is particularly dangerous when nobody can agree about what the enemy is. Much of the actual stress experienced by frontier organizations comes from a fairly comprehensible mixture of competitive pressure and techlash hostility. Those forces are intense, but neither is an existential adversary.</p>



<p class="wp-block-paragraph">The clearest live adversarial problem involving AI is much more ordinary: humans using AI against other humans. Criminal applications are already real and deserve serious attention. Military applications are rapidly becoming real as well, and the relevant strategic picture is much broader than a stylized US-versus-China AI race. Smaller powers and nonstate actors can use cheap cognitive capability to lower engineering barriers that previously required deeper technical institutions. Recent reporting, for example, describes AI assistance being used in weapons-engineering work by actors in Houthi-controlled Yemen. That strikes me as the kind of development around which one can build a concrete threat model.</p>



<p class="wp-block-paragraph">Longer-term military diffusion is clearly a serious concern. But much of the fear actually shaping frontier behavior seems aimed somewhere else entirely: toward vague runaway “AGIs,” “superintelligences,” and a metaphysically capacious notion of “alignment” inherited from philosophical traditions I find largely unpersuasive.</p>



<p class="wp-block-paragraph">There is a useful analogy with climate change. Climate change produces actual physical stressors: more extreme weather, unstable agricultural conditions, infrastructure damage, wildfire risk, and so on. Those generate concrete political and humanitarian problems, including displacement and unmanaged refugee flows, while longer-term adaptation requires things like shoreline management, wildfire regimes, agricultural relocation, and preparedness for changing disease ecologies. Yet parts of climate politics have preferred to identify an ultimate metaphysical adversary called Capitalism, Markets, or Growth and an equally totalizing remedy called degrowth. Heterogeneous problems with different timescales and mechanisms get collapsed into one grand theory.</p>



<p class="wp-block-paragraph">Parts of AI safety discourse increasingly strike me the same way. Competitive instability, cybercrime, weapons proliferation, institutional disruption, labor-market effects, and exhausted frontier personnel are all real and different problems. “Unaligned superintelligence” turns them into a single theological object. Once that happens, every stressor becomes evidence for the same threat model.</p>



<p class="wp-block-paragraph">This is a kind of threat-model collapse. Adaptation is usually plural; apocalypse is singular. Real technological transitions produce dozens of mismatched rates and local failure modes, requiring different responses at different tempos. If criminals are the problem, work on security and law enforcement. If weapons diffusion is the problem, work on doctrine and proliferation. If operators are exhausted, rotate them. If schools lack experimental knowledge, run experiments. If power is scarce, build infrastructure. “Align superintelligence” is not a substitute for any of those things.</p>



<p class="wp-block-paragraph">Pacing would give us something valuable here beyond safety: enough time to discriminate among threats.</p>



<h2 class="wp-block-heading"><strong>Proof abundance, understanding scarcity</strong></h2>



<p class="wp-block-paragraph">Mathematics may already offer a miniature preview of what happens when one part of a knowledge-production system accelerates far beyond the others.</p>



<p class="wp-block-paragraph">Terence Tao has recently <a href="https://teorth.github.io/tao-web/ai-views.html?utm_source=chatgpt.com" target="_blank" rel="noopener">distinguished</a> three stages of mathematical work: generation, verification, and digestion. AI is rapidly making the first two cheaper. Models can generate candidate proofs, while formal systems such as Lean can increasingly verify them. But digestion remains stubbornly slow. Somebody still has to understand what the proof is doing, relate it to existing mathematics, extract reusable techniques, explain it, teach it, and use the resulting understanding to generate better questions. Tao describes the resulting condition as an “impedance mismatch.”</p>



<p class="wp-block-paragraph">This is particularly interesting in light of what I have elsewhere called the <a href="https://contraptions.venkateshrao.com/p/the-curiously-playable-universe" target="_blank" rel="noopener">curiously playable universe</a>: the apparently expanding set of domains that can be transformed into sufficiently explicit games that AI can optimize effectively within them. Anything that begins to resemble a CAD system, a formal proof environment, or an evolutionary optimization problem over a sufficiently well-defined parameter space becomes potentially tractable to extraordinarily capable models. More of the world appears to be playable than we previously thought.</p>



<p class="wp-block-paragraph">But playability has an important pathology. A highly playable domain supplies a scoreboard, and once AI becomes extraordinarily good at optimizing the scoreboard, the relationship between winning the game and advancing the larger domain can weaken. Solving a theorem is valuable partly because, historically, getting to the solution usually required acquiring understanding along the way. If an AI can helicopter directly to the summit, to borrow Tao’s analogy, the summit has still been reached, but nobody necessarily learned the trails, landmarks, terrain, or neighboring geography encountered during the climb.</p>



<p class="wp-block-paragraph">Tao and two dozen other Fields Medalists recently made essentially this point in a declaration strikingly titled “<a href="https://mathandai.org/" target="_blank" rel="noopener">A Severe Misalignment of AI in Mathematics</a>.” Their pointed use of <em>misalignment</em> is almost the reverse of its standard AI-safety meaning. The problem they identify is not that AI has developed alien goals. It is that the incentives of AI companies to demonstrate spectacular problem-solving performance are becoming misaligned with the goals of mathematics itself. Solving difficult problems has historically served as a proxy for mathematical understanding and progress. Once AI can optimize the proxy directly, the correlation can break.</p>



<p class="wp-block-paragraph">The recent Navier–Stokes episode illustrates the issue. Enormous amounts of inference can now be directed at a famous open problem, candidate constructions produced, and formal verification generated at extraordinary speed. Yet that does not automatically produce a corresponding increase in comprehensible, reusable mathematical knowledge. The pipeline is something like problem selection → generation → verification → exposition → digestion → canonicalization → better questions. Increasing the bandwidth of generation and verification by orders of magnitude while leaving the downstream stages roughly unchanged creates a queue.</p>



<p class="wp-block-paragraph">Proof generation becomes abundant. Understanding becomes scarce.</p>



<p class="wp-block-paragraph">That is frontier pacing in miniature. Maximum local throughput does not imply maximum system throughput. Indeed, beyond a certain point it can create congestion.</p>



<h2 class="wp-block-heading"><strong>Orientation beats equipment</strong></h2>



<p class="wp-block-paragraph">There is a strategic implication here for organizations outside the frontier labs. The natural reaction to rapidly advancing models is to assume that whoever possesses the strongest model necessarily possesses an overwhelming advantage. If a frontier model can turn increasingly playable engineering problems into few-shot solutions, and if most of the necessary input information exists somewhere in public literature, then organizations can easily conclude that whatever intellectual lead they possess is temporary. Why bother competing with organizations that possess better models, more compute, more money, and privileged access to the frontier?</p>



<p class="wp-block-paragraph">But this risks confusing equipment superiority with orientation superiority.</p>



<p class="wp-block-paragraph">Boyd repeatedly emphasized that superior orientation could overcome substantial equipment disadvantages. (“We’d still have won if we’d swapped equipment.”) The relevant analogy today is something like centaur chess. Your model does not necessarily have to outthink their model. Your humans have to out-orient their humans.</p>



<p class="wp-block-paragraph">This becomes increasingly true as frontier capabilities bunch together above the threshold required for a particular task. In my own work, I increasingly find that I can use the strongest model to formulate or specify a solution and then hand most of the execution to a weaker model. For many problems, even that is overkill. I would readily bet on a well-oriented person using a slightly weaker model against a poorly oriented person using the strongest available model.</p>



<p class="wp-block-paragraph">And the frontier labs have no automatic orientation advantage. Quite the contrary: They are simultaneously fighting an extraordinary number of battles under extreme strategic distraction. They are building models, securing compute, raising capital, negotiating with governments, managing safety factions, defending themselves against critics, competing for talent, building consumer products, selling enterprise software, contemplating hardware and robotics, responding to geopolitical pressure, and trying to decide what sort of companies they are becoming. They may have a model advantage while suffering an orientation disadvantage. There is no reason to assume that organizations exceptionally good at building foundation models are exceptionally good at everything their models can be applied to.</p>



<p class="wp-block-paragraph">This is another reason pacing can be strategically productive. It creates room for orientation. In an environment saturated with FUD and “resistance is futile” rhetoric, organizations can lose before competing because they assume frontier capability automatically determines every downstream contest. It doesn’t. Superior orientation does.</p>



<h2 class="wp-block-heading"><strong>From one S-curve to the next</strong></h2>



<p class="wp-block-paragraph">Put all of this together and Amodei’s proposal starts to look different. His concern is that capability is outrunning safety. I think capability may be outrunning almost everything.</p>



<p class="wp-block-paragraph">It is outrunning our ability to deploy it productively. It is outrunning classroom experimentation, organizational adaptation, security practice, mathematical digestion, physical infrastructure, and human attention. It may be outrunning our ability to distinguish actual threats from theological ones.</p>



<p class="wp-block-paragraph">And it is almost certainly outrunning the decompression and recovery cycles of some of the people charged with making the most consequential decisions about it.</p>



<p class="wp-block-paragraph">An accelerationist should care about every one of these things precisely because an accelerationist wants acceleration.</p>



<p class="wp-block-paragraph">The mistake is to identify acceleration with the derivative of a single visible variable: benchmark scores, parameter counts, inference budgets, training compute, or whatever happens to define the current frontier. Technological progress is a coupled system. Accelerating one component beyond the absorption capacity of its complements eventually stops accelerating the system. The problem becomes especially acute near the top of an S-curve, where enormous resources can be consumed eking out diminishing improvements while the exploration necessary to find the next curve is crowded out.</p>



<p class="wp-block-paragraph">There is a useful precedent in the history of the PC industry. For years, processor clock frequency functioned as the wonderfully simple consumer metric for progress: 486 MHz was better than 400 MHz; 1 GHz was better than 800 MHz; higher number, faster computer. Manufacturers had every reason to compete on the legible scalar, and consumers learned to buy it. Eventually this became the “<a href="https://www.theguardian.com/technology/2002/feb/28/onlinesupplement3" target="_blank" rel="noopener">megahertz myth</a>.” Different architectures could do very different amounts of useful work per clock cycle, while pushing frequency upward ran increasingly hard into heat and power constraints. By the mid-2000s, the industry was moving toward multicore designs and a more complicated understanding of performance in which throughput, architecture, workload, thermal limits and performance per watt all mattered. Intel itself acknowledged at the time that as computer usage diversified, factors other than clock speed were becoming increasingly important to platform performance.</p>



<p class="wp-block-paragraph">AI benchmark culture looks increasingly like the early stages of the same mistake. A benchmark is useful because it compresses a complicated question into a number. When capability is scarce and improvements are large, the number may track value surprisingly well. As systems become overpowered for more uses, however, the proxy begins to detach from what customers actually care about. A model that goes from 87 to 91 on some benchmark may represent an impressive scientific achievement while producing essentially zero additional value for a company whose relevant workload was already handled adequately at 75.</p>



<p class="wp-block-paragraph">This suggests a path to frontier pacing that does not require a council of frontier CEOs deciding how quickly everyone is allowed to move.</p>



<p class="wp-block-paragraph"><em>Customers can simply become harder to impress.</em></p>



<p class="wp-block-paragraph">Enterprise buyers can demand demonstrated improvements on their actual workloads rather than accepting leaderboard gains as evidence of value. Developers can (and already do) route work to the cheapest model that clears the capability threshold rather than reflexively calling the smartest one. Researchers can value useful scientific infrastructure, explanation and reusable knowledge rather than merely celebrating another famous benchmark or theorem knocked down. Investors can become less impressed by capital expenditure whose primary justification is preserving position on a frontier whose marginal economic value is falling. Users can decline to upgrade when the previous generation is already good enough. Current enterprise behavior already points in this direction: Cheaper and open-weight models are becoming attractive precisely because many workloads do not require frontier intelligence, while buyers increasingly demand measurable returns rather than capability in the abstract.</p>



<p class="wp-block-paragraph">None of these mechanisms requires anybody to agree upon a socially optimal rate of AI development. They simply improve the feedback signal facing producers. The market stops saying “more intelligence, at almost any price” and begins saying “show me what this additional intelligence is for.”</p>



<p class="wp-block-paragraph">There are supply-side versions too. As power becomes a binding constraint, performance per watt and useful inference per dollar should matter more than sheer training scale. As inference proliferates toward edge devices and private deployments, latency, reliability, privacy and local controllability become competitive dimensions. As organizations discover that weaker models can execute plans produced by stronger ones, heterogeneous model portfolios should compete with monolithic frontier consumption. Open-weight and decentralized systems can keep proprietary labs honest by making “good enough” intelligence cheap and difficult to monopolize. A mature AI market should develop more dimensions of performance precisely as the mature processor market did.</p>



<p class="wp-block-paragraph">This is the sort of pacing I would prefer: not a speed limit but a richer scoreboard.</p>



<p class="wp-block-paragraph">The objective should therefore not be maximum speed. It should be managed tempo for maximum actual progress. Sometimes that means sprinting. Sometimes it means dwelling at a capability level while applications, institutions, infrastructure, science, and humans catch up. Sometimes it means letting one subsystem race ahead while another rests.</p>



<p class="wp-block-paragraph">Sometimes it means deliberately leaving expensive capability unused.</p>



<p class="wp-block-paragraph">That last possibility may be the hardest one for AI culture to accept, because we are still psychologically adapting to the idea that intelligence might actually be abundant.</p>



<p class="wp-block-paragraph">A true sense of abundance does not require you to max out the bounty. Nobody hyperventilates because free oxygen might disappear before they get their fair share. When something is genuinely abundant, you can waste it. You can use a frontier model for a trivial question. You can use a weaker model because it is good enough. You can leave tokens unused. You can spend a week doing something that doesn’t involve AI. You can allow an extraordinarily powerful model to sit idle while you think.</p>



<p class="wp-block-paragraph">The mark of abundance is waste, including nonuse.</p>



<p class="wp-block-paragraph">Compulsive token-maxxing is in this sense still a scarcity behavior. So is compulsive benchmark-maxxing, compute-maxxing, and capability-maxxing. Train now because somebody else will. Deploy now because the window might close. Consume all the intelligence available because leaving any unused feels like falling behind. An industry behaving this way may possess an abundance of intelligence without yet having developed an abundance mentality.</p>



<p class="wp-block-paragraph">Slack is not necessarily the enemy of acceleration. Slack is where people recover, where institutions adapt, where strange experiments happen, where understanding catches up with proof, where neglected complements receive attention, and where somebody finally notices that the old S-curve is flattening and another one is waiting nearby.</p>



<p class="wp-block-paragraph">So yes, pace the frontier. But don’t turn the frontier labs into a cartel to do it. Let safety work catch up, but also let customers become bored with vanity benchmarks. Let exhausted researchers sleep. Let mathematicians digest their proofs. Let schools figure out what to do with the models they already have. Let robotics catch up. Let organizations learn to orient themselves in a world where intelligence is cheap. Let markets discover that efficiency, reliability, privacy, integration and domain-specific usefulness sometimes matter more than another few points on a benchmark. Let us discover which risks are real, which bottlenecks have moved, and which parts of the world turn out to be playable.</p>



<p class="wp-block-paragraph">Then, when the situation calls for it, accelerate again.</p>



<p class="wp-block-paragraph">Slow is smooth. Smooth is fast.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI</title>
		<link>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/</link>
				<comments>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/#respond</comments>
				<pubDate>Tue, 22 Sep 2026 11:00:29 +0000</pubDate>
					<dc:creator><![CDATA[Nicole Butterfield]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Innovation & Disruption]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19763</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-Sovereignty—Bargaining-with-Big-Tech-and-the-Promise-of-Full-Stack-Open-Source-AI.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-Sovereignty—Bargaining-with-Big-Tech-and-the-Promise-of-Full-Stack-Open-Source-AI-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The early rapid expansion of AI capabilities that focused on frontier models was largely ushered into the world by a few powerful, US-based AI labs. Open-weight models released from labs in China, early on from DeepSeek, and later from Moonshot, Z.ai, and others, have in part disrupted that dominance. But growing concerns about the concentration [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The early rapid expansion of AI capabilities that focused on frontier models was largely ushered into the world by a few powerful, US-based AI labs. Open-weight models released from labs in China, early on from DeepSeek, and later from Moonshot, <a href="http://z.ai" target="_blank" rel="noopener">Z.ai</a>, and others, have in part disrupted that dominance. But growing concerns about the concentration of power have led to discussions about the need for at least some level of <em>AI sovereignty</em>.</p>



<p class="wp-block-paragraph">AI sovereignty doesn’t necessarily imply total control of your AI stack. It holds the promise of having more localized security and privacy, better adherence to local jurisprudence (for example EU AI and data laws), more dependable service, and potentially more culturally specific outputs from the AI technologies used within a specified border or region. Negotiating interdependence is not necessarily a problem, but having an array of tools beyond just open-weight models and options in those negotiations beyond just commercial offerings is imperative.</p>



<p class="wp-block-paragraph">Unsurprisingly, the same AI labs that gave rise to the need for AI sovereignty are also pushing their own solutions to the problem. While local institutions consider and even adopt some of these initial “sovereignty” offerings from these labs, open source AI technologies may offer a more promising horizon, more flexibility, and a means to manage dependencies. Tim O’Reilly argues open source AI is a potential opening for <a href="https://www.oreilly.com/radar/ai-sovereignty-and-the-architecture-of-participation/" target="_blank" rel="noopener">greater participation</a> in the future of AI’s development.</p>



<p class="wp-block-paragraph">From the underlying chip technology that is necessary for model training and inference to the cloud and data infrastructure that enables model development, companies such as NVIDIA, OpenAI, Google, Microsoft, and AWS have begun to stake out their own territory to maintain relevance within the global push toward AI sovereignty. Stanford University’s Human-Centered Artificial Intelligence Lab (HAI) lays out the different approaches and offerings these labs have developed in its report <em><a href="https://hai.stanford.edu/policy/the-commercial-landscape-of-ai-sovereignty-offerings" target="_blank" rel="noopener">The Commercial Landscape of AI Sovereignty Offerings</a></em>. It argues that while these labs “promise that countries will own their AI stack, [they also] deepen dependencies on U.S. Big Tech.”</p>



<p class="wp-block-paragraph">There are some non-US-based commercial alternatives that offer their own “full stack” solutions or AI sovereignty for specific layers of the stack. Companies in Europe, the Gulf region, Asia, and elsewhere are positioning themselves as local alternatives to US tech oligarchs. According to the <a href="https://hai.stanford.edu/policy/the-commercial-landscape-of-ai-sovereignty-offerings" target="_blank" rel="noopener">HAI report</a>, “Many of the most mature and advanced companies are actively backed by their governments. In these cases, sovereignty is not just a marketing claim but a stated policy objective, with governments directing funding, structuring procurement, and, in some cases, selecting specific companies to build out domestic AI capacity on their behalf.”</p>



<p class="wp-block-paragraph">These types of collaboration can both enable independence from US labs but may also create openings for political intervention. Claims about censorship and control of Chinese models emerged quickly after DeepSeek’s initial 2025 release. More recently, there have been <a href="https://apnews.com/article/artificial-intelligence-chatbots-censorship-bias-free-speech-fed8fdbf90751c10fe77b77832e0ffba" target="_blank" rel="noopener">probes into how US models</a> may limit certain types of discourse. Moreover, the HAI report points out that often the offerings of these alternative providers still rely on the underlying technologies, specifically chips and cloud infra, of the US labs.</p>



<p class="wp-block-paragraph">The proliferation of commercial offerings provides the space for diversification or potential leverage to negotiate better terms for collaboration, even with the dominant players. Open source AI technologies also play an important role in creating opportunities for even more diversification and greater sovereignty. As HAI argues, “Sovereignty strategies that do not consider the role of open-source AI risk normalizing fragmentation and political overreach.”</p>



<p class="wp-block-paragraph">In order for open source AI to counter the diversification of commercial sovereignty offerings, these technologies must also proliferate beyond open-weight models. Arguing for a “<a href="https://www.oreilly.com/radar/ai-sovereignty-and-the-architecture-of-participation/" target="_blank" rel="noopener">federated system</a>” of open source AI that enables sovereignty based on an “architecture of participation,” Tim O’Reilly writes that “the right infrastructure to let us satisfy both goals [of being everywhere and allowing everyone to have a say] will be a federation of models, a federation of protocols and code, and a federation of capacity. We need an architecture of participation all the way down the stack, and all the way up.” A key technology in the expansion of the open source AI stack these days are agent harnesses.</p>



<p class="wp-block-paragraph">In a recent article, Mozilla CTO <a href="https://www.oreilly.com/people/raffi-krikorian/" target="_blank" rel="noopener">Raffi Krikorian</a> argues that “the orchestration layer above the [model] weights is where capability is concentrating, and closed labs are already welding it shut”; therefore, it’s imperative to build on open harnesses, not just models. Commercial offerings that have dominated thus far include Claude Code and Codex. OpenClaw offered an initial disruption and promise for open source in late 2025, though the creator was quickly absorbed into OpenAI’s organization. While big tech labs continue to absorb when, who, and what they can, NousResearch’s self-improving Hermes agent harness has also garnered substantial attention now with over 230,000 stars on GitHub. More recently harnesses such as Pi and DeepSeek Harness are expanding that open source offering, heeding Krikorian’s call.</p>



<p class="wp-block-paragraph">Beyond agent harnesses, some of the strongest open source projects are developing in the less visible layers. Inference engines such as vLLM, SGLang, llama.cpp, and ONNX Runtime make it possible to serve a range of models efficiently across data centers, regional clouds, personal computers, and edge devices. Ray, which was developed by researchers at UC Berkeley, distributes demanding AI workloads. Ollama lowers the barrier to running models locally. Together, these projects give institutions more freedom to change models, hardware, and hosting providers without rebuilding an entire system around another company’s proprietary platform.</p>



<p class="wp-block-paragraph">Other fast-growing projects are filling out the data, interoperability, and accountability layers of the stack. The Model Context Protocol and Agent2Agent Protocol offer open standards through which agents can connect to tools and to one another. Qdrant, Chroma, Milvus, and LanceDB provide open infrastructure for storing and retrieving institutional knowledge. MLflow, Opik, and OpenLLMetry allow developers to evaluate, trace, and monitor AI applications without surrendering operational data to a closed dashboard. The <a href="https://www.aipotluck.org/map/gap-map" target="_blank" rel="noopener">AI Potluck Gap Map</a> classifies inference, deployment, and agent protocols as mature open ecosystems but identifies resiliency gaps in storage and observability, where fewer fully open projects occupy the leading tier. These gaps point toward an important investment agenda. Sovereignty will depend not on finding a single open replacement for Big Tech but on sustaining interoperable public alternatives across every consequential layer of the stack.</p>



<p class="wp-block-paragraph">Even still, the looming threat of acquisitions and absorption of open source is persistent. The fintech company <a href="https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter" target="_blank" rel="noopener">Stripe recently bought OpenRouter</a>, a platform that allows developers to access various models and has become a primary hub for accessing and routing open-weight models in particular. NVIDIA <a href="https://arstechnica.com/ai/2026/08/report-nvidia-to-acquire-ai-model-repository-hugging-face-for-13-billion/" target="_blank" rel="noopener">has acquired Hugging Face</a>, one of the key players for the open source AI ecosystem, and it has also just settled a <a href="https://finance.yahoo.com/technology/ai/articles/nvidia-7-billion-poolside-deal-224313075.html" target="_blank" rel="noopener">licensing deal with Poolside</a>, which develops open-weight coding models, purportedly to avoid the oversight of complete acquisition.</p>



<p class="wp-block-paragraph">AI sovereignty may mean the necessity of “calibrating interdependence” with US Big Tech solutions or replacing a foreign dependency with a domestic one for the time being. But it must also entail building the technical capacity, open infrastructure, and participatory institutions needed to preserve genuine choice across the entire AI stack. Open source will continue to be vulnerable to commercial absorption, and efforts to counter that must become more robust.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Build a Sports Concierge Agent with Chester Ismay</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/#respond</comments>
				<pubDate>Mon, 21 Sep 2026 17:01:35 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19754</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Turn sports schedules and personal preferences into one weekly recommendation]]></custom:subtitle>
		
				<description><![CDATA[Chester Ismay, a data science educator and AI consultant, created schedule viewers to keep up with the sports he follows, including the WNBA, NFL, NBA, and Premier League. But he still had to decide which games deserved his attention each week. In this episode of Zero to Agent in 30 Minutes, Chester built a sports [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Chester Ismay, a data science educator and AI consultant, created schedule viewers to keep up with the sports he follows, including the WNBA, NFL, NBA, and Premier League. But he still had to decide which games deserved his attention each week.</p>



<p class="wp-block-paragraph">In this episode of <em>Zero to Agent in 30 Minutes</em>, Chester built a sports concierge agent to surface the games he should watch. It reads his preferences and current schedules, then sends a weekly summary to his phone. He took the audience <a href="https://ismayc.github.io/never-miss-a-game/" target="_blank" rel="noopener">through the setup</a>, which combines a prompt file, limited tool permissions, a schedule, and notifications.</p>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/W1_p5WRstFw?si=YRSKXXwW7H1h6tEw" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<h2 class="wp-block-heading"><strong>How to create your own sports concierge</strong></h2>



<ol class="wp-block-list">
<li><strong>Define your preferences.</strong> Chester started with a structured preferences file that identifies the teams he follows and adds context about why he follows them. He built <a href="https://ismayc.github.io/never-miss-a-game/preferences-editor.html" target="_blank" rel="noopener">a simple web interface</a> for editing those preferences rather than working directly with the underlying JSON.</li>



<li><strong>Give the agent access to current schedule data.</strong> His existing sports viewers pull schedule information from sources such as ESPN and store it in repositories on GitHub. A read-schedules tool, run with Node, pulls the latest schedule files and combines them with Chester’s preferences, giving the agent the information it needs without requiring it to search for each game on its own.</li>



<li><strong>Write the agent’s policy.</strong> Chester spent most of the walkthrough on <a href="https://raw.githubusercontent.com/ismayc/never-miss-a-game/refs/heads/complete/concierge.md" target="_blank" rel="noopener">CONCIERGE.md</a>, the file that defines the agent’s job. The policy sets the goal, identifies the data sources, and defines the rules for deciding which games to recommend. It also specifies how to handle finished tournaments and duplicate matchups, along with the expected output and delivery format. Chester noted that when the output misses the mark, he goes back to the policy and adds more detail to the instructions or adjusts them to better match with the goals of the project.</li>



<li><strong>Run the agent with limited permissions.</strong> Chester used Claude Code to read the files, apply the policy, and generate the weekly recommendations. He configured permissions so the agent worked only with the files and tools required for the task.</li>



<li><strong>Schedule delivery and check the results.</strong> Chester used <a href="https://www.launchd.info/" target="_blank" rel="noopener">launchd</a> on his Mac to run the concierge every Wednesday and <a href="https://ntfy.sh/" target="_blank" rel="noopener">ntfy</a> to send the result to his phone. He checks the recommendations against the underlying schedules and uses tests and multiple data sources to catch errors. Time zone handling required another adjustment. Games could fall on the wrong day when the system defaulted to UTC, so Chester added explicit time zone instructions.</li>
</ol>



<h3 class="wp-block-heading"><strong>Coming up next</strong></h3>



<p class="wp-block-paragraph">Next week, AI engineer Sajal Sharma gives an agent its own computer in the cloud using services such as E2B and Scrapybara. He’ll demonstrate how sandboxing lets an agent install packages, run code, drive a browser, and control a remote desktop.</p>



<p class="wp-block-paragraph"><em>Follow along with</em> Zero to Agent in 30 Minutes <em>on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O’Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Post-training Process OpenAI Used for ChatGPT</title>
		<link>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/</link>
				<comments>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/#respond</comments>
				<pubDate>Mon, 21 Sep 2026 10:55:18 +0000</pubDate>
					<dc:creator><![CDATA[Sharon Zhou]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19751</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-post-training-process-OpenAI-used-for-ChatGPT.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-post-training-process-OpenAI-used-for-ChatGPT-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Walk through the pipeline that made ChatGPT such an improvement over GPT-3]]></custom:subtitle>
		
				<description><![CDATA[This is the third post in a four-part series about post-training. If you missed them, read part 1 and part 2. The final post, on implementing your own pipeline, will be coming October 7. Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>This is the third post in a four-part series about post-training. If you missed them, read <a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">part 1</a> and <a href="https://www.oreilly.com/radar/the-two-pillars-of-post-training-reinforcement-learning-and-supervised-fine-tuning/" target="_blank" rel="noopener">part 2</a>. The final post, on implementing your own pipeline, will be coming</em> <em>October 7</em>.</p>
</blockquote>



<p class="wp-block-paragraph">Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished in the frontier models you know and love.</p>



<p class="wp-block-paragraph">Remember our prompt “<a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">Why do people like golden retrievers?</a>” GPT-3 would often answer nonsensically. But that all changed in November 2022, with the launch of ChatGPT. Now “Why do people like golden retrievers?” actually returned a reasonable response like “Because they are affectionate, patient, and make excellent family pets” no matter who was typing (with no weird formatting tricks to consider). Anyone who could send a text message could get a response back on any topic.</p>



<p class="wp-block-paragraph">I’ll cover some of those behavior changes below, then take you through ChatGPT’s training pipeline as described in OpenAI’s InstructGPT paper.</p>



<h2 class="wp-block-heading">Conversational and helpful</h2>



<p class="wp-block-paragraph">The most visible impact of post-training is models that can chat with you and hold a relatively long conversation. This sounds simple, but it’s not.</p>



<p class="wp-block-paragraph">Being conversational means more than responding to a question with an answer. The model needs to recognize when a question is ambiguous and ask for clarification or make the right assumptions in a quick response. It should adjust its tone and detail level to the context, for example being brief for a quick factual question, but thorough for a learning-oriented one. The model should be coherent across multiturn conversations without losing the thread. It also needs to handle messy real-world inputs: You attach a giant PDF and ask it to find one specific clause, and it should either find it or tell you it can’t, not hallucinate an answer.</p>



<h2 class="wp-block-heading">Safety and alignment</h2>



<p class="wp-block-paragraph">Post-training is also the primary mechanism for making models safe. Safety in this context means a few things:</p>



<ul class="wp-block-list">
<li>Refusing to generate harmful content (like instructions for creating weapons, when asked)</li>



<li>Avoiding biased or discriminatory outputs</li>



<li>Not making up information when unsure (hallucination reduction)</li>



<li>Respecting user privacy</li>
</ul>



<p class="wp-block-paragraph">However, you can define safety rules in whatever way you want and teach the model to abide by them, within the limits of what your reward signals can capture. If you think cats are unsafe, because you’re a dog person, you can teach the model that in post-training—as long as you can properly encode that into a reward signal.</p>



<p class="wp-block-paragraph">Safety is often in tension with helpfulness. On one extreme, a model that’s too conservative will refuse reasonable requests, something that has frustrated many users. On the other extreme, a model that’s too permissive will comply with harmful ones. Navigating this trade-off is difficult. Ultimately, it comes down to determining where to draw the line, which is (as of today) a human decision within labs. Post-training is the tool to implement wherever the line is drawn.</p>



<h2 class="wp-block-heading">Tool use and function calling</h2>



<p class="wp-block-paragraph">Tool use is one of the most practically important capabilities enabled by post-training. Tools include search engines, APIs, calculators, databases, and code interpreters. Being able to hit a search engine alone allows the model to not hallucinate, given its own knowledge cutoff. Tools are extremely useful ways for models to interact with the world, and are fundamental components in building agents.</p>



<p class="wp-block-paragraph">Tool use is a set of new behaviors. The model needs to recognize when a user’s request would benefit from an external tool. It needs to know which tools are available to it, and not hallucinate a tool. It needs to formulate a correct API call with the right parameters. It needs to interpret the results that come back and incorporate them into a natural language response. It needs to do all of this seamlessly, without the user needing to know the details of the underlying tool.</p>



<p class="wp-block-paragraph">This is taught almost entirely through SFT, at least initially. The training data includes many examples of conversations where the model correctly decides to invoke a tool that it has access to, constructs the right call, and processes the result. RL can further improve tool use by rewarding the model for correct tool invocations and penalizing unnecessary or incorrect ones.</p>



<p class="wp-block-paragraph">As an example of tool use, let’s say you’re building a veterinary appointment scheduling assistant. A user asks: “My golden retriever has been limping since yesterday. Can I see Dr. Patel this afternoon?” A pretrained model might generate plausible but fictional appointment times. A post-trained model with tool use instead calls the clinic’s scheduling API, checks Dr. Patel’s availability, and responds: “Dr. Patel has an opening at 3:15pm today. I’ve tentatively held it for you. Should I confirm?” The model needed to decide if the user’s intent was urgent, select the right tool, construct the API call with the right veterinarian and time constraints, and present the result conversationally.</p>



<p class="wp-block-paragraph">Tool use has expanded through the <a href="https://www.anthropic.com/news/model-context-protocol" target="_blank" rel="noopener">Model Context Protocol</a> (MCP), a lightweight standard for connecting models to external services like Gmail, GitHub, or a company’s internal databases. Rather than building custom integrations for each tool, MCP provides a standard interface that any API can plug into, and different frontier models have now included learning MCP in their post-training recipes. Agentic frameworks take this further by allowing models to chain multiple tool calls together to accomplish common multistep tasks more easily.</p>



<h2 class="wp-block-heading">Reasoning (“thinking”)</h2>



<p class="wp-block-paragraph">Reasoning models, or models that are trained to “think” before they answer, are an exciting result of post-training. Rather than producing an immediate response, these models generate an internal chain of thought, working through the problem step-by-step, before arriving at a final answer. As a result, their answers are more often correct than nonreasoning models that might guess at an answer.</p>



<p class="wp-block-paragraph">This capability has an interesting relationship with pretraining and post-training. The raw ability to reason is latent in pretrained models; they’ve been trained on text that includes mathematical proofs, logical arguments, scientific analyses, and code with comments explaining the logic. But pretrained models don’t default to reasoning. They default to pattern-matching, which often produces plausible-looking but incorrect answers.</p>



<p class="wp-block-paragraph">Reasoning models dramatically outperform standard models on tasks that require multistep logic: mathematical problem-solving, complex coding, scientific analysis, and planning. The improvements are not incremental. On the 2024 AIME exam, GPT-4o was only able to get 12% of problems correct on average. OpenAI’s o1 reasoning model solved 74% off the bat, with a single attempt. With 1,000 attempts and a learned scoring function to rerank the attempts, <a href="https://openai.com/index/learning-to-reason-with-llms/" target="_blank" rel="noopener">it reached 93%</a>, a result placing it among the top 500 students who took the AIME math exam in the US.</p>



<p class="wp-block-paragraph">More capable reasoning requires more compute, both during training and at inference time. Scaling laws meet post-training. Models that have learned to spend more inference (test-time) compute on reasoning tend to reach better answers and therefore exhibit higher intelligence. For some frontier reasoning models, the RL post-training phase uses as much compute as the entire pretraining phase.</p>



<p class="wp-block-paragraph">The cost is not only in post-training compute but also in inference (test-time) tokens and latency. Reasoning takes up a lot of tokens and can result in a longer time to get a response back to the user. But the type of request matters. For a quick factual question, you don’t need reasoning. For a complex technical problem, the extra latency is well worth it. This is something that model providers can modulate during post-training.</p>



<h2 class="wp-block-heading">The classic ChatGPT pipeline</h2>



<p class="wp-block-paragraph">As I mentioned above, the first post-training pipeline that captured global attention was ChatGPT’s, and it drew on the pipeline described in the <a href="https://arxiv.org/abs/2203.02155" target="_blank" rel="noopener">InstructGPT paper</a>. While modern systems use more advanced approaches today, this classic pipeline remains the conceptual foundation for nearly all alignment methods.</p>



<p class="wp-block-paragraph">The pipeline has three stages, each building on the previous one:</p>



<ol class="wp-block-list">
<li>Supervised fine-tuning (SFT) on human demonstrations</li>



<li>Training a reward model on human preference comparisons</li>



<li>Reinforcement learning with human feedback (RLHF) to optimize the main model using the reward model</li>
</ol>



<h3 class="wp-block-heading">Stage 1: SFT on demonstrations</h3>



<p class="wp-block-paragraph">The first stage is straightforward and teaches the model to follow instructions and behave like an assistant.</p>



<p class="wp-block-paragraph">OpenAI contracted ~40 human labelers to label their data, and they were careful to filter for people who were good at identifying harmful outputs. The labelers had to write ideal responses to prompts. But what’s interesting is that the prompts came from two sources: (1) prompts submitted by real users through the OpenAI API and (2) prompts that labelers wrote themselves. The users had to write prompts too, because these were the days before ChatGPT. There weren’t that many real users with instruction-like prompts through the API to collect.</p>



<p class="wp-block-paragraph">The prompts were diverse and mostly in English. There’s also an extensive data cleaning pipeline to remove duplicates and remove sensitive PII (personally identifiable information). Importantly, they split the training, validation, and test sets by human labeler. This is to avoid data leakage that could happen within a single user’s data between training and validation/testing.</p>



<p class="wp-block-paragraph">The resulting SFT dataset had ~13,000 prompts, all with human-labeled responses. The base model was GPT-3 at the time, a pretrained model without any post-training. Using SFT, they trained GPT-3 for 16 epochs, which was effective for the final RLHF model. This was interesting, because for the SFT stage alone, the model overfit after just 1 epoch, but ultimately SFT was an intermediate stage so they picked the best checkpoint for the final RLHF model. They also mixed in 10% pretraining data during this phase, because it would help the next RL phase.</p>



<p class="wp-block-paragraph">At this point, this SFT model could already be pretty useful: It could have a conversation and follow instructions, which is leaps and bounds beyond the pretrained GPT-3 checkpoint.</p>



<h3 class="wp-block-heading">Stage 2: Preference data and reward modeling</h3>



<p class="wp-block-paragraph">This next stage is training the reward model. The reward model needs to grade millions of responses during RL training. In the original method, OpenAI’s team mainly trained the reward model on responses from the SFT model. However, as the policy model is trained in the RL loop and generates new, and likely better, responses from its evolving checkpoints, the reward model needs to stay robust. As a result, they also continually updated the reward model using responses from new RL checkpoints over time.</p>



<p class="wp-block-paragraph">To train the reward model in InstructGPT’s RLHF pipeline, OpenAI needed pairwise comparisons of two model responses from one prompt, and a label for which one is better. For example, given “What’s 2+2?” and the responses are “4” and “Yes,” the label should say “4” is better than “Yes.” Note again that these are responses from the SFT model (and later, the RL-ed models during the RL training loop), not the pretrained base model. So labeling can only happen after you’ve SFT-ed your model. If you need to retrain that model, you likely need to relabel to make sure the reward model is trained on the right distribution of data pairs.</p>



<h4 class="wp-block-heading">Reward model training</h4>



<p class="wp-block-paragraph">The reward model was small at 6B parameters, for both efficiency and stability, and included a head that outputted a scalar reward. They had tried multiple sizes, but found this was more stable than using the original 175B main model. It was also more compute efficient, as the reward model would take up extra compute, for both inference and training, on top of training the main model itself. More recently, reward models have become a lot larger, but note that they don’t have to be the same model or same size model as the main model.</p>



<p class="wp-block-paragraph">To train the model, the loss was a cross-entropy loss that represented the log odds that someone would prefer one option over the other, in the pairwise comparison. This was done by taking the difference between the rewards of the preferred and unpreferred options. So in the example “What’s 2+2?,” if the reward model correctly assigns “4” a high reward and “Yes” a small reward, then the difference would be high and positive, and the loss would be small. However, if the reward model were to incorrectly assign “Yes” a higher reward than “4,” the difference would be high and negative, and the loss would be huge—discouraging it from outputting this result again.</p>



<p class="wp-block-paragraph">One of the big challenges in training the reward model was overfitting, and OpenAI found that training for only 1 epoch would help prevent that.</p>



<h4 class="wp-block-heading">Reward model data</h4>



<p class="wp-block-paragraph">The simplest way to get preference pairs is to generate two responses per prompt and have a labeler tell you which one was better. To make more efficient use of each prompt, instead the model would generate not 2 but 4–9 different responses per prompt that human labelers would rank from best to worst.</p>



<p class="wp-block-paragraph">Rankings can be transformed into pairwise comparisons, so it was an efficient way to collect those preference pairs. A ranking of N responses yields N-choose-2 pairs. For example, a ranking of 4 responses results in 6 pairs, a ranking of 9 results in 36 pairs. That means with 33K prompts and 4–9 responses ranked per prompt, there would be 200K–1.2M pairwise comparisons used to train a separate reward model. That’s a lot of data, from relatively efficient data labeling.</p>



<p class="wp-block-paragraph">This is a relatively efficient use of human annotations. Just compare it to SFT. It’s easier, cheaper, faster, and more reliable (higher agreement between people) than writing good responses from scratch, so this stage of human labeling wasn’t as tedious as in SFT.</p>



<p class="wp-block-paragraph">However, using the pairs from rankings wasn’t straightforward in training. The reward model would overfit if they mixed the pairs randomly, even in just 1 epoch, because the pairs for a single prompt were highly correlated with each other. So instead, they would train all the pairs from the same prompt as one element in a batch, and normalize it. This was also computationally more efficient to run and score all the N responses at once together, e.g., just score 9 times and reuse those calculations in this pass, rather than 36 times for each pair if mixed into the dataset.</p>



<p class="wp-block-paragraph">A quick note on terminology. This data is often called <em>preference data</em>, because it’s about collecting human preferences. The reward model can also be called a <em>preference model</em>.</p>



<h3 class="wp-block-heading">Stage 3: RLHF (reinforcement learning with human feedback)</h3>



<p class="wp-block-paragraph">At this point, you have an SFT model that can follow instructions, and a reward model that can score responses. The goal of RL is to continue training the SFT model to produce responses that the reward model scores highly. If the reward model is any good, the resulting model will produce responses that humans would prefer.</p>



<p class="wp-block-paragraph">The RL algorithm used was PPO. As you learned previously about RL terminology, the SFT model is the “policy” that takes actions (generating tokens) in an environment (the conversation). The reward model provides the reward after the policy generates a complete response, and a critic model calculates the expected reward, a baseline estimate that offers a more stable overall reward signal in training.</p>



<p class="wp-block-paragraph">Here are the critical steps. I’ve covered some of them before and will dive into others in detail later on in this section.</p>



<ol class="wp-block-list">
<li><strong>Sample a prompt</strong>. The prompt comes from the dataset of 31,000 prompts that were gathered from users organically using the API. No need for human labels.</li>



<li><strong>Generate a response</strong>. Then, the current policy generates a response. The current policy is the SFT model in the beginning, but as the policy updates, it’s a new model that generates responses to be graded. A single prompt-response pair is called a “rollout.” In practice, this all happens in a batch of rollouts.</li>



<li><strong>Calculate the reward</strong>. The reward model grades the full response with a reward. For every token position in the full response, they subtract a per-token KL divergence penalty between the current policy and the original SFT model. The per-token KL penalty and the reward model score added at the final token make up the per-token reward signal.</li>



<li><strong>Calculate the advantage</strong>. The critic estimates the expected reward for the full response, at each token position in generation (so with partial knowledge of the full response). The critic’s expected rewards and the per-token reward signals are combined, using an algorithm called GAE (<a href="https://arxiv.org/abs/1506.02438" target="_blank" rel="noopener">Generalized Advantage Estimation</a>), to estimate how much better the reward was compared to expected. This is the advantage. In InstructGPT, the critic was initialized with the same weights as the reward model, giving it a head start on estimating expected reward. </li>



<li><strong>Update the policy</strong>. Then, PPO updates the policy model’s weights with the advantage, pushing it towards rollouts with higher advantage and away from ones with lower advantage.</li>



<li><strong>(Optional) Mix in pretraining data and the pretraining objective in the policy model’s loss function</strong> to prevent catastrophic forgetting. </li>



<li><strong>Update the critic</strong> to better predict expected future rewards at each token position, by using the actual per-token rewards (from the reward model and KL penalty) as the targets in training.</li>



<li><strong>(Optional) Update the reward model</strong>. Collect new ranking data on the current best policy and train a new reward model. In practice, OpenAI did collect some data from the PPO models, but most was from the original SFT model.</li>



<li><strong>Repeat!</strong> This RL loop repeats over many prompts and many updates, with the latest policy always generating its responses.</li>
</ol>



<h4 class="wp-block-heading">The KL penalty</h4>



<p class="wp-block-paragraph">If you just let the model maximize the reward model’s score with no constraints, it finds weird, degenerate outputs that exploit quirks in the reward model to get high scores without actually being good responses. This is called “reward hacking,” and it’s one of the central problems in RLHF.</p>



<p class="wp-block-paragraph">To address this, OpenAI added a KL divergence penalty between the RL policy and the original SFT model (“reference policy”) in the reward calculation. In AI, KL divergence is a common method of measuring how different two probability distributions are. In this case, it would measure how different the RL policy is from the old policy and penalize being too far from it, essentially telling the model: You can optimize for higher reward, but you can’t drift too far from where you started. If the RL model starts producing outputs that look nothing like what the SFT model would produce, the penalty helps to pull it back by making the reward for those outputs lower.</p>



<p class="wp-block-paragraph">The total reward for a response becomes the reward model’s score minus the KL divergence from the SFT model, with a coefficient term that weighs how much to care about the KL divergence. If the coefficient is too low, you’re saying that you don’t need to penalize drift from the reference policy, and you’ll get reward hacking. Too high and the model barely changes from the SFT checkpoint.</p>



<p class="wp-block-paragraph">They also mixed in a significant amount of pretraining data during the RL phase, adding a pretraining loss alongside the RL objective. This was to prevent the model from degrading on general tasks from pretraining, like knowledge recall or coherent long-form creative text, as it optimized for reward. This is sometimes called the “<a href="https://arxiv.org/abs/2309.06256" target="_blank" rel="noopener">alignment tax</a>,” where you trade-off alignment for general capabilities, a type of “<a href="https://arxiv.org/abs/1612.00796" target="_blank" rel="noopener">catastrophic forgetting</a>.” This is an active area of research.</p>



<p class="wp-block-paragraph">So the final RL objective combined three things: (1) maximize the reward model’s score on prompted responses, (2) stay close to the SFT model via the KL penalty, and (3) maintain performance on pretraining data. This means improving on the things humans care about without losing what the model already knew how to do from SFT.</p>



<h4 class="wp-block-heading">Practical details</h4>



<p class="wp-block-paragraph">The critic reduces noise and makes training stable enough to make PPO work practically. In practice, OpenAI initialized the critic from the 6B-parameter reward model, since it’s already trained to predict reward and gives the critic a head start as it is further trained in the RL loop.</p>



<p class="wp-block-paragraph">The RL training was computationally expensive and involved running several models simultaneously: the policy model (the main model being trained), the critic (estimating the reward as a baseline, also being trained), the reward model (grading responses), and a copy of the SFT model (for computing KL divergence).</p>



<p class="wp-block-paragraph">That’s four models in memory at once. That’s a lot of GPU memory, especially when the policy and SFT models are 175B parameters! In addition to weights, the policy and value models also needed their gradients, optimizer states, and cached activations for backpropagation because they were being trained, which can actually multiply the per-model memory cost by 3-4 times. This is another reason the reward and critic models were kept at 6B.</p>



<p class="wp-block-paragraph">PPO also requires generating fresh rollouts during training, which is much slower than SFT where you already have all the data upfront. Each PPO training step also requires grading each rollout with the reward model, computing advantages with the critic, and updating both the policy and the critic. All these moving parts make the system harder to tune and debug compared to SFT, and harder to parallelize than pretraining.</p>



<p class="wp-block-paragraph">Hyperparameters like learning rate, the KL penalty coefficient term, the number of rollouts per batch, and the clipping ratio all matter and interact with each other. The whole system depends on the quality and representativeness of your human annotations. Many RL training runs fail or produce degenerate results.</p>



<h2 class="wp-block-heading">Getting it right</h2>



<p class="wp-block-paragraph">Getting all this right takes significant engineering effort and experience, but the first step is deeply understanding the pieces. Modern methods have addressed many of these issues, but the ideas from InstructGPT remain the foundation of post-training today.</p>



<p class="wp-block-paragraph">Ultimately, human evaluators compare all the models. People preferred the RLHF model’s outputs over the SFT model’s, and the SFT model’s over base GPT-3’s. Each stage of the pipeline added a large improvement in the model’s response quality.</p>



<p class="wp-block-paragraph">RLHF was extremely effective. In experiments, human evaluators preferred even a tiny 1.3B parameter RLHF model over the 175B parameter SFT model, most of the time. That’s a model over 100x smaller, trained with RL, beating a much larger model trained only with SFT. This made a strong case that how you train matters as much as how big your model is. Overall, the largest RLHF model still beat the smaller RLHF model.</p>



<p class="wp-block-paragraph">The RLHF model was also better at following explicit constraints in instructions, less likely to produce harmful outputs, and hallucinated less, though it didn’t eliminate hallucinations as you may remember when you first used ChatGPT (and even now).</p>



<p class="wp-block-paragraph">One caveat worth noting: The labelers who evaluated the final model were the same population who created the training data. When they tested with held-out labelers who hadn’t been involved in data creation, preferences for the RLHF model were still positive but less dramatic. The model was, to some degree, optimized for the preferences of a specific group of people. This means if you create the data to follow your preferences, the model will optimize for those.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>How to Get from AI-Assisted to AI Native</title>
		<link>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/</link>
				<comments>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/#respond</comments>
				<pubDate>Fri, 18 Sep 2026 21:38:53 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19736</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-as-an-enterprise-operating-system.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-as-an-enterprise-operating-system-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[AI as an Enterprise Operating System, a keynote from Ai4 2026]]></custom:subtitle>
		
				<description><![CDATA[When considering the history of AI, Richard Sutton observed that brute force and compute scale has always trumped human expertise, and when you look for it, you can see this “bitter lesson” play out throughout tech history. In his keynote at Ai4 2026, Tim O’Reilly explains why grappling with the bitter lesson is the forge [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/5EcTnkCY5ww?si=Ft-Khy2zizSbj1g4" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph"> </p>



<p class="wp-block-paragraph">When considering the history of AI, Richard Sutton observed that brute force and compute scale has always trumped human expertise, and when you look for it, you can see this “bitter lesson” play out throughout tech history. In his keynote at Ai4 2026, Tim O’Reilly explains why grappling with the bitter lesson is the forge of effective AI corporate strategy, as companies figure out what to embrace and what to let go of. Drawing on his <a href="https://www.oreilly.com/radar/ai-as-an-enterprise-operating-system/" target="_blank" rel="noopener">recent conversations</a> with Trail of Bits CEO Dan Guido, Tim argues that AI&#8217;s business impact actually hinges on organizational adoption—the hard, unglamorous work of restructuring workflows, data, and incentives around what AI can do. Trail of Bits has modeled that process and <a href="https://trailofbits.com/?item=https-github-com-trailofbits-publications-blob-master-presentations-how-20we-20m" target="_blank" rel="noopener">documented it in a playbook</a> other companies can use. Here, Tim shares some of the practices, like capability ladders, shared config repos, and company-wide hackathons, that helped Trail of Bits make AI a structural component of its business. This doesn’t mean that AI-native companies “sit back and let the progress of AI carry us forward.” Human expertise still matters, and it’s often the differentiator that helps organizations rise above their competitors. As Tim concludes, “The world is full of great problems. And so if AI takes away and makes easy something small, celebrate it and go work on something big with the new powers that we’ve been given.”</p>



<h2 class="wp-block-heading">Takeaways</h2>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=egYeWccrMdXhFu-N&amp;t=153" target="_blank" rel="noopener">02.33</a></strong> <strong>The bitter lesson is real, and it can catch any of us.</strong><br>The bitter lesson is Richard Sutton’s contention that human expertise doesn’t really matter, that it will eventually be outmatched by computing scale. O’Reilly’s <em>Whole Internet User’s Guide &amp; Catalog</em> was the first catalog of websites and the first site on the web to have advertising. It grew into Global Network Navigator, which was the first web portal. But O’Reilly’s products were manually curated. Yahoo came along and expanded on these ideas, but O’Reilly and Yahoo were both beaten by Google, which simply threw a bunch of compute at the problem. Now ChatGPT has changed the game again.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=E8D9kntfeiYL_plj&amp;t=387" target="_blank" rel="noopener">06.27</a></strong> <strong>AI-native workflows require a different mindset.</strong><br>When O’Reilly set out to develop a product that assessed learners’ capabilities and gave them a skill path to level up, the team used AI as an assistant, to write quiz questions, for instance. But LLM chatbots can already identify skills when given context about a developer. Evolving toward an AI-native skill path builder meant reconceptualizing the product as a more interactive experience that reflects where capabilities are today. However, even the most well-thought-out workflow can be hindered by gaps in access or knowledge. As Trail of Bits CEO Dan Guido says, “You have to build a system in which expertise compounds.”</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=wPVC_gop7qBWgaGO&amp;t=628" target="_blank" rel="noopener">10.28</a></strong> <strong>AI adoption is a human problem.</strong><br>Moving up the framework for AI adoption from AI-assisted to AI-augmented to AI-native isn’t just a technical challenge. It’s psychological. Only 5% of Dan’s staff was actually on board when he started the transformation; 70% were just quietly going through the motions, and 20% were actively resistant. He traces this to a handful of biases: self-enhancing bias, opacity, intolerance for imperfection, and above all, identity threat, the fear that AI won&#8217;t just replace the work someone does but who they are. Getting teams on board requires the organization to reframe AI as a tool that enhances identity, not something that will take it away.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=Ro9rr9iaMN9T8_M-&amp;t=1004" target="_blank" rel="noopener">16.44</a></strong> <strong>A status ladder helps team members understand where they&#8217;re at and where to focus next. Hackathons compound that knowledge across the company.</strong><br>Trail of Bits has a three-level status ladder: not engaged with AI or actively resisting it, experimenting with AI, and building AI that strengthens the organization&#8217;s overall capability. Level zero isn&#8217;t treated as a skill gap. It&#8217;s treated as working against the company&#8217;s goals, and the other two levels get a more detailed capability matrix broken out by department, since what a security auditor does with AI looks nothing like what someone in accounting does. O&#8217;Reilly is building its own version of this, drawing on the technical and business skill data it already has across its platform. Trail of Bits runs a hackathon every two months, each with a stated objective and learning goals announced a week ahead. Success is measured not by what got shipped but by where people land on the capability ladder afterward. Then the work gets fed into a shared skill repo, giving the entire company a set of reusable artifacts, and what one hackathon turns up becomes something the next one can build on.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=rHHKVmNEYeiOm8ts&amp;t=1477" target="_blank" rel="noopener">24.37</a></strong> <strong>Turn scar tissue into infrastructure.</strong><br>Drew Breunig talks about the problem of <a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">prompt debt</a>: prompts that grow more complex and more tuned to one specific model until they&#8217;re no longer portable. Trail of Bits flattens this complexity by turning every failure into a global, copy-pasted fix hosted in a company-wide repository. They’ve also standardized the safety net, with sandboxes for different needs and a seven-day cooldown on every new package from outside that gets installed—rules the whole company follows. To make this all work, employees need the chance to try things out and iterate on their failures. Dan says the only real mistake he made was not giving people enough unstructured time to experiment.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=AEUjP4JHTL2c0jn-&amp;t=1964" target="_blank" rel="noopener">32.44</a></strong> <strong>Human expertise still matters.</strong><br>AI can make companies more productive, but it’s not a magic weapon. It’s a <a href="https://www.oreilly.com/radar/writing-with-ai/" target="_blank" rel="noopener">medium</a> that people can use to share or extend their unique expertise and perspective. O’Reilly’s mission is to share the knowledge of innovators: You can think of the company as a matching marketplace for people who have expertise and people who need it. Agents offer a valuable new means of getting that expertise to customers in the tools they’re using to make business decisions. O&#8217;Reilly CTO Andrew Odewahn has noted that faster local decision-making has splintered central planning, so it&#8217;s harder than ever to get the big-picture view a good corporate decision needs. O&#8217;Reilly&#8217;s Expert MCP server lets customers access our content and use it to increase organizational intelligence. For instance, you can ask an AI tool to analyze a team’s workload and write a hiring case based on how O&#8217;Reilly&#8217;s own experts would review the request, and you’ll get a grounded argument with solutions authenticated by citations from actual practitioners. O&#8217;Reilly is building this capability into an organization-wide grounding layer it calls O&#8217;Reilly Expert Intelligence. It’s in beta now, and you can <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noopener">check it out</a>.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 112/113 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-09-29 18:05:49 by W3 Total Cache
-->