<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Thu, 24 Sep 2026 10:54:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>When Software Subscriptions Become Public Policy</title>
		<link>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/</link>
				<comments>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/#respond</comments>
				<pubDate>Thu, 24 Sep 2026 10:54:23 +0000</pubDate>
					<dc:creator><![CDATA[Steve Hayes]]></dc:creator>
						<category><![CDATA[Executive Briefing]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19795</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/When-software-subscriptions-become-public-policy-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave? There was something wonderfully circular about hearing Dan Gookin tell me that, after he [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>My conversation with Dan Gookin, the original For Dummies author and now mayor of Coeur d’Alene, Idaho, started as a publishing reunion. It ended with a much larger question: What happens when the software you rent becomes infrastructure you can’t leave?</em></p>



<p class="wp-block-paragraph">There was something wonderfully circular about hearing Dan Gookin tell me that, after he was first elected to public office, he bought a copy of <em>Robert’s Rules For Dummies</em>. Dan wrote <em>DOS For Dummies</em>, the 1991 book that launched the <em>For Dummies</em> series, and went on to write more than 180 technology books with over 12 million copies in print. I spent nearly three decades acquiring books in that program before joining O’Reilly this summer. Dan saw the news of my new role and reached out. We caught up on publishing, books, and AI. The part of the conversation that stuck with me had nothing to do with any of those topics. Dan is now the mayor of Coeur d’Alene, Idaho, and he’s begun thinking about the city’s software the same way he’s been thinking about his own.</p>



<p class="wp-block-paragraph">Dan was elected to the Coeur d’Alene City Council in 2011 and became mayor in 2025. He was quick to draw a distinction between the two jobs. “Council, you can be a little bit more extreme,” he told me. “It’s more rhetoric based. This is administrative. It’s management.” A council member can object to a budget line. A mayor has to make sure the systems that line funds keep working. For a city of nearly 58,000 people, that covers everything from the police network to the water bill.</p>



<h2 class="wp-block-heading"><strong>Cutting the leash</strong></h2>



<p class="wp-block-paragraph">Near the end of our conversation, Dan mentioned a project he’s been working on personally. He’s begun weaning himself off Microsoft and what he calls an increasingly subscription-based, cloud-based, AI-heavy environment, moving his computing to Linux and running his own servers. This move isn’t entirely ideological. There is money attached. Dan told me his Adobe subscription for software he uses to produce content costs about $800 a year. He believes he can replace most of it on Linux. What’s surprised him is how much he can replace. “It’s interesting to see how quickly it can be substituted,” he said, and how quickly he can “cut that chain.&nbsp;.&nbsp;.or cut that leash.”</p>



<p class="wp-block-paragraph">Dan sees a similar economic model when he gets to work. “Our software subscription just for our city, our size, is $700,000 a year and growing,” he told me. The example he kept returning to was the city’s financial software. “We used to own our finance software,” he explained. “Now the finance software is leased and it’s cloud-based.” Then he asked the question that ought to appear in a lot more software procurement meetings: “So we don’t even own our own data?”</p>



<p class="wp-block-paragraph">His concern isn’t literally that the city has forfeited legal ownership of its records. It’s about practical control. What happens if the vendor is hacked, or the city decides to switch? Can it get all its data back? “Do they just give you a binary dump or do they actually give you the data,” he asked. That question has stuck with me since we chatted.</p>



<h2 class="wp-block-heading"><strong>Reversibility is an architectural property</strong></h2>



<p class="wp-block-paragraph">Cloud contracts have a word for part of this issue. In cloud services agreements, <strong>reversibility</strong> refers to a customer’s ability to retrieve its data and unwind a vendor relationship. The United Nations Commission on International Trade Law (UNCITRAL) even includes it in its glossary of cloud contract terms. That idea belongs in municipal software evaluation too, not as an exit clause buried in a contract but as a standing column in the evaluation spreadsheet, next to features, price, and security.</p>



<p class="wp-block-paragraph">When you evaluate software, you compare features, price, implementation time, security, and, increasingly, AI capabilities. But what’s the exit cost? Can an organization export its data in a documented, useful format that another application can consume? How much institutional knowledge has quietly migrated from the organization to the vendor? And if the vendor raises prices sharply at renewal, gets acquired, or simply stops serving your needs, how long would it take to leave?</p>



<p class="wp-block-paragraph">These questions are properties of the system and subscription software has raised their stakes. When software came in a box, skipping an upgrade didn’t make the version you owned disappear. That world of proprietary formats and dominant platforms had plenty of lock-in. But possession still meant something.</p>



<p class="wp-block-paragraph">Dan and I talked about how alien that world now seems. Software once arrived with manuals. Today it may not even arrive. You authenticate to it. If you don’t know how to do something, you ask an AI instead of consulting a manual. That change highlights a step from possessing tools to maintaining permission to use them. For an individual, that might mean Photoshop is a monthly subscription now. For an organization, recurring access can become an architectural dependency. For a government, that dependency is borne by taxpayers.</p>



<h2 class="wp-block-heading"><strong>Should cities write their own software?</strong></h2>



<p class="wp-block-paragraph">Dan takes the argument one step further. “For $700,000 a year,” he said, “we could hire a couple of programmers just on contract and have them code our own stuff and then we own it again.” I’m not convinced that math works out. Two programmers likely can’t effectively reproduce a mature municipal financial system, endpoint protection, records management, and specialized public-safety applications, let alone the compliance work and vendor support that come with a modern city’s technology stack. AI could help accelerate the process, but liability and security issues surrounding AI-enabled development likely add more risk than a government is willing to accept in the name of software ownership.</p>



<p class="wp-block-paragraph">Building software also creates its own long-term bills, ones that don’t fit easily within most city budgets. Code must be maintained, security vulnerabilities patched, and staff and frameworks eventually replaced. An application written in-house can become every bit as difficult to escape as one bought from a vendor. On the flip side, the headaches of replacing a “good enough” proprietary system with a packaged one that fits most, but not all, of an organization’s needs can outweigh the cost of keeping the old system running. The familiar build-versus-buy analysis exists for good reasons.</p>



<p class="wp-block-paragraph">But Dan’s question still matters, even if his proposed fix isn’t right for every case. At what point does the cost of renting capability justify rebuilding some capability of your own? Perhaps more importantly, which capabilities should an organization insist on controlling, even if renting them is cheaper? There is no universal answer, including “the vendor handles it.”</p>



<h2 class="wp-block-heading"><strong>This isn’t an argument against the cloud</strong></h2>



<p class="wp-block-paragraph">It would be easy to turn Dan’s experiment into a familiar prescription. Move to Linux. Embrace open source. Bring everything back on premises. Escape the cloud. That’s too simple. Cloud and SaaS products solve real problems, shifting maintenance to specialists and giving a city of 58,000 residents access to capabilities it could never economically build or maintain for itself.</p>



<p class="wp-block-paragraph">The key question doesn’t boil down to cloud versus on premises, or proprietary versus open source. It becomes a decision about whether you accept dependency you’ve consciously chosen or dependency you’ve acquired by default. An organization may rationally decide to rent a critical service indefinitely. But it should know where the data lives, how it comes back, what replacing the service would require, and which internal skills have atrophied because the vendor now supplies them. Revisiting those answers periodically, rather than treating last year’s renewal as the rationale for next year’s, is a step that’s easy to skip when nobody’s asking the questions in the first place.</p>



<h2 class="wp-block-heading"><strong>From hobbyists to city hall</strong></h2>



<p class="wp-block-paragraph">Dan suspects the search for alternatives will happen outside big organizations first. He compared it to the early personal computer movement. Hobbyists and enthusiasts experiment first, long before organizations decide the ideas are practical. His hunch is executives will eventually look at how much of their budgets go to recurring subscriptions and ask a simpler question: How much are programmers? Again, I don’t think the answer will be “hire programmers and cancel SaaS.” But more organizations will ask the question behind the question. What are they paying for convenience? What are they paying for capability? And what are they paying because leaving has become too difficult?</p>



<p class="wp-block-paragraph">Software can grow to define an organization rather than serve it, especially when the cost of paying for or maintaining a tool outgrows the tool’s value. That shift becomes a problem when systems are so deeply embedded that replacing them feels impossible or when years of subscriptions leave an organization unable to perform a basic function on its own.</p>



<p class="wp-block-paragraph">Dan’s new job has given that shift a different scale. He told me the biggest adjustment from council member to mayor was realizing that the job is administration and management. He’s less interested in ceremonial appearances than in answering email, returning calls, setting meetings, and, in his words, getting stuff done. Software is part of getting stuff done. So is knowing when to buy it and when to build it. The latest addition to that list may be knowing that the tool with the most impressive new capability is not as valuable as the one you can still leave.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-software-subscriptions-become-public-policy/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Will TypeSafe’s Jev Change How We Build AI Applications?</title>
		<link>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/</link>
				<comments>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 16:28:07 +0000</pubDate>
					<dc:creator><![CDATA[Laurie Voss]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19782</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Will-TypeSafes-Jev-change-how-we-build-AI-applications-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Arize’s blog and is being reposted here with the author’s permission. This week the AI community was in uproar about Jev from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on <a href="https://arize.com/blog/typesafe-jev-llm-judge/" target="_blank" rel="noopener">Arize’s blog</a> and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">This week the AI community was in uproar about <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener">Jev</a> from TypeSafe, not just a new model but a new kind of model: one that classifies, scores, and routes but can’t write a sentence. The reason for the fuss is simple. It’s radically faster and cheaper than using an LLM to perform the same task (up to 200x faster and 400x cheaper if TypeSafe’s numbers are to be trusted). In one small independent test, a general-purpose model spent about 910 output tokens reasoning its way to each yes-or-no answer while Jev spent 85, and it doesn’t even bill for them.</p>



<p class="wp-block-paragraph">That’s potentially a really big deal. An enormous share of LLM-powered components in AI applications today are being asked to make decisions: pass or fail, route A or route B, which of five labels to pick. In particular, that’s something that LLM-as-a-judge evaluations are doing all the time, so it really made our ears perk up at Arize AI. This post is about how we got here, what this new kind of model buys you, what you lose, and what choices you should be making about your application’s architecture as a result.</p>



<h2 class="wp-block-heading">TypeSafe shipped a model that can’t write, only decide</h2>



<p class="wp-block-paragraph">Here’s how Jev works. You send it some data that represents a state (a support ticket or an agent trace or a JSON blob) plus a list of typed questions (Choose one of these options; Score this on a scale; Is this statement true?). It doesn’t generate a token stream. It returns <a href="https://docs.typesafe.ai/concepts/system-one.md" target="_blank" rel="noopener">typed answers with probability distributions</a> in a single parallel pass, in 70 ms to 500 ms, at $0.042 per million input tokens. No free-form text comes back.</p>



<p class="wp-block-paragraph">The training method used to create Jev is what TypeSafe calls Reinforcement Learning for Calibrated Decisions or RLCD, described in their <a href="https://docs.typesafe.ai/introduction/machine-learning-primer.md" target="_blank" rel="noopener">primer</a> as training the model so that a higher stated probability means a higher chance the answer is right. (You’d think that’s always what a higher stated probability should mean, but read on for surprising facts about how LLMs work.)</p>



<p class="wp-block-paragraph">TypeSafe also claims Jev “can’t hallucinate,” but that really feels like an overreach. Jev can’t return an answer outside the schema you gave it. Within that schema, it could still be giving the wrong answer, although its probability score should give you a clue if it’s not confident.</p>



<p class="wp-block-paragraph">And the whole thing is incredibly fast and incredibly cheap: 40x to 200x faster and 40x to 400x cheaper, depending on the task, are TypeSafe’s numbers from TypeSafe’s evals. Of course, we know better than to take a vendor’s word for these things, so Arize will be running our own benchmarks just as soon as we can. But other people have already started doing that.</p>



<h2 class="wp-block-heading">On the early data, Jev is mid-tier intelligence at a two-orders-of-magnitude discount</h2>



<p class="wp-block-paragraph">TypeSafe’s <a href="https://evals.typesafe.ai/" target="_blank" rel="noopener">published evals</a> run four decision workflows, one of which is reviewing a finished agent trace to decide whether a human needs to look at it. Averaged across the four workflows, Jev lands at 68% accuracy at $0.0004 and 0.4 seconds per case. GPT-5.6 Terra is at 68% for $0.03 and 10 seconds. Opus 5 is at 73% for $0.18 and 38 seconds. That’s five points behind Opus 5, but on the other hand it’s 440x cheaper. That’s a very interesting cost-benefit trade-off, and such a radical one that it may change how we architect our applications.</p>



<p class="wp-block-paragraph">The independent data so far is small but it points the same way. <a href="https://every.to/" target="_blank" rel="noopener">Every</a>’s head of evals ran <a href="https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds" target="_blank" rel="noopener">777 judgments in under 0.7 seconds</a> for about a quarter of a cent. A UK events site, NearHere, tested listing moderation and got <a href="https://nearhere.events/blog/typesafe-jev-mistral-gemini-event-validation" target="_blank" rel="noopener">96% from Jev against 86% from Gemini Flash-Lite</a>, 58x cheaper per decision. That’s where the 910-versus-85 token count I mentioned earlier came from. And a developer ran Jev zero-shot over <a href="https://github.com/bitnovus/jev-spam-eval" target="_blank" rel="noopener">18,514 spam emails</a>, getting a result that was a statistical tie versus a classifier trained on the labels.</p>



<p class="wp-block-paragraph">These are small samples and early data but hey, the thing was released a few days ago.</p>



<h2 class="wp-block-heading">We used LLM judges because nothing else worked without labels</h2>



<p class="wp-block-paragraph">Why are we using LLMs to make decisions in the first place? The reason is simple: They are able to do it without huge, expensive training sets, which is what most ML solutions prior to LLMs required. Here’s the options on the table now:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059040199&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059040199" class="aligncenter size-full wp-lightbox-container"><img fetchpriority="high" decoding="async" width="2034" height="740" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png" alt="ML solutions table" class="wp-image-19783" style="aspect-ratio:2.748898678414097" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18.png 2034w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-300x109.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-767x279.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1600x582.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-18-1536x559.png 1536w" sizes="(max-width: 2034px) 100vw, 2034px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The first two rows are the old-school options, which need a training dataset: hundreds to thousands of labeled examples before you get a single prediction, and then a training run, and then someone to maintain it. Nobody building a first version of an AI product has that data or that kind of time. The LLM as a judge, on the other hand, just asks for a paragraph-long prompt. The decision was easy.</p>



<p class="wp-block-paragraph">But the results weren’t without trade-offs. On a <a href="https://www.latent.space/p/benchmarks-201" target="_blank" rel="noopener"><em>Latent Space</em> episode in July 2024</a>, Clémentine Fourrier of Hugging Face laid out what LLM judges are bad at: They prefer their own model family, and they can’t score on a continuous scale. Asked what benchmark she wished existed, Fourrier said, “Nobody’s evaluating model calibration at the moment.” With the release of Jev, the need for that benchmark is even greater, because real progress seems to have been made.</p>



<h2 class="wp-block-heading">Jev gives you zero-shot probabilities without training and without a generator</h2>



<p class="wp-block-paragraph">Jev takes the same plain-English criteria you’d put in a judge prompt, needs no labels, and returns a probability. Zero-shot and autoregressive text generation are no longer tied together. We were paying for the second to get the first, and it turns out you don’t have to.</p>



<p class="wp-block-paragraph">The spam evaluation I mentioned earlier is an impressive demonstration of how attractive this new offering is. With zero labeled examples and a simply well-written definition of spam, Jev hit 98.3% accuracy. A TF-IDF logistic regression trained on about 14,800 labeled emails hit 98.4%. The two disagreed on 466 emails and split them almost evenly, with no statistically meaningful difference between the two. So a classifier from 2003, trained on a dataset, only ties a decision model trained on nothing. It’s early data that’s yet to be reproduced, but if it holds up, that’s an amazing new capability unlocked.</p>



<p class="wp-block-paragraph">But there are still some trade-offs you’re making.</p>



<h2 class="wp-block-heading">A radically cheaper decision loses you some things</h2>



<p class="wp-block-paragraph">The biggest loss is the explanation. TypeSafe’s docs say plainly that System One models don’t generate explanations of their reasoning, and NearHere’s test noted the same thing: a category and probabilities came back, nothing else.</p>



<p class="wp-block-paragraph">Depending on your use case, that could matter a lot. LLM judge explanations are an incredibly valuable tool that tells you not just what was wrong, but why. That provides real signal that can be fed en masse back to a coding agent and used to automatically improve your software. Jev on the other hand just gives you a probability, which leaves you with much less directional signal of how to improve.</p>



<p class="wp-block-paragraph">Of course, at these prices, you can do both: run Jev on every single trace for broad, comparably accurate measurement and monitoring, and then take samples of failures and rerun them through an LLM judge to get your directional signal. That involves changing how you work, which is why I say that this may require rearchitecting your systems.</p>



<h2 class="wp-block-heading">To automate a decision, you need to know which 5% to hand to a human</h2>



<p class="wp-block-paragraph">TypeSafe’s launch post makes the point that a model that’s right 95% of the time but can’t tell you when it’s in the other 5% can’t automate anything, because a person still has to review all of it. That’s an important point because it highlights a problem with LLM judges.</p>



<p class="wp-block-paragraph">We evaluate LLM judges by accuracy against a gold set. Accuracy tells you how many errors to expect, but not where they will be. If your LLM application is making decisions for you, it feeds three things: a threshold that decides when to act, an escalation path that decides when to ask a human, and a drift monitor that decides when the world has changed under it. All three need a probability score, but LLM judges don’t provide reliable probabilities. A <a href="https://arxiv.org/abs/2508.06225" target="_blank" rel="noopener">2025 study of 14 models on JudgeBench</a> found judges clustering their predictions at 90% to 100% confidence while landing well below that in accuracy, and argued for exactly this shift from accuracy-centric to confidence-driven evaluation.</p>



<p class="wp-block-paragraph">The same small spam evaluation test shows what a usable probability looks like. Of the emails Jev scored under 0.1, 0.1% were spam. Of those scored 0.9 or above, 99.9% were. In the 0.5 to 0.6 band, only 38% were. That curve tells you where to set your threshold and how much human review you’re buying: Sending the 4.6% of emails scored between 0.3 and 0.7 to a person left the rest at 99.5% accuracy. Your overconfident LLM judge can’t get you there.</p>



<p class="wp-block-paragraph">Another metric to consider is tokens per decision. A component that spends thousands of output tokens to emit one of five labels is telling you it’s the wrong tool for the job. UkisAI’s <a href="https://huggingface.co/ukisai/Swift-Qwen3.8-27b" target="_blank" rel="noopener">Swift-Qwen3.8-27B</a> cut 58% of its reasoning tokens on GPQA-Diamond and lost 0.1 points of accuracy. A lot of these tokens aren’t making a critical difference to accuracy.</p>



<p class="wp-block-paragraph">In <a href="https://arize.com/products/ax/?utm_source=lvoss&amp;utm_medium=linkedin&amp;utm_campaign=devrel&amp;utm_content=Have%20we%20been%20using%20the%20wrong%20kind%20of%20model%20to%20make%20decisions%3F" target="_blank" rel="noopener">Arize AX</a>, eval labels, the judge’s explanation, and the token count and cost of the judge call sit on the same trace, so tokens per decision is a column you can sort by rather than a number you have to figure out.</p>



<h2 class="wp-block-heading">Cheap decisions change the math of how you build and measure AI applications</h2>



<p class="wp-block-paragraph">As I mentioned earlier, at $0.0004 and 0.4 seconds a decision, you can stop sampling. You can check every output, every tool call, and every agent step as it happens. For some use cases that’s a total game changer.</p>



<p class="wp-block-paragraph">But it might require that you rearchitect how your application works to make the most of it. Take the work and decompose into many small typed questions; only call the expensive LLM generator when text actually needs to be written. That’s a stack where the decision layer is something you can version, measure, and swap independently of the model that writes the words, and it’s the first time decision-making has been cheap and fast enough to make that practical without requiring training data.</p>



<p class="wp-block-paragraph">So go count how many of your LLM calls end in one of five labels. Then work out what you’d check, and how often, if each of those calls cost a fraction of a cent and came back with a probability you could trust.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/will-typesafes-jev-change-how-we-build-ai-applications/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>MCP Is Not Just Another API Standard</title>
		<link>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/</link>
				<comments>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/#respond</comments>
				<pubDate>Wed, 23 Sep 2026 10:17:25 +0000</pubDate>
					<dc:creator><![CDATA[Balaji Venkatasubramaniyar]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Operations]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19772</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/MCP-is-not-just-another-API-standard-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What Model Context Protocol actually changes for agentic systems, and what it doesn’t]]></custom:subtitle>
		
				<description><![CDATA[Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water through pipes you already designed. MCP changes who’s holding the wrench. Once you’ve felt that shift in a real production system, the “just another API standard” framing stops making sense.</p>



<p class="wp-block-paragraph">This piece is the long version of that argument. It walks through what MCP’s primitives actually are and why they’re the right primitives, where the standard genuinely collapses integration work that used to be duplicated per framework, where the abstraction leaks in ways that only show up once you’re past the demo, and what the protocol’s own 2026 evolution tells you about where the real pain has been. Nearly everything worth knowing about building on MCP falls out of understanding these pieces and how they interact.</p>



<h2 class="wp-block-heading"><strong>What MCP actually standardizes</strong></h2>



<p class="wp-block-paragraph">Strip away the framing and MCP is a JSON-RPC-based protocol that lets a client (the thing driving an LLM) talk to a server that exposes capabilities, over a small, fixed set of primitives:</p>



<p class="wp-block-paragraph"><strong>Tools</strong> are callable functions. Each one has a name, a description, and a JSON Schema describing its inputs. This is the primitive most people mean when they say “MCP,” and it’s the one doing the heavy lifting in most production deployments: “look up an order,” “run a query,” “create a ticket,” etc.</p>



<p class="wp-block-paragraph"><strong>Resources</strong> are readable context, addressed by URI, that a client can pull in without the model having to call a function to get it: a file, a record, or a document, for instance. Think of this as the read side of the interface, separate from the “do something” side that tools represent.</p>



<p class="wp-block-paragraph"><strong>Prompts</strong> are reusable templates a server offers to the client, so common workflows don’t have to be respecified from scratch every time.</p>



<p class="wp-block-paragraph">On top of those three, the spec defines capabilities that flow the other direction, from server back to client: Sampling lets a server ask the client’s model to generate text on its behalf; elicitation, added in the 2025-06-18 revision, lets a server pause and ask the human for more input mid-task; and roots let a server learn which directories or URIs it’s actually allowed to touch.</p>



<p class="wp-block-paragraph">None of these primitives are individually novel. What’s novel is that they’re the same five primitives regardless of which model, which framework, or which vendor is on the client side. That’s the entire value proposition in one sentence, and it’s also the source of everything that goes right and everything that goes wrong when you build on top of it.</p>



<h2 class="wp-block-heading"><strong>Where the old model breaks down</strong></h2>



<p class="wp-block-paragraph">Before MCP, wiring an LLM into an enterprise system meant writing tool-calling code for that specific model, that specific framework, that specific integration. Every agent framework had its own function-calling convention: its own way of describing a schema, its own way of parsing a model’s intent to call something, and its own error-handling contract. Every system you wanted to expose needed its own adapter written to whichever dialect that framework spoke. Add a second framework to your stack and you don’t get twice the work. You get a second, incompatible copy of the same logic, maintained by whoever drew the short straw.</p>



<p class="wp-block-paragraph">MCP replaces that with one contract, written once, usable by any compliant client regardless of which model sits behind it. That’s the part every MCP explainer gets right, and it’s a real, measurable win. I’ve watched it collapse from a maintenance burden that used to scale with the number of frameworks a team happened to be supporting that quarter down to something that scales with the number of systems, full stop.</p>



<p class="wp-block-paragraph">But the more consequential change is where the integration decision gets made. A traditional API integration is an agreement two systems make in advance. You negotiate a contract: endpoints, payloads, auth, versioning, and both sides build to it, because a project plan said this integration should exist. The plan predates the code.</p>



<p class="wp-block-paragraph">An MCP server doesn’t get that luxury. It has no idea which agent will call it, in what sequence, alongside which other servers, in service of what goal a human typed into a chat box 30 seconds ago. The plan doesn’t exist as a concrete thing until the agent composes one, at runtime, out of whatever tools happen to be available to it. That’s not a stylistic difference from the old model. It’s a different category of integration problem, because the party doing the composing isn’t your code anymore. It’s a model, reasoning over natural-language descriptions you wrote weeks or months earlier, with no idea what context it would eventually be reasoning inside of.</p>



<h2 class="wp-block-heading"><strong>The description is the interface now</strong></h2>



<p class="wp-block-paragraph">A tool’s JSON Schema tells the agent what parameters it takes and what shape they need to be. That part is mechanical, and MCP handles it well. The tool’s name and description tell the agent when to use it at all, and whether to prefer it over some other tool that does something adjacent. Those are two different jobs, and only one of them is solved by a well-formed schema.</p>



<p class="wp-block-paragraph">Picture two versions of the same tool description. The first is technically correct and nothing more:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current status of a record given its ID."
}</code></pre>



<p class="wp-block-paragraph">An agent reading that has no idea when this is the right tool versus three other tools that also return some kind of status, no idea what “record” means in this system, and no idea whether IDs are case-sensitive, numeric, or prefixed. The second version spells out the domain the tool operates in, gives the ID format explicitly, states what the returned status values mean, and flags the one adjacent tool this one is commonly confused with and why they’re different:</p>



<pre class="wp-block-code"><code>{
  "name": "get_status",
  "description": "Returns the current fulfillment status for an order record. 
IDs are numeric order numbers (e.g. 48213), not SKUs or customer IDs. Status 
values are one of: pending, processing, shipped, delivered, cancelled. Use this 
instead of get_shipment_status, which returns carrier tracking events rather 
than the order's internal state."
}
</code></pre>



<p class="wp-block-paragraph">That’s a longer description, and it will feel like overexplaining to the engineer writing it, because the engineer already knows all of this. The agent doesn’t. It’s encountering the tool for the first time, with a handful of tokens to decide whether it’s the right call, and no colleague to ask.</p>



<p class="wp-block-paragraph">I’ve watched teams ship a technically correct MCP server that agents used badly, or avoided entirely in favor of a worse but better-described alternative, purely because of this gap. The failure mode isn’t a stack trace. It’s an agent confidently calling the wrong tool, or the right tool with an assumption baked in that happened to be wrong for this case, and nobody notices until the output looks slightly off downstream. Writing tool descriptions well is closer to technical writing and product design than it is to backend engineering, and it’s not a skill most integration teams (mine included, early on) walked in the door with.</p>



<h2 class="wp-block-heading"><strong>Composition is emergent, and that cuts both ways</strong></h2>



<p class="wp-block-paragraph">The entire appeal of MCP is that an agent can combine tools from servers that never agreed to work together, in combinations their respective authors never planned for. That’s also the risk, and it’s structural, not a bug you fix with better testing.</p>



<p class="wp-block-paragraph">In a traditional integration, the sequencing logic (call A, then check its result, then decide whether to call B or C) lives in a script that a human wrote and a reviewer read. You can unit test it. In an MCP-based agent, that same sequencing logic lives in the model’s runtime reasoning, generated fresh for each task based on the goal it was given and whatever tools happen to be available in that session. You can’t unit test a decision that doesn’t exist until the moment it’s made.</p>



<p class="wp-block-paragraph">A tool that behaves correctly in isolation, with the exact inputs its author tested against, can still produce a bad outcome the first time an agent calls it third instead of first in a chain, or passes it a value that came from a different server’s output rather than a human’s direct input. This is qualitatively different from a normal integration bug, because it doesn’t show up in code review, and it won’t show up in testing unless your test suite happens to exercise that specific, unplanned chain of calls. It shows up in production, once, when a particular combination finally occurs. That’s exactly the kind of failure mode that’s cheap to dismiss as an edge case until it happens to the wrong customer.</p>



<h2 class="wp-block-heading"><strong>The protocol is catching up to its own success, in specific and telling ways</strong></h2>



<p class="wp-block-paragraph">To be fair to MCP, it isn’t standing still, and the shape of its evolution tells you a lot about where the real production pain has been. The July 28, 2026 specification is the largest revision since the protocol’s November 2024 launch, and every major change in it traces back to something that broke, or nearly broke, at scale.</p>



<p class="wp-block-paragraph">The protocol core is now stateless. The original design tracked sessions with an MCP-Session-Id header, workable for a single server instance but painful the moment you’re running behind a normal horizontally scaled fleet and discover that “any instance can answer any request” and “sticky session state” don’t coexist. Removing protocol-level sessions means the same request can be served by any instance behind ordinary load-balancing infrastructure, which sounds unglamorous right up until you’re the one who has to explain in an incident review why a routine deploy dropped a chunk of in-flight sessions.</p>



<p class="wp-block-paragraph">Tasks formalize long-running work. A lot of real enterprise work (document processing, multistep approvals, anything involving a human in the loop) doesn’t complete inside a single request/response cycle. Before this extension existed, teams hand-rolled this with polling loops and webhook callbacks, each implementation slightly different, each one a source of its own edge cases. Tasks turn that into a first-class protocol concept.</p>



<p class="wp-block-paragraph">MCP Apps let a server return interactive UI, not just structured data. That matters the moment a “tool” is something a human needs to actually look at and approve before it fires, which in any environment with real consequences attached is often.</p>



<p class="wp-block-paragraph">Authorization was hardened to align with OAuth 2.1 and OpenID Connect. This one isn’t novel so much as overdue, and the gap it closes was a real one; see the governance section below.</p>



<p class="wp-block-paragraph">A formal deprecation policy now governs the legacy HTTP+SSE transport, with a 12-month offramp. That’s the kind of unglamorous governance maturity a protocol only earns after it’s been run in production long enough for someone to need it.</p>



<p class="wp-block-paragraph">None of this is exciting reading. All of it is the sound of a two-year-old protocol absorbing genuine operational scar tissue, which is a far better signal about its trajectory than raw adoption numbers. And the adoption numbers are themselves striking: The official registry tracks close to 10,000 distinct servers, Tier 1 SDK downloads run into the tens of millions monthly, and both the TypeScript and Python SDKs have individually crossed a billion total downloads. Competitors of the protocol’s original author adopted it within months. That combination, real scale plus a spec that keeps changing in response to real production failure modes, is a much stronger signal of durability than either fact alone.</p>



<h2 class="wp-block-heading"><strong>Governance hasn’t caught up</strong></h2>



<p class="wp-block-paragraph">Here’s what I’d want any team to weigh before connecting MCP to anything that matters. Independent security research through 2026 paints a specific, and specifically uncomfortable, picture of the current ecosystem.</p>



<p class="wp-block-paragraph">Scans across thousands of publicly registered servers have found the large majority carrying file-operation patterns prone to path traversal. A meaningful share of tested servers are vulnerable to command injection or server-side request forgery, and there are documented, disclosed cases of tool description poisoning, where the attack lives in the text a model reads to decide what to do rather than in the code the tool actually executes. A closely related failure mode, configuration poisoning, targets the server’s operational baseline directly: stealthy permission changes or altered defaults that persist across sessions and are hard to catch in a normal code review because the malicious logic lives in configuration state, not application code. Multiple high-severity vulnerabilities, including at least one missing-authentication flaw in a major vendor’s own production package, have already been disclosed and patched.</p>



<p class="wp-block-paragraph">None of that is a reason to avoid MCP. It’s a reason to treat it the way you’d treat any protocol that hands an autonomous caller real privileges inside your systems: skeptically, and with the controls in place before the agent gets access rather than after an incident teaches you why you needed them. In practice that means an explicit, enforced allowlist of vetted tools per agent rather than open discovery of whatever happens to be reachable; authentication on every remote endpoint with no quiet exception carved out for “internal” traffic; centralized, immutable audit logging of every tool call an agent makes; and secrets pulled dynamically from a real secrets manager rather than sitting in a server’s local config where a configuration-poisoning attack can find them. None of this is exotic. It’s the same discipline any experienced integration team already applies to systems with real privileges, applied here to a caller that can now improvise its own sequence of actions.</p>



<p class="wp-block-paragraph">Here’s what I’d tell a team starting today.</p>



<p class="wp-block-paragraph">Treat the tool description as reviewed engineering output, not documentation you write last and skim once. Test it against how an agent actually behaves when given it, not just against whether a human reviewer nods along.</p>



<p class="wp-block-paragraph">Assume composition you didn’t plan for will eventually happen, and design tools to fail safely and legibly when it does, rather than assuming a chain of calls you never tested simply won’t occur.</p>



<p class="wp-block-paragraph">Put governance in front of capability, not after it. The allowlist, the auth, the audit log, and the secrets manager are the entry price given where the current vulnerability data sits, not optional hardening for later.</p>



<p class="wp-block-paragraph">And build against the current specification baseline, not whichever example repository you copied six months ago. The stateless core and the authorization changes in the July 2026 spec aren’t cosmetic; targeting an older baseline today is technical debt you’re taking on knowingly, on day one.</p>



<p class="wp-block-paragraph">MCP earned the “not just another API standard” framing honestly. It didn’t get there by being a cleaner REST, or a nicer SDK, or a better-documented function-calling convention: all real, all incremental. It got there by changing who, or what, is actually doing the integration work at runtime. The parts of that job the protocol doesn’t standardize (how well you describe a capability, how safely your tools behave when composed in ways you never anticipated, and how seriously you take governance before you grant an agent real privileges) are exactly the parts worth taking seriously before you bet production traffic on it.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/mcp-is-not-just-another-api-standard/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Accelerationist Case for Frontier Pacing</title>
		<link>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/</link>
				<comments>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/#respond</comments>
				<pubDate>Tue, 22 Sep 2026 15:57:06 +0000</pubDate>
					<dc:creator><![CDATA[Venkatesh Rao]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19768</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-accelerationist-case-for-frontier-pacing.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-accelerationist-case-for-frontier-pacing-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Slow AI is smooth AI, smooth AI is fast AI]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Venkatesh Rao’s Substack, Contraptions, and is being republished here with the author’s permission. The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://contraptions.venkateshrao.com/p/the-accelerationist-case-for-frontier" target="_blank" rel="noopener">Venkatesh Rao’s Substack, Contraptions</a>, and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment year) and launched my brief and entirely undistinguished athletic career. By my senior year, however, my 50m time had improved to about 38s (not enough to get me off water-boy duty since the team had several exceptional swimmers with much better times). Interestingly though, it was <em>easier</em> for me to swim faster at the end of my career than it was to swim slower in the beginning. The reason was that in the interim, the coach had significantly improved my stroke and breathing technique. It was all about managed pacing, not raw intensity of effort.</p>



<p class="wp-block-paragraph">There is a fairly deep literature behind this apparently mundane lesson. Daniel Chambliss’s classic 1989 paper “<a href="https://www.jstor.org/stable/202063?seq=1#page_scan_tab_contents" target="_blank" rel="noopener">The Mundanity of Excellence</a>,” based on years of fieldwork studying competitive swimmers all the way from local clubs to the Olympic level, argued that excellence is primarily qualitative rather than quantitative. Elite swimmers do not simply do more of what mediocre swimmers do, or do it harder. They organize their activity differently: Strokes, turns, training habits, attention, and countless other small practices combine into a qualitatively different way of swimming. The route to excellence is not therefore reducible to maximizing effort along some obvious scalar dimension.</p>



<p class="wp-block-paragraph">The same insight is condensed in a maxim common in military and special operations circles: “Slow is smooth, smooth is fast.” In activities where speed really matters, trying to go fast naively is often an excellent way to go slowly.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img decoding="async" width="1456" height="971" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.jpeg" alt="" class="wp-image-19769" style="aspect-ratio:1.5009380863039399;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29.jpeg 1456w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-300x200.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-29-768x512.jpeg 768w" sizes="(max-width: 1456px) 100vw, 1456px" /></figure>
</div>


<p class="wp-block-paragraph">I never really stopped thinking about this problem. My 2011 book <em><a href="https://books.venkateshrao.com/configurancy-tempo/" target="_blank" rel="noopener">Tempo</a></em> grew partly out of a long-standing interest in pacing across performance domains: how people experience time while making decisions, how rhythms of action emerge, and how timing relates to effectiveness. One of the ideas that has stuck with me since then is that tempo is something to be managed rather than maximized. There is no universally correct speed. There are only tempos appropriate or inappropriate to the dynamics of the situation.</p>



<p class="wp-block-paragraph">Which brings me, somewhat unexpectedly, to Dario Amodei.</p>



<p class="wp-block-paragraph">Amodei <a href="https://darioamodei.com/post/we-must-pace-the-frontier" target="_blank" rel="noopener">recently made the case</a> that frontier AI development should be deliberately paced. His argument is primarily a safety argument. AI capabilities, he believes, are advancing quickly enough that the processes required to understand, evaluate, align, secure, and safely operate them are having trouble keeping up. This is not quite the old proposal for an AI “pause.” Pacing means continuing to advance the frontier while deliberately managing its rate, allowing safety work and institutional capacity to remain within striking distance of capability. Sam Altman has now endorsed the basic proposition, and Demis Hassabis has made closely related arguments about frontier capabilities outrunning scientific understanding and governance capacity. Elon Musk, more tersely, has said that Amodei is right.</p>



<p class="wp-block-paragraph">There is an obvious cynical reading of this emerging consensus. The leading frontier labs have powerful economic reasons to want a regulated frontier. A regime that requires enormous compliance budgets, restricts open-weight releases, discourages foreign models, imposes burdens that startups cannot afford, or legitimizes coordination among a small number of incumbents could turn “safety” into a remarkably effective mechanism for protectionism and regulatory capture. That suspicion is not paranoid. Open models increasingly constitute a competitive threat to proprietary frontier providers, and the politics around regulating them <a href="https://www.washingtonpost.com/wp-intelligence/ai-tech-brief/2026/07/20/ai-tech-brief-open-source-debate/" target="_blank" rel="noopener">already feature</a> explicit accusations of regulatory capture. There is an additional awkwardness: Coordinated pacing among nominal competitors looks uncomfortably like coordinated restriction of output, enough so that the legality of such arrangements under antitrust law <a href="https://www.wired.com/story/openai-wants-to-know-if-an-ai-industry-slowdown-would-even-be-legal/" target="_blank" rel="noopener">is already being debated</a>.</p>



<p class="wp-block-paragraph">I don’t think we need to resolve the question of motives. Perhaps these CEOs are sincerely terrified. Perhaps they are sincerely terrified <em>and</em> understand perfectly well that the regulations they favor would strengthen their competitive positions. Perhaps the mixture varies by person, company, and day of the week. It doesn’t matter much for my argument. The proposition that the frontier should be paced is worth considering independently of the political economy of the people proposing it.</p>



<p class="wp-block-paragraph">I also don’t share enough of Amodei’s safety premises to make his argument my own. In particular, I think a great deal of contemporary concern about runaway AGI, superintelligence, and “alignment” is badly framed, and often borders on the theological. But I increasingly agree with his conclusion.</p>



<p class="wp-block-paragraph">In fact, I think there is a strong case for frontier pacing even if you are an accelerationist and your objective is simply to make technological progress happen as fast as possible. I am not myself an accelerationist. My preferred framing is closer to managed tempo. But if I were one, I would still favor pacing the frontier right now, for a simple reason: <em>Maximizing the instantaneous velocity of the AI capability frontier is no longer obviously maximizing the rate of technological progress.</em></p>



<p class="wp-block-paragraph">There is, however, an important difference between my conclusion and the emerging frontier consensus. Their natural solution is coordination at the top: labs agreeing upon thresholds, governments blessing the coordination, evaluators policing it, and eventually perhaps international agreements extending it.</p>



<p class="wp-block-paragraph">In other words, cartelization, hopefully of a benign sort.</p>



<p class="wp-block-paragraph">I would prefer to see how much frontier pacing can be produced from the bottom up through ordinary market mechanisms. The distinction matters. The objective should not be to decide administratively how fast AI is allowed to improve. It should be to stop artificially rewarding frontier velocity after frontier velocity has ceased to be the most important form of progress.</p>



<h2 class="wp-block-heading"><strong>Getting inside the loop</strong></h2>



<p class="wp-block-paragraph">A useful way to understand the distinction comes from another idea that startup culture has borrowed, and mostly misunderstood, from the military: John Boyd’s OODA loop. OODA theory says that you win by “getting inside the adversary’s decision cycle,” which is usually glossed as making decisions <em>faster</em> than the other guy. If you observe, orient, decide, and act faster than he can, the story goes, you eventually overwhelm him.</p>



<p class="wp-block-paragraph">But <em>inside</em> does not mean <em>faster</em>. The objective is to operate within the decision dynamics of the system you are engaging in a way that lets you shape them. Against a human adversary, that may indeed sometimes involve accelerating until his ability to orient collapses psychologically. But it may also require waiting, withholding action, changing rhythm, or deliberately slowing down. In nonadversarial situations, the goal may not be collapse at all but harmonization for resonant support.</p>



<p class="wp-block-paragraph">What matters is the right tempo at the right phase, not speed for the sake of speed.</p>



<p class="wp-block-paragraph">Something analogous applies to scientific and technological progress. There is no enemy psychology to collapse, but there are still loops to get inside: observation, experimentation, interpretation, investment, construction, deployment, feedback, learning, and recombination. The useful question is not how rapidly one component of that system can be made to move. It is whether the tempo of development allows those loops to close. If one subsystem changes faster than the surrounding system can observe, understand, absorb, and respond to it, pushing that subsystem still faster can reduce rather than increase effective progress. It can induce fragility and collapse.</p>



<p class="wp-block-paragraph">This, I think, is approximately where AI is now.</p>



<p class="wp-block-paragraph">The simplest evidence is personal and almost embarrassingly mundane. Frontier AI is already overpowered for nearly everything I use it for. In my most advanced projects I may use the strongest model available (Fable for my coding projects) to plan an approach or make critical strategic decisions, but I can generally hand the resulting specification to a cheaper model (such as Opus or Sonnet) to do the routine work. For ordinary uses I don’t need anything close to the frontier. In ChatGPT, I no longer even know exactly which model I am talking to much of the time. Whatever the “think harder” control does is sufficient model selection for my purposes.</p>



<p class="wp-block-paragraph">The situation increasingly reminds me of smartphones. There was a period when getting the newest iPhone produced a noticeable improvement in everyday life. Eventually the hardware got good enough that the upgrade cycle ceased to matter much. I kept an iPhone XS for almost a decade before replacing it with a 16. The frontier continued advancing; I simply fell off the frontier because my demand curve had stopped following it.</p>



<p class="wp-block-paragraph">Something similar is beginning to happen with AI, except that the supply curve is moving incomparably faster. Six months ago I routinely maxed out token allotments. Now I don’t. Some weeks I barely use coding agents. This isn’t because I’ve become less interested in AI. It is because my own capacity to productively absorb AI output has become the constraint. I have projects to think about, things to read, people to talk to, and work to do in domains where AI cannot help me yet, or perhaps ever. I am already pacing myself at my own tiny personal frontier.</p>



<p class="wp-block-paragraph">That is a significant change in the technological situation. The binding constraint is migrating.</p>



<h2 class="wp-block-heading"><strong>When the bottleneck moves</strong></h2>



<p class="wp-block-paragraph">Broader AI deployment is increasingly blocked by things other than model intelligence. Robotics has long been constrained by actuators, power, reliability, dexterity, manufacturing, and the sheer recalcitrance of the physical world. Those constraints are beginning to move, but making the model smarter does not make them disappear. AI in education is constrained less by whether a model can explain calculus than by our lack of sufficiently rich classroom experimentation about what happens when students and teachers actually use these systems. Current mid-tier models are probably capable enough to power almost any educational experiment worth trying in a high-school or undergraduate classroom. We do not need another order of magnitude of intelligence before conducting them.</p>



<p class="wp-block-paragraph">This pattern should become more common as AI improves. Once intelligence ceases to be scarce, its complements become more important. Model capability can be abundant while classroom knowledge is scarce. Model capability can be abundant while actuators are scarce. Model capability can be abundant while electrical infrastructure is scarce. It can be abundant while organizational competence, human attention, scientific understanding, military doctrine, security practices, and good judgment are scarce.</p>



<p class="wp-block-paragraph">This is not peculiar to AI. Capability-maxxing the coolest new weapon is bad military doctrine. The United States has enjoyed extraordinary technological superiority over its adversaries for decades and has nevertheless repeatedly discovered that superior equipment does not automatically produce strategic success. Logistics, doctrine, morale, training, political understanding, industrial capacity, and orientation matter. A force that neglects those complements because it possesses the best weapons can become remarkably fragile. From Vietnam to Iran, the US military has been repeatedly forced to relearn the lesson.</p>



<p class="wp-block-paragraph">AI may now be entering the same regime. Fragility from neglect of everything non-AI is becoming a bigger risk than failure to token-max.</p>



<p class="wp-block-paragraph">There is a further reason to suspect that continuing to redline the existing frontier may yield diminishing returns. The major labs increasingly appear to be competing along broadly the same technological S-curve. One suggestive sign is that they run into the same supply constraints: HBM, electrical power, data-center capacity, capital, and access to sufficiently large clusters. When a technological system moves onto a genuinely different S-curve, its important bottlenecks often change as well. If everybody’s problem is how to secure more of the same scarce inputs to do more of the same basic thing, that is at least suggestive that everybody is climbing the same sigmoid.</p>



<p class="wp-block-paragraph">As I learned as a freshman swimmer, near the upper portion of an S-curve, pushing harder can become exactly the wrong acceleration strategy. You expend increasing resources for decreasing gains while starving exploration of the attention required to discover the next curve. Moving faster <em>along</em> an S-curve is not the same thing as accelerating technological evolution. Sometimes you have to back off the incumbent trajectory long enough to notice what the next trajectory is.</p>



<h2 class="wp-block-heading"><strong>A cognitive ergonomics crisis</strong></h2>



<p class="wp-block-paragraph">There is also a more immediate bottleneck that the AI industry seems reluctant to acknowledge: the humans at the frontier.</p>



<p class="wp-block-paragraph">Startup people have been LARPing war for decades. This is one reason concepts like OODA became popular in startup culture in the first place. “War mode” usually means working extremely hard under conditions of strong personal financial incentives: long hours, high urgency, extreme focus, centralized authority, and a willingness to sacrifice ordinary organizational niceties. I’ve been around startup culture for decades, and I suspect the AI boom may be the first time the conditions have actually become meaningfully war-like.</p>



<p class="wp-block-paragraph">AI frontier people are visibly unprepared for it.</p>



<p class="wp-block-paragraph">Actual militaries and other frontline risk professions take the human consequences of sustained high-stress operations seriously. Soldiers, firefighters, emergency medical personnel, disaster responders, surgeons, pilots, and others operating in consequential environments develop elaborate practices around training, emotional regulation, redundancy, rotations, decompression, mandatory rest, checklists, after-action review, and recovery. These practices exist because motivation does not repeal physiology. Judgment deteriorates. Attention narrows. People make stupid mistakes. Emotional reactions become harder to regulate. Creativity disappears. Eventually people break.</p>



<p class="wp-block-paragraph">Does frontier AI look like an industry managing itself accordingly?</p>



<p class="wp-block-paragraph">From the outside, it looks closer to the opposite. People at the frontier have been operating under extraordinary pressure for several years with little respite. The cognitive ergonomics of their working conditions are a disaster. (We have a <a href="https://protocol-institute.org/projects/project/?slug=cognitive-ergonomics" target="_blank" rel="noopener">project going</a> at the Protocol Institute led by <a href="https://open.substack.com/users/17195021-timber-stinson-schroff?utm_source=mentions" target="_blank" rel="noopener">Timber Stinson-Schroff</a> to study this—contact him if you’re interested in participating in or supporting it.)</p>



<p class="wp-block-paragraph">Competitive pressure, enormous amounts of capital, geopolitical attention, hostile public scrutiny, internal ideological battles, rapidly changing technology, and the conviction among some participants that their daily work may determine the fate of humanity are not normal occupational stressors. The rate of dumb, unforced errors appears to be rising. The quality of frontier discourse has, in my view, visibly deteriorated. People I once assumed were much smarter than me increasingly seem to be missing obvious things while becoming susceptible again to bad ideas I thought they had outgrown.</p>



<p class="wp-block-paragraph">Tired people catch colds more easily; they catch bad ideas more easily too.</p>



<p class="wp-block-paragraph">This produces a peculiar inversion of the conventional AI safety model. We normally imagine increasingly unreliable or dangerous AIs surrounded by reliable human supervisors. But what if we are increasingly producing extremely capable AIs surrounded by progressively less reliable humans?</p>



<p class="wp-block-paragraph">“Human in the loop” is not much of a safety guarantee if the human has been metaphorically deployed aboard an aircraft carrier in a war zone for eight months without relief.</p>



<p class="wp-block-paragraph">Some of the public testimony emerging from frontier organizations should perhaps be interpreted through this lens. I do not want to diagnose particular people from afar, and testimony from people who have worked closely with frontier systems should obviously be taken seriously. But when someone emerges from prolonged immersion at the frontier sounding psychologically shattered, there are at least two possible kinds of information in the signal. One concerns the technology. The other concerns what prolonged immersion at the frontier does to the observer. Frontier workers are sensors, but the sensors themselves are being perturbed by the phenomenon they are measuring.</p>



<p class="wp-block-paragraph">We have seen versions of this going back at least to the Blake Lemoine episode at Google, when sustained interaction with LaMDA led him to conclude that the system was sentient. More recently, former frontier employees have emerged making extraordinarily grave predictions about where AI is heading, that ill-prepared, tech-hostile journalists are eagerly amplifying with lurid headlines.</p>



<p class="wp-block-paragraph">The correct response need not be either “believe them and stop AI” or “they’re crazy and should be ignored.” Sometimes a sensible response to someone coming back from the front sounding shell-shocked and exhibiting symptoms of PTSD is: This person needs a vacation. We rotate soldiers partly because the testimony of exhausted soldiers matters.</p>



<p class="wp-block-paragraph">The largest near-term AI safety concern may therefore be exhausted frontline humans supervising overpowered AIs.</p>



<h2 class="wp-block-heading"><strong>Fighting the wrong enemy</strong></h2>



<p class="wp-block-paragraph">Exhaustion is particularly dangerous when nobody can agree about what the enemy is. Much of the actual stress experienced by frontier organizations comes from a fairly comprehensible mixture of competitive pressure and techlash hostility. Those forces are intense, but neither is an existential adversary.</p>



<p class="wp-block-paragraph">The clearest live adversarial problem involving AI is much more ordinary: humans using AI against other humans. Criminal applications are already real and deserve serious attention. Military applications are rapidly becoming real as well, and the relevant strategic picture is much broader than a stylized US-versus-China AI race. Smaller powers and nonstate actors can use cheap cognitive capability to lower engineering barriers that previously required deeper technical institutions. Recent reporting, for example, describes AI assistance being used in weapons-engineering work by actors in Houthi-controlled Yemen. That strikes me as the kind of development around which one can build a concrete threat model.</p>



<p class="wp-block-paragraph">Longer-term military diffusion is clearly a serious concern. But much of the fear actually shaping frontier behavior seems aimed somewhere else entirely: toward vague runaway “AGIs,” “superintelligences,” and a metaphysically capacious notion of “alignment” inherited from philosophical traditions I find largely unpersuasive.</p>



<p class="wp-block-paragraph">There is a useful analogy with climate change. Climate change produces actual physical stressors: more extreme weather, unstable agricultural conditions, infrastructure damage, wildfire risk, and so on. Those generate concrete political and humanitarian problems, including displacement and unmanaged refugee flows, while longer-term adaptation requires things like shoreline management, wildfire regimes, agricultural relocation, and preparedness for changing disease ecologies. Yet parts of climate politics have preferred to identify an ultimate metaphysical adversary called Capitalism, Markets, or Growth and an equally totalizing remedy called degrowth. Heterogeneous problems with different timescales and mechanisms get collapsed into one grand theory.</p>



<p class="wp-block-paragraph">Parts of AI safety discourse increasingly strike me the same way. Competitive instability, cybercrime, weapons proliferation, institutional disruption, labor-market effects, and exhausted frontier personnel are all real and different problems. “Unaligned superintelligence” turns them into a single theological object. Once that happens, every stressor becomes evidence for the same threat model.</p>



<p class="wp-block-paragraph">This is a kind of threat-model collapse. Adaptation is usually plural; apocalypse is singular. Real technological transitions produce dozens of mismatched rates and local failure modes, requiring different responses at different tempos. If criminals are the problem, work on security and law enforcement. If weapons diffusion is the problem, work on doctrine and proliferation. If operators are exhausted, rotate them. If schools lack experimental knowledge, run experiments. If power is scarce, build infrastructure. “Align superintelligence” is not a substitute for any of those things.</p>



<p class="wp-block-paragraph">Pacing would give us something valuable here beyond safety: enough time to discriminate among threats.</p>



<h2 class="wp-block-heading"><strong>Proof abundance, understanding scarcity</strong></h2>



<p class="wp-block-paragraph">Mathematics may already offer a miniature preview of what happens when one part of a knowledge-production system accelerates far beyond the others.</p>



<p class="wp-block-paragraph">Terence Tao has recently <a href="https://teorth.github.io/tao-web/ai-views.html?utm_source=chatgpt.com" target="_blank" rel="noopener">distinguished</a> three stages of mathematical work: generation, verification, and digestion. AI is rapidly making the first two cheaper. Models can generate candidate proofs, while formal systems such as Lean can increasingly verify them. But digestion remains stubbornly slow. Somebody still has to understand what the proof is doing, relate it to existing mathematics, extract reusable techniques, explain it, teach it, and use the resulting understanding to generate better questions. Tao describes the resulting condition as an “impedance mismatch.”</p>



<p class="wp-block-paragraph">This is particularly interesting in light of what I have elsewhere called the <a href="https://contraptions.venkateshrao.com/p/the-curiously-playable-universe" target="_blank" rel="noopener">curiously playable universe</a>: the apparently expanding set of domains that can be transformed into sufficiently explicit games that AI can optimize effectively within them. Anything that begins to resemble a CAD system, a formal proof environment, or an evolutionary optimization problem over a sufficiently well-defined parameter space becomes potentially tractable to extraordinarily capable models. More of the world appears to be playable than we previously thought.</p>



<p class="wp-block-paragraph">But playability has an important pathology. A highly playable domain supplies a scoreboard, and once AI becomes extraordinarily good at optimizing the scoreboard, the relationship between winning the game and advancing the larger domain can weaken. Solving a theorem is valuable partly because, historically, getting to the solution usually required acquiring understanding along the way. If an AI can helicopter directly to the summit, to borrow Tao’s analogy, the summit has still been reached, but nobody necessarily learned the trails, landmarks, terrain, or neighboring geography encountered during the climb.</p>



<p class="wp-block-paragraph">Tao and two dozen other Fields Medalists recently made essentially this point in a declaration strikingly titled “<a href="https://mathandai.org/" target="_blank" rel="noopener">A Severe Misalignment of AI in Mathematics</a>.” Their pointed use of <em>misalignment</em> is almost the reverse of its standard AI-safety meaning. The problem they identify is not that AI has developed alien goals. It is that the incentives of AI companies to demonstrate spectacular problem-solving performance are becoming misaligned with the goals of mathematics itself. Solving difficult problems has historically served as a proxy for mathematical understanding and progress. Once AI can optimize the proxy directly, the correlation can break.</p>



<p class="wp-block-paragraph">The recent Navier–Stokes episode illustrates the issue. Enormous amounts of inference can now be directed at a famous open problem, candidate constructions produced, and formal verification generated at extraordinary speed. Yet that does not automatically produce a corresponding increase in comprehensible, reusable mathematical knowledge. The pipeline is something like problem selection → generation → verification → exposition → digestion → canonicalization → better questions. Increasing the bandwidth of generation and verification by orders of magnitude while leaving the downstream stages roughly unchanged creates a queue.</p>



<p class="wp-block-paragraph">Proof generation becomes abundant. Understanding becomes scarce.</p>



<p class="wp-block-paragraph">That is frontier pacing in miniature. Maximum local throughput does not imply maximum system throughput. Indeed, beyond a certain point it can create congestion.</p>



<h2 class="wp-block-heading"><strong>Orientation beats equipment</strong></h2>



<p class="wp-block-paragraph">There is a strategic implication here for organizations outside the frontier labs. The natural reaction to rapidly advancing models is to assume that whoever possesses the strongest model necessarily possesses an overwhelming advantage. If a frontier model can turn increasingly playable engineering problems into few-shot solutions, and if most of the necessary input information exists somewhere in public literature, then organizations can easily conclude that whatever intellectual lead they possess is temporary. Why bother competing with organizations that possess better models, more compute, more money, and privileged access to the frontier?</p>



<p class="wp-block-paragraph">But this risks confusing equipment superiority with orientation superiority.</p>



<p class="wp-block-paragraph">Boyd repeatedly emphasized that superior orientation could overcome substantial equipment disadvantages. (“We’d still have won if we’d swapped equipment.”) The relevant analogy today is something like centaur chess. Your model does not necessarily have to outthink their model. Your humans have to out-orient their humans.</p>



<p class="wp-block-paragraph">This becomes increasingly true as frontier capabilities bunch together above the threshold required for a particular task. In my own work, I increasingly find that I can use the strongest model to formulate or specify a solution and then hand most of the execution to a weaker model. For many problems, even that is overkill. I would readily bet on a well-oriented person using a slightly weaker model against a poorly oriented person using the strongest available model.</p>



<p class="wp-block-paragraph">And the frontier labs have no automatic orientation advantage. Quite the contrary: They are simultaneously fighting an extraordinary number of battles under extreme strategic distraction. They are building models, securing compute, raising capital, negotiating with governments, managing safety factions, defending themselves against critics, competing for talent, building consumer products, selling enterprise software, contemplating hardware and robotics, responding to geopolitical pressure, and trying to decide what sort of companies they are becoming. They may have a model advantage while suffering an orientation disadvantage. There is no reason to assume that organizations exceptionally good at building foundation models are exceptionally good at everything their models can be applied to.</p>



<p class="wp-block-paragraph">This is another reason pacing can be strategically productive. It creates room for orientation. In an environment saturated with FUD and “resistance is futile” rhetoric, organizations can lose before competing because they assume frontier capability automatically determines every downstream contest. It doesn’t. Superior orientation does.</p>



<h2 class="wp-block-heading"><strong>From one S-curve to the next</strong></h2>



<p class="wp-block-paragraph">Put all of this together and Amodei’s proposal starts to look different. His concern is that capability is outrunning safety. I think capability may be outrunning almost everything.</p>



<p class="wp-block-paragraph">It is outrunning our ability to deploy it productively. It is outrunning classroom experimentation, organizational adaptation, security practice, mathematical digestion, physical infrastructure, and human attention. It may be outrunning our ability to distinguish actual threats from theological ones.</p>



<p class="wp-block-paragraph">And it is almost certainly outrunning the decompression and recovery cycles of some of the people charged with making the most consequential decisions about it.</p>



<p class="wp-block-paragraph">An accelerationist should care about every one of these things precisely because an accelerationist wants acceleration.</p>



<p class="wp-block-paragraph">The mistake is to identify acceleration with the derivative of a single visible variable: benchmark scores, parameter counts, inference budgets, training compute, or whatever happens to define the current frontier. Technological progress is a coupled system. Accelerating one component beyond the absorption capacity of its complements eventually stops accelerating the system. The problem becomes especially acute near the top of an S-curve, where enormous resources can be consumed eking out diminishing improvements while the exploration necessary to find the next curve is crowded out.</p>



<p class="wp-block-paragraph">There is a useful precedent in the history of the PC industry. For years, processor clock frequency functioned as the wonderfully simple consumer metric for progress: 486 MHz was better than 400 MHz; 1 GHz was better than 800 MHz; higher number, faster computer. Manufacturers had every reason to compete on the legible scalar, and consumers learned to buy it. Eventually this became the “<a href="https://www.theguardian.com/technology/2002/feb/28/onlinesupplement3" target="_blank" rel="noopener">megahertz myth</a>.” Different architectures could do very different amounts of useful work per clock cycle, while pushing frequency upward ran increasingly hard into heat and power constraints. By the mid-2000s, the industry was moving toward multicore designs and a more complicated understanding of performance in which throughput, architecture, workload, thermal limits and performance per watt all mattered. Intel itself acknowledged at the time that as computer usage diversified, factors other than clock speed were becoming increasingly important to platform performance.</p>



<p class="wp-block-paragraph">AI benchmark culture looks increasingly like the early stages of the same mistake. A benchmark is useful because it compresses a complicated question into a number. When capability is scarce and improvements are large, the number may track value surprisingly well. As systems become overpowered for more uses, however, the proxy begins to detach from what customers actually care about. A model that goes from 87 to 91 on some benchmark may represent an impressive scientific achievement while producing essentially zero additional value for a company whose relevant workload was already handled adequately at 75.</p>



<p class="wp-block-paragraph">This suggests a path to frontier pacing that does not require a council of frontier CEOs deciding how quickly everyone is allowed to move.</p>



<p class="wp-block-paragraph"><em>Customers can simply become harder to impress.</em></p>



<p class="wp-block-paragraph">Enterprise buyers can demand demonstrated improvements on their actual workloads rather than accepting leaderboard gains as evidence of value. Developers can (and already do) route work to the cheapest model that clears the capability threshold rather than reflexively calling the smartest one. Researchers can value useful scientific infrastructure, explanation and reusable knowledge rather than merely celebrating another famous benchmark or theorem knocked down. Investors can become less impressed by capital expenditure whose primary justification is preserving position on a frontier whose marginal economic value is falling. Users can decline to upgrade when the previous generation is already good enough. Current enterprise behavior already points in this direction: Cheaper and open-weight models are becoming attractive precisely because many workloads do not require frontier intelligence, while buyers increasingly demand measurable returns rather than capability in the abstract.</p>



<p class="wp-block-paragraph">None of these mechanisms requires anybody to agree upon a socially optimal rate of AI development. They simply improve the feedback signal facing producers. The market stops saying “more intelligence, at almost any price” and begins saying “show me what this additional intelligence is for.”</p>



<p class="wp-block-paragraph">There are supply-side versions too. As power becomes a binding constraint, performance per watt and useful inference per dollar should matter more than sheer training scale. As inference proliferates toward edge devices and private deployments, latency, reliability, privacy and local controllability become competitive dimensions. As organizations discover that weaker models can execute plans produced by stronger ones, heterogeneous model portfolios should compete with monolithic frontier consumption. Open-weight and decentralized systems can keep proprietary labs honest by making “good enough” intelligence cheap and difficult to monopolize. A mature AI market should develop more dimensions of performance precisely as the mature processor market did.</p>



<p class="wp-block-paragraph">This is the sort of pacing I would prefer: not a speed limit but a richer scoreboard.</p>



<p class="wp-block-paragraph">The objective should therefore not be maximum speed. It should be managed tempo for maximum actual progress. Sometimes that means sprinting. Sometimes it means dwelling at a capability level while applications, institutions, infrastructure, science, and humans catch up. Sometimes it means letting one subsystem race ahead while another rests.</p>



<p class="wp-block-paragraph">Sometimes it means deliberately leaving expensive capability unused.</p>



<p class="wp-block-paragraph">That last possibility may be the hardest one for AI culture to accept, because we are still psychologically adapting to the idea that intelligence might actually be abundant.</p>



<p class="wp-block-paragraph">A true sense of abundance does not require you to max out the bounty. Nobody hyperventilates because free oxygen might disappear before they get their fair share. When something is genuinely abundant, you can waste it. You can use a frontier model for a trivial question. You can use a weaker model because it is good enough. You can leave tokens unused. You can spend a week doing something that doesn’t involve AI. You can allow an extraordinarily powerful model to sit idle while you think.</p>



<p class="wp-block-paragraph">The mark of abundance is waste, including nonuse.</p>



<p class="wp-block-paragraph">Compulsive token-maxxing is in this sense still a scarcity behavior. So is compulsive benchmark-maxxing, compute-maxxing, and capability-maxxing. Train now because somebody else will. Deploy now because the window might close. Consume all the intelligence available because leaving any unused feels like falling behind. An industry behaving this way may possess an abundance of intelligence without yet having developed an abundance mentality.</p>



<p class="wp-block-paragraph">Slack is not necessarily the enemy of acceleration. Slack is where people recover, where institutions adapt, where strange experiments happen, where understanding catches up with proof, where neglected complements receive attention, and where somebody finally notices that the old S-curve is flattening and another one is waiting nearby.</p>



<p class="wp-block-paragraph">So yes, pace the frontier. But don’t turn the frontier labs into a cartel to do it. Let safety work catch up, but also let customers become bored with vanity benchmarks. Let exhausted researchers sleep. Let mathematicians digest their proofs. Let schools figure out what to do with the models they already have. Let robotics catch up. Let organizations learn to orient themselves in a world where intelligence is cheap. Let markets discover that efficiency, reliability, privacy, integration and domain-specific usefulness sometimes matter more than another few points on a benchmark. Let us discover which risks are real, which bottlenecks have moved, and which parts of the world turn out to be playable.</p>



<p class="wp-block-paragraph">Then, when the situation calls for it, accelerate again.</p>



<p class="wp-block-paragraph">Slow is smooth. Smooth is fast.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-accelerationist-case-for-frontier-pacing/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI</title>
		<link>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/</link>
				<comments>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/#respond</comments>
				<pubDate>Tue, 22 Sep 2026 11:00:29 +0000</pubDate>
					<dc:creator><![CDATA[Nicole Butterfield]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Innovation & Disruption]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19763</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-Sovereignty—Bargaining-with-Big-Tech-and-the-Promise-of-Full-Stack-Open-Source-AI.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-Sovereignty—Bargaining-with-Big-Tech-and-the-Promise-of-Full-Stack-Open-Source-AI-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The early rapid expansion of AI capabilities that focused on frontier models was largely ushered into the world by a few powerful, US-based AI labs. Open-weight models released from labs in China, early on from DeepSeek, and later from Moonshot, Z.ai, and others, have in part disrupted that dominance. But growing concerns about the concentration [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The early rapid expansion of AI capabilities that focused on frontier models was largely ushered into the world by a few powerful, US-based AI labs. Open-weight models released from labs in China, early on from DeepSeek, and later from Moonshot, <a href="http://z.ai" target="_blank" rel="noopener">Z.ai</a>, and others, have in part disrupted that dominance. But growing concerns about the concentration of power have led to discussions about the need for at least some level of <em>AI sovereignty</em>.</p>



<p class="wp-block-paragraph">AI sovereignty doesn’t necessarily imply total control of your AI stack. It holds the promise of having more localized security and privacy, better adherence to local jurisprudence (for example EU AI and data laws), more dependable service, and potentially more culturally specific outputs from the AI technologies used within a specified border or region. Negotiating interdependence is not necessarily a problem, but having an array of tools beyond just open-weight models and options in those negotiations beyond just commercial offerings is imperative.</p>



<p class="wp-block-paragraph">Unsurprisingly, the same AI labs that gave rise to the need for AI sovereignty are also pushing their own solutions to the problem. While local institutions consider and even adopt some of these initial “sovereignty” offerings from these labs, open source AI technologies may offer a more promising horizon, more flexibility, and a means to manage dependencies. Tim O’Reilly argues open source AI is a potential opening for <a href="https://www.oreilly.com/radar/ai-sovereignty-and-the-architecture-of-participation/" target="_blank" rel="noopener">greater participation</a> in the future of AI’s development.</p>



<p class="wp-block-paragraph">From the underlying chip technology that is necessary for model training and inference to the cloud and data infrastructure that enables model development, companies such as NVIDIA, OpenAI, Google, Microsoft, and AWS have begun to stake out their own territory to maintain relevance within the global push toward AI sovereignty. Stanford University’s Human-Centered Artificial Intelligence Lab (HAI) lays out the different approaches and offerings these labs have developed in its report <em><a href="https://hai.stanford.edu/policy/the-commercial-landscape-of-ai-sovereignty-offerings" target="_blank" rel="noopener">The Commercial Landscape of AI Sovereignty Offerings</a></em>. It argues that while these labs “promise that countries will own their AI stack, [they also] deepen dependencies on U.S. Big Tech.”</p>



<p class="wp-block-paragraph">There are some non-US-based commercial alternatives that offer their own “full stack” solutions or AI sovereignty for specific layers of the stack. Companies in Europe, the Gulf region, Asia, and elsewhere are positioning themselves as local alternatives to US tech oligarchs. According to the <a href="https://hai.stanford.edu/policy/the-commercial-landscape-of-ai-sovereignty-offerings" target="_blank" rel="noopener">HAI report</a>, “Many of the most mature and advanced companies are actively backed by their governments. In these cases, sovereignty is not just a marketing claim but a stated policy objective, with governments directing funding, structuring procurement, and, in some cases, selecting specific companies to build out domestic AI capacity on their behalf.”</p>



<p class="wp-block-paragraph">These types of collaboration can both enable independence from US labs but may also create openings for political intervention. Claims about censorship and control of Chinese models emerged quickly after DeepSeek’s initial 2025 release. More recently, there have been <a href="https://apnews.com/article/artificial-intelligence-chatbots-censorship-bias-free-speech-fed8fdbf90751c10fe77b77832e0ffba" target="_blank" rel="noopener">probes into how US models</a> may limit certain types of discourse. Moreover, the HAI report points out that often the offerings of these alternative providers still rely on the underlying technologies, specifically chips and cloud infra, of the US labs.</p>



<p class="wp-block-paragraph">The proliferation of commercial offerings provides the space for diversification or potential leverage to negotiate better terms for collaboration, even with the dominant players. Open source AI technologies also play an important role in creating opportunities for even more diversification and greater sovereignty. As HAI argues, “Sovereignty strategies that do not consider the role of open-source AI risk normalizing fragmentation and political overreach.”</p>



<p class="wp-block-paragraph">In order for open source AI to counter the diversification of commercial sovereignty offerings, these technologies must also proliferate beyond open-weight models. Arguing for a “<a href="https://www.oreilly.com/radar/ai-sovereignty-and-the-architecture-of-participation/" target="_blank" rel="noopener">federated system</a>” of open source AI that enables sovereignty based on an “architecture of participation,” Tim O’Reilly writes that “the right infrastructure to let us satisfy both goals [of being everywhere and allowing everyone to have a say] will be a federation of models, a federation of protocols and code, and a federation of capacity. We need an architecture of participation all the way down the stack, and all the way up.” A key technology in the expansion of the open source AI stack these days are agent harnesses.</p>



<p class="wp-block-paragraph">In a recent article, Mozilla CTO <a href="https://www.oreilly.com/people/raffi-krikorian/" target="_blank" rel="noopener">Raffi Krikorian</a> argues that “the orchestration layer above the [model] weights is where capability is concentrating, and closed labs are already welding it shut”; therefore, it’s imperative to build on open harnesses, not just models. Commercial offerings that have dominated thus far include Claude Code and Codex. OpenClaw offered an initial disruption and promise for open source in late 2025, though the creator was quickly absorbed into OpenAI’s organization. While big tech labs continue to absorb when, who, and what they can, NousResearch’s self-improving Hermes agent harness has also garnered substantial attention now with over 230,000 stars on GitHub. More recently harnesses such as Pi and DeepSeek Harness are expanding that open source offering, heeding Krikorian’s call.</p>



<p class="wp-block-paragraph">Beyond agent harnesses, some of the strongest open source projects are developing in the less visible layers. Inference engines such as vLLM, SGLang, llama.cpp, and ONNX Runtime make it possible to serve a range of models efficiently across data centers, regional clouds, personal computers, and edge devices. Ray, which was developed by researchers at UC Berkeley, distributes demanding AI workloads. Ollama lowers the barrier to running models locally. Together, these projects give institutions more freedom to change models, hardware, and hosting providers without rebuilding an entire system around another company’s proprietary platform.</p>



<p class="wp-block-paragraph">Other fast-growing projects are filling out the data, interoperability, and accountability layers of the stack. The Model Context Protocol and Agent2Agent Protocol offer open standards through which agents can connect to tools and to one another. Qdrant, Chroma, Milvus, and LanceDB provide open infrastructure for storing and retrieving institutional knowledge. MLflow, Opik, and OpenLLMetry allow developers to evaluate, trace, and monitor AI applications without surrendering operational data to a closed dashboard. The <a href="https://www.aipotluck.org/map/gap-map" target="_blank" rel="noopener">AI Potluck Gap Map</a> classifies inference, deployment, and agent protocols as mature open ecosystems but identifies resiliency gaps in storage and observability, where fewer fully open projects occupy the leading tier. These gaps point toward an important investment agenda. Sovereignty will depend not on finding a single open replacement for Big Tech but on sustaining interoperable public alternatives across every consequential layer of the stack.</p>



<p class="wp-block-paragraph">Even still, the looming threat of acquisitions and absorption of open source is persistent. The fintech company <a href="https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter" target="_blank" rel="noopener">Stripe recently bought OpenRouter</a>, a platform that allows developers to access various models and has become a primary hub for accessing and routing open-weight models in particular. NVIDIA <a href="https://arstechnica.com/ai/2026/08/report-nvidia-to-acquire-ai-model-repository-hugging-face-for-13-billion/" target="_blank" rel="noopener">has acquired Hugging Face</a>, one of the key players for the open source AI ecosystem, and it has also just settled a <a href="https://finance.yahoo.com/technology/ai/articles/nvidia-7-billion-poolside-deal-224313075.html" target="_blank" rel="noopener">licensing deal with Poolside</a>, which develops open-weight coding models, purportedly to avoid the oversight of complete acquisition.</p>



<p class="wp-block-paragraph">AI sovereignty may mean the necessity of “calibrating interdependence” with US Big Tech solutions or replacing a foreign dependency with a domestic one for the time being. But it must also entail building the technical capacity, open infrastructure, and participatory institutions needed to preserve genuine choice across the entire AI stack. Open source will continue to be vulnerable to commercial absorption, and efforts to counter that must become more robust.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/ai-sovereignty-bargaining-with-big-tech-and-the-promise-of-full-stack-open-source-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Build a Sports Concierge Agent with Chester Ismay</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/#respond</comments>
				<pubDate>Mon, 21 Sep 2026 17:01:35 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19754</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Turn sports schedules and personal preferences into one weekly recommendation]]></custom:subtitle>
		
				<description><![CDATA[Chester Ismay, a data science educator and AI consultant, created schedule viewers to keep up with the sports he follows, including the WNBA, NFL, NBA, and Premier League. But he still had to decide which games deserved his attention each week. In this episode of Zero to Agent in 30 Minutes, Chester built a sports [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Chester Ismay, a data science educator and AI consultant, created schedule viewers to keep up with the sports he follows, including the WNBA, NFL, NBA, and Premier League. But he still had to decide which games deserved his attention each week.</p>



<p class="wp-block-paragraph">In this episode of <em>Zero to Agent in 30 Minutes</em>, Chester built a sports concierge agent to surface the games he should watch. It reads his preferences and current schedules, then sends a weekly summary to his phone. He took the audience <a href="https://ismayc.github.io/never-miss-a-game/" target="_blank" rel="noopener">through the setup</a>, which combines a prompt file, limited tool permissions, a schedule, and notifications.</p>



<iframe width="800" height="450" src="https://www.youtube.com/embed/W1_p5WRstFw?si=YRSKXXwW7H1h6tEw" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<h2 class="wp-block-heading"><strong>How to create your own sports concierge</strong></h2>



<ol class="wp-block-list">
<li><strong>Define your preferences.</strong> Chester started with a structured preferences file that identifies the teams he follows and adds context about why he follows them. He built <a href="https://ismayc.github.io/never-miss-a-game/preferences-editor.html" target="_blank" rel="noopener">a simple web interface</a> for editing those preferences rather than working directly with the underlying JSON.</li>



<li><strong>Give the agent access to current schedule data.</strong> His existing sports viewers pull schedule information from sources such as ESPN and store it in repositories on GitHub. A read-schedules tool, run with Node, pulls the latest schedule files and combines them with Chester’s preferences, giving the agent the information it needs without requiring it to search for each game on its own.</li>



<li><strong>Write the agent’s policy.</strong> Chester spent most of the walkthrough on <a href="https://raw.githubusercontent.com/ismayc/never-miss-a-game/refs/heads/complete/concierge.md" target="_blank" rel="noopener">CONCIERGE.md</a>, the file that defines the agent’s job. The policy sets the goal, identifies the data sources, and defines the rules for deciding which games to recommend. It also specifies how to handle finished tournaments and duplicate matchups, along with the expected output and delivery format. Chester noted that when the output misses the mark, he goes back to the policy and adds more detail to the instructions or adjusts them to better match with the goals of the project.</li>



<li><strong>Run the agent with limited permissions.</strong> Chester used Claude Code to read the files, apply the policy, and generate the weekly recommendations. He configured permissions so the agent worked only with the files and tools required for the task.</li>



<li><strong>Schedule delivery and check the results.</strong> Chester used <a href="https://www.launchd.info/" target="_blank" rel="noopener">launchd</a> on his Mac to run the concierge every Wednesday and <a href="https://ntfy.sh/" target="_blank" rel="noopener">ntfy</a> to send the result to his phone. He checks the recommendations against the underlying schedules and uses tests and multiple data sources to catch errors. Time zone handling required another adjustment. Games could fall on the wrong day when the system defaulted to UTC, so Chester added explicit time zone instructions.</li>
</ol>



<h3 class="wp-block-heading"><strong>Coming up next</strong></h3>



<p class="wp-block-paragraph">Next week, AI engineer Sajal Sharma gives an agent its own computer in the cloud using services such as E2B and Scrapybara. He’ll demonstrate how sandboxing lets an agent install packages, run code, drive a browser, and control a remote desktop.</p>



<p class="wp-block-paragraph"><em>Follow along with</em> Zero to Agent in 30 Minutes <em>on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O’Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-sports-concierge-agent-with-chester-ismay/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Post-training Process OpenAI Used for ChatGPT</title>
		<link>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/</link>
				<comments>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/#respond</comments>
				<pubDate>Mon, 21 Sep 2026 10:55:18 +0000</pubDate>
					<dc:creator><![CDATA[Sharon Zhou]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19751</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-post-training-process-OpenAI-used-for-ChatGPT.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-post-training-process-OpenAI-used-for-ChatGPT-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Walk through the pipeline that made ChatGPT such an improvement over GPT-3]]></custom:subtitle>
		
				<description><![CDATA[This is the third post in a four-part series about post-training. If you missed them, read part 1 and part 2. The final post, on implementing your own pipeline, will be coming October 7. Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>This is the third post in a four-part series about post-training. If you missed them, read <a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">part 1</a> and <a href="https://www.oreilly.com/radar/the-two-pillars-of-post-training-reinforcement-learning-and-supervised-fine-tuning/" target="_blank" rel="noopener">part 2</a>. The final post, on implementing your own pipeline, will be coming</em> <em>October 7</em>.</p>
</blockquote>



<p class="wp-block-paragraph">Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished in the frontier models you know and love.</p>



<p class="wp-block-paragraph">Remember our prompt “<a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">Why do people like golden retrievers?</a>” GPT-3 would often answer nonsensically. But that all changed in November 2022, with the launch of ChatGPT. Now “Why do people like golden retrievers?” actually returned a reasonable response like “Because they are affectionate, patient, and make excellent family pets” no matter who was typing (with no weird formatting tricks to consider). Anyone who could send a text message could get a response back on any topic.</p>



<p class="wp-block-paragraph">I’ll cover some of those behavior changes below, then take you through ChatGPT’s training pipeline as described in OpenAI’s InstructGPT paper.</p>



<h2 class="wp-block-heading">Conversational and helpful</h2>



<p class="wp-block-paragraph">The most visible impact of post-training is models that can chat with you and hold a relatively long conversation. This sounds simple, but it’s not.</p>



<p class="wp-block-paragraph">Being conversational means more than responding to a question with an answer. The model needs to recognize when a question is ambiguous and ask for clarification or make the right assumptions in a quick response. It should adjust its tone and detail level to the context, for example being brief for a quick factual question, but thorough for a learning-oriented one. The model should be coherent across multiturn conversations without losing the thread. It also needs to handle messy real-world inputs: You attach a giant PDF and ask it to find one specific clause, and it should either find it or tell you it can’t, not hallucinate an answer.</p>



<h2 class="wp-block-heading">Safety and alignment</h2>



<p class="wp-block-paragraph">Post-training is also the primary mechanism for making models safe. Safety in this context means a few things:</p>



<ul class="wp-block-list">
<li>Refusing to generate harmful content (like instructions for creating weapons, when asked)</li>



<li>Avoiding biased or discriminatory outputs</li>



<li>Not making up information when unsure (hallucination reduction)</li>



<li>Respecting user privacy</li>
</ul>



<p class="wp-block-paragraph">However, you can define safety rules in whatever way you want and teach the model to abide by them, within the limits of what your reward signals can capture. If you think cats are unsafe, because you’re a dog person, you can teach the model that in post-training—as long as you can properly encode that into a reward signal.</p>



<p class="wp-block-paragraph">Safety is often in tension with helpfulness. On one extreme, a model that’s too conservative will refuse reasonable requests, something that has frustrated many users. On the other extreme, a model that’s too permissive will comply with harmful ones. Navigating this trade-off is difficult. Ultimately, it comes down to determining where to draw the line, which is (as of today) a human decision within labs. Post-training is the tool to implement wherever the line is drawn.</p>



<h2 class="wp-block-heading">Tool use and function calling</h2>



<p class="wp-block-paragraph">Tool use is one of the most practically important capabilities enabled by post-training. Tools include search engines, APIs, calculators, databases, and code interpreters. Being able to hit a search engine alone allows the model to not hallucinate, given its own knowledge cutoff. Tools are extremely useful ways for models to interact with the world, and are fundamental components in building agents.</p>



<p class="wp-block-paragraph">Tool use is a set of new behaviors. The model needs to recognize when a user’s request would benefit from an external tool. It needs to know which tools are available to it, and not hallucinate a tool. It needs to formulate a correct API call with the right parameters. It needs to interpret the results that come back and incorporate them into a natural language response. It needs to do all of this seamlessly, without the user needing to know the details of the underlying tool.</p>



<p class="wp-block-paragraph">This is taught almost entirely through SFT, at least initially. The training data includes many examples of conversations where the model correctly decides to invoke a tool that it has access to, constructs the right call, and processes the result. RL can further improve tool use by rewarding the model for correct tool invocations and penalizing unnecessary or incorrect ones.</p>



<p class="wp-block-paragraph">As an example of tool use, let’s say you’re building a veterinary appointment scheduling assistant. A user asks: “My golden retriever has been limping since yesterday. Can I see Dr. Patel this afternoon?” A pretrained model might generate plausible but fictional appointment times. A post-trained model with tool use instead calls the clinic’s scheduling API, checks Dr. Patel’s availability, and responds: “Dr. Patel has an opening at 3:15pm today. I’ve tentatively held it for you. Should I confirm?” The model needed to decide if the user’s intent was urgent, select the right tool, construct the API call with the right veterinarian and time constraints, and present the result conversationally.</p>



<p class="wp-block-paragraph">Tool use has expanded through the <a href="https://www.anthropic.com/news/model-context-protocol" target="_blank" rel="noopener">Model Context Protocol</a> (MCP), a lightweight standard for connecting models to external services like Gmail, GitHub, or a company’s internal databases. Rather than building custom integrations for each tool, MCP provides a standard interface that any API can plug into, and different frontier models have now included learning MCP in their post-training recipes. Agentic frameworks take this further by allowing models to chain multiple tool calls together to accomplish common multistep tasks more easily.</p>



<h2 class="wp-block-heading">Reasoning (“thinking”)</h2>



<p class="wp-block-paragraph">Reasoning models, or models that are trained to “think” before they answer, are an exciting result of post-training. Rather than producing an immediate response, these models generate an internal chain of thought, working through the problem step-by-step, before arriving at a final answer. As a result, their answers are more often correct than nonreasoning models that might guess at an answer.</p>



<p class="wp-block-paragraph">This capability has an interesting relationship with pretraining and post-training. The raw ability to reason is latent in pretrained models; they’ve been trained on text that includes mathematical proofs, logical arguments, scientific analyses, and code with comments explaining the logic. But pretrained models don’t default to reasoning. They default to pattern-matching, which often produces plausible-looking but incorrect answers.</p>



<p class="wp-block-paragraph">Reasoning models dramatically outperform standard models on tasks that require multistep logic: mathematical problem-solving, complex coding, scientific analysis, and planning. The improvements are not incremental. On the 2024 AIME exam, GPT-4o was only able to get 12% of problems correct on average. OpenAI’s o1 reasoning model solved 74% off the bat, with a single attempt. With 1,000 attempts and a learned scoring function to rerank the attempts, <a href="https://openai.com/index/learning-to-reason-with-llms/" target="_blank" rel="noopener">it reached 93%</a>, a result placing it among the top 500 students who took the AIME math exam in the US.</p>



<p class="wp-block-paragraph">More capable reasoning requires more compute, both during training and at inference time. Scaling laws meet post-training. Models that have learned to spend more inference (test-time) compute on reasoning tend to reach better answers and therefore exhibit higher intelligence. For some frontier reasoning models, the RL post-training phase uses as much compute as the entire pretraining phase.</p>



<p class="wp-block-paragraph">The cost is not only in post-training compute but also in inference (test-time) tokens and latency. Reasoning takes up a lot of tokens and can result in a longer time to get a response back to the user. But the type of request matters. For a quick factual question, you don’t need reasoning. For a complex technical problem, the extra latency is well worth it. This is something that model providers can modulate during post-training.</p>



<h2 class="wp-block-heading">The classic ChatGPT pipeline</h2>



<p class="wp-block-paragraph">As I mentioned above, the first post-training pipeline that captured global attention was ChatGPT’s, and it drew on the pipeline described in the <a href="https://arxiv.org/abs/2203.02155" target="_blank" rel="noopener">InstructGPT paper</a>. While modern systems use more advanced approaches today, this classic pipeline remains the conceptual foundation for nearly all alignment methods.</p>



<p class="wp-block-paragraph">The pipeline has three stages, each building on the previous one:</p>



<ol class="wp-block-list">
<li>Supervised fine-tuning (SFT) on human demonstrations</li>



<li>Training a reward model on human preference comparisons</li>



<li>Reinforcement learning with human feedback (RLHF) to optimize the main model using the reward model</li>
</ol>



<h3 class="wp-block-heading">Stage 1: SFT on demonstrations</h3>



<p class="wp-block-paragraph">The first stage is straightforward and teaches the model to follow instructions and behave like an assistant.</p>



<p class="wp-block-paragraph">OpenAI contracted ~40 human labelers to label their data, and they were careful to filter for people who were good at identifying harmful outputs. The labelers had to write ideal responses to prompts. But what’s interesting is that the prompts came from two sources: (1) prompts submitted by real users through the OpenAI API and (2) prompts that labelers wrote themselves. The users had to write prompts too, because these were the days before ChatGPT. There weren’t that many real users with instruction-like prompts through the API to collect.</p>



<p class="wp-block-paragraph">The prompts were diverse and mostly in English. There’s also an extensive data cleaning pipeline to remove duplicates and remove sensitive PII (personally identifiable information). Importantly, they split the training, validation, and test sets by human labeler. This is to avoid data leakage that could happen within a single user’s data between training and validation/testing.</p>



<p class="wp-block-paragraph">The resulting SFT dataset had ~13,000 prompts, all with human-labeled responses. The base model was GPT-3 at the time, a pretrained model without any post-training. Using SFT, they trained GPT-3 for 16 epochs, which was effective for the final RLHF model. This was interesting, because for the SFT stage alone, the model overfit after just 1 epoch, but ultimately SFT was an intermediate stage so they picked the best checkpoint for the final RLHF model. They also mixed in 10% pretraining data during this phase, because it would help the next RL phase.</p>



<p class="wp-block-paragraph">At this point, this SFT model could already be pretty useful: It could have a conversation and follow instructions, which is leaps and bounds beyond the pretrained GPT-3 checkpoint.</p>



<h3 class="wp-block-heading">Stage 2: Preference data and reward modeling</h3>



<p class="wp-block-paragraph">This next stage is training the reward model. The reward model needs to grade millions of responses during RL training. In the original method, OpenAI’s team mainly trained the reward model on responses from the SFT model. However, as the policy model is trained in the RL loop and generates new, and likely better, responses from its evolving checkpoints, the reward model needs to stay robust. As a result, they also continually updated the reward model using responses from new RL checkpoints over time.</p>



<p class="wp-block-paragraph">To train the reward model in InstructGPT’s RLHF pipeline, OpenAI needed pairwise comparisons of two model responses from one prompt, and a label for which one is better. For example, given “What’s 2+2?” and the responses are “4” and “Yes,” the label should say “4” is better than “Yes.” Note again that these are responses from the SFT model (and later, the RL-ed models during the RL training loop), not the pretrained base model. So labeling can only happen after you’ve SFT-ed your model. If you need to retrain that model, you likely need to relabel to make sure the reward model is trained on the right distribution of data pairs.</p>



<h4 class="wp-block-heading">Reward model training</h4>



<p class="wp-block-paragraph">The reward model was small at 6B parameters, for both efficiency and stability, and included a head that outputted a scalar reward. They had tried multiple sizes, but found this was more stable than using the original 175B main model. It was also more compute efficient, as the reward model would take up extra compute, for both inference and training, on top of training the main model itself. More recently, reward models have become a lot larger, but note that they don’t have to be the same model or same size model as the main model.</p>



<p class="wp-block-paragraph">To train the model, the loss was a cross-entropy loss that represented the log odds that someone would prefer one option over the other, in the pairwise comparison. This was done by taking the difference between the rewards of the preferred and unpreferred options. So in the example “What’s 2+2?,” if the reward model correctly assigns “4” a high reward and “Yes” a small reward, then the difference would be high and positive, and the loss would be small. However, if the reward model were to incorrectly assign “Yes” a higher reward than “4,” the difference would be high and negative, and the loss would be huge—discouraging it from outputting this result again.</p>



<p class="wp-block-paragraph">One of the big challenges in training the reward model was overfitting, and OpenAI found that training for only 1 epoch would help prevent that.</p>



<h4 class="wp-block-heading">Reward model data</h4>



<p class="wp-block-paragraph">The simplest way to get preference pairs is to generate two responses per prompt and have a labeler tell you which one was better. To make more efficient use of each prompt, instead the model would generate not 2 but 4–9 different responses per prompt that human labelers would rank from best to worst.</p>



<p class="wp-block-paragraph">Rankings can be transformed into pairwise comparisons, so it was an efficient way to collect those preference pairs. A ranking of N responses yields N-choose-2 pairs. For example, a ranking of 4 responses results in 6 pairs, a ranking of 9 results in 36 pairs. That means with 33K prompts and 4–9 responses ranked per prompt, there would be 200K–1.2M pairwise comparisons used to train a separate reward model. That’s a lot of data, from relatively efficient data labeling.</p>



<p class="wp-block-paragraph">This is a relatively efficient use of human annotations. Just compare it to SFT. It’s easier, cheaper, faster, and more reliable (higher agreement between people) than writing good responses from scratch, so this stage of human labeling wasn’t as tedious as in SFT.</p>



<p class="wp-block-paragraph">However, using the pairs from rankings wasn’t straightforward in training. The reward model would overfit if they mixed the pairs randomly, even in just 1 epoch, because the pairs for a single prompt were highly correlated with each other. So instead, they would train all the pairs from the same prompt as one element in a batch, and normalize it. This was also computationally more efficient to run and score all the N responses at once together, e.g., just score 9 times and reuse those calculations in this pass, rather than 36 times for each pair if mixed into the dataset.</p>



<p class="wp-block-paragraph">A quick note on terminology. This data is often called <em>preference data</em>, because it’s about collecting human preferences. The reward model can also be called a <em>preference model</em>.</p>



<h3 class="wp-block-heading">Stage 3: RLHF (reinforcement learning with human feedback)</h3>



<p class="wp-block-paragraph">At this point, you have an SFT model that can follow instructions, and a reward model that can score responses. The goal of RL is to continue training the SFT model to produce responses that the reward model scores highly. If the reward model is any good, the resulting model will produce responses that humans would prefer.</p>



<p class="wp-block-paragraph">The RL algorithm used was PPO. As you learned previously about RL terminology, the SFT model is the “policy” that takes actions (generating tokens) in an environment (the conversation). The reward model provides the reward after the policy generates a complete response, and a critic model calculates the expected reward, a baseline estimate that offers a more stable overall reward signal in training.</p>



<p class="wp-block-paragraph">Here are the critical steps. I’ve covered some of them before and will dive into others in detail later on in this section.</p>



<ol class="wp-block-list">
<li><strong>Sample a prompt</strong>. The prompt comes from the dataset of 31,000 prompts that were gathered from users organically using the API. No need for human labels.</li>



<li><strong>Generate a response</strong>. Then, the current policy generates a response. The current policy is the SFT model in the beginning, but as the policy updates, it’s a new model that generates responses to be graded. A single prompt-response pair is called a “rollout.” In practice, this all happens in a batch of rollouts.</li>



<li><strong>Calculate the reward</strong>. The reward model grades the full response with a reward. For every token position in the full response, they subtract a per-token KL divergence penalty between the current policy and the original SFT model. The per-token KL penalty and the reward model score added at the final token make up the per-token reward signal.</li>



<li><strong>Calculate the advantage</strong>. The critic estimates the expected reward for the full response, at each token position in generation (so with partial knowledge of the full response). The critic’s expected rewards and the per-token reward signals are combined, using an algorithm called GAE (<a href="https://arxiv.org/abs/1506.02438" target="_blank" rel="noopener">Generalized Advantage Estimation</a>), to estimate how much better the reward was compared to expected. This is the advantage. In InstructGPT, the critic was initialized with the same weights as the reward model, giving it a head start on estimating expected reward. </li>



<li><strong>Update the policy</strong>. Then, PPO updates the policy model’s weights with the advantage, pushing it towards rollouts with higher advantage and away from ones with lower advantage.</li>



<li><strong>(Optional) Mix in pretraining data and the pretraining objective in the policy model’s loss function</strong> to prevent catastrophic forgetting. </li>



<li><strong>Update the critic</strong> to better predict expected future rewards at each token position, by using the actual per-token rewards (from the reward model and KL penalty) as the targets in training.</li>



<li><strong>(Optional) Update the reward model</strong>. Collect new ranking data on the current best policy and train a new reward model. In practice, OpenAI did collect some data from the PPO models, but most was from the original SFT model.</li>



<li><strong>Repeat!</strong> This RL loop repeats over many prompts and many updates, with the latest policy always generating its responses.</li>
</ol>



<h4 class="wp-block-heading">The KL penalty</h4>



<p class="wp-block-paragraph">If you just let the model maximize the reward model’s score with no constraints, it finds weird, degenerate outputs that exploit quirks in the reward model to get high scores without actually being good responses. This is called “reward hacking,” and it’s one of the central problems in RLHF.</p>



<p class="wp-block-paragraph">To address this, OpenAI added a KL divergence penalty between the RL policy and the original SFT model (“reference policy”) in the reward calculation. In AI, KL divergence is a common method of measuring how different two probability distributions are. In this case, it would measure how different the RL policy is from the old policy and penalize being too far from it, essentially telling the model: You can optimize for higher reward, but you can’t drift too far from where you started. If the RL model starts producing outputs that look nothing like what the SFT model would produce, the penalty helps to pull it back by making the reward for those outputs lower.</p>



<p class="wp-block-paragraph">The total reward for a response becomes the reward model’s score minus the KL divergence from the SFT model, with a coefficient term that weighs how much to care about the KL divergence. If the coefficient is too low, you’re saying that you don’t need to penalize drift from the reference policy, and you’ll get reward hacking. Too high and the model barely changes from the SFT checkpoint.</p>



<p class="wp-block-paragraph">They also mixed in a significant amount of pretraining data during the RL phase, adding a pretraining loss alongside the RL objective. This was to prevent the model from degrading on general tasks from pretraining, like knowledge recall or coherent long-form creative text, as it optimized for reward. This is sometimes called the “<a href="https://arxiv.org/abs/2309.06256" target="_blank" rel="noopener">alignment tax</a>,” where you trade-off alignment for general capabilities, a type of “<a href="https://arxiv.org/abs/1612.00796" target="_blank" rel="noopener">catastrophic forgetting</a>.” This is an active area of research.</p>



<p class="wp-block-paragraph">So the final RL objective combined three things: (1) maximize the reward model’s score on prompted responses, (2) stay close to the SFT model via the KL penalty, and (3) maintain performance on pretraining data. This means improving on the things humans care about without losing what the model already knew how to do from SFT.</p>



<h4 class="wp-block-heading">Practical details</h4>



<p class="wp-block-paragraph">The critic reduces noise and makes training stable enough to make PPO work practically. In practice, OpenAI initialized the critic from the 6B-parameter reward model, since it’s already trained to predict reward and gives the critic a head start as it is further trained in the RL loop.</p>



<p class="wp-block-paragraph">The RL training was computationally expensive and involved running several models simultaneously: the policy model (the main model being trained), the critic (estimating the reward as a baseline, also being trained), the reward model (grading responses), and a copy of the SFT model (for computing KL divergence).</p>



<p class="wp-block-paragraph">That’s four models in memory at once. That’s a lot of GPU memory, especially when the policy and SFT models are 175B parameters! In addition to weights, the policy and value models also needed their gradients, optimizer states, and cached activations for backpropagation because they were being trained, which can actually multiply the per-model memory cost by 3-4 times. This is another reason the reward and critic models were kept at 6B.</p>



<p class="wp-block-paragraph">PPO also requires generating fresh rollouts during training, which is much slower than SFT where you already have all the data upfront. Each PPO training step also requires grading each rollout with the reward model, computing advantages with the critic, and updating both the policy and the critic. All these moving parts make the system harder to tune and debug compared to SFT, and harder to parallelize than pretraining.</p>



<p class="wp-block-paragraph">Hyperparameters like learning rate, the KL penalty coefficient term, the number of rollouts per batch, and the clipping ratio all matter and interact with each other. The whole system depends on the quality and representativeness of your human annotations. Many RL training runs fail or produce degenerate results.</p>



<h2 class="wp-block-heading">Getting it right</h2>



<p class="wp-block-paragraph">Getting all this right takes significant engineering effort and experience, but the first step is deeply understanding the pieces. Modern methods have addressed many of these issues, but the ideas from InstructGPT remain the foundation of post-training today.</p>



<p class="wp-block-paragraph">Ultimately, human evaluators compare all the models. People preferred the RLHF model’s outputs over the SFT model’s, and the SFT model’s over base GPT-3’s. Each stage of the pipeline added a large improvement in the model’s response quality.</p>



<p class="wp-block-paragraph">RLHF was extremely effective. In experiments, human evaluators preferred even a tiny 1.3B parameter RLHF model over the 175B parameter SFT model, most of the time. That’s a model over 100x smaller, trained with RL, beating a much larger model trained only with SFT. This made a strong case that how you train matters as much as how big your model is. Overall, the largest RLHF model still beat the smaller RLHF model.</p>



<p class="wp-block-paragraph">The RLHF model was also better at following explicit constraints in instructions, less likely to produce harmful outputs, and hallucinated less, though it didn’t eliminate hallucinations as you may remember when you first used ChatGPT (and even now).</p>



<p class="wp-block-paragraph">One caveat worth noting: The labelers who evaluated the final model were the same population who created the training data. When they tested with held-out labelers who hadn’t been involved in data creation, preferences for the RLHF model were still positive but less dramatic. The model was, to some degree, optimized for the preferences of a specific group of people. This means if you create the data to follow your preferences, the model will optimize for those.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>How to Get from AI-Assisted to AI Native</title>
		<link>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/</link>
				<comments>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/#respond</comments>
				<pubDate>Fri, 18 Sep 2026 21:38:53 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19736</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-as-an-enterprise-operating-system.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/AI-as-an-enterprise-operating-system-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[AI as an Enterprise Operating System, a keynote from Ai4 2026]]></custom:subtitle>
		
				<description><![CDATA[When considering the history of AI, Richard Sutton observed that brute force and compute scale has always trumped human expertise, and when you look for it, you can see this “bitter lesson” play out throughout tech history. In his keynote at Ai4 2026, Tim O’Reilly explains why grappling with the bitter lesson is the forge [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/5EcTnkCY5ww?si=Ft-Khy2zizSbj1g4" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph"> </p>



<p class="wp-block-paragraph">When considering the history of AI, Richard Sutton observed that brute force and compute scale has always trumped human expertise, and when you look for it, you can see this “bitter lesson” play out throughout tech history. In his keynote at Ai4 2026, Tim O’Reilly explains why grappling with the bitter lesson is the forge of effective AI corporate strategy, as companies figure out what to embrace and what to let go of. Drawing on his <a href="https://www.oreilly.com/radar/ai-as-an-enterprise-operating-system/" target="_blank" rel="noopener">recent conversations</a> with Trail of Bits CEO Dan Guido, Tim argues that AI&#8217;s business impact actually hinges on organizational adoption—the hard, unglamorous work of restructuring workflows, data, and incentives around what AI can do. Trail of Bits has modeled that process and <a href="https://trailofbits.com/?item=https-github-com-trailofbits-publications-blob-master-presentations-how-20we-20m" target="_blank" rel="noopener">documented it in a playbook</a> other companies can use. Here, Tim shares some of the practices, like capability ladders, shared config repos, and company-wide hackathons, that helped Trail of Bits make AI a structural component of its business. This doesn’t mean that AI-native companies “sit back and let the progress of AI carry us forward.” Human expertise still matters, and it’s often the differentiator that helps organizations rise above their competitors. As Tim concludes, “The world is full of great problems. And so if AI takes away and makes easy something small, celebrate it and go work on something big with the new powers that we’ve been given.”</p>



<h2 class="wp-block-heading">Takeaways</h2>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=egYeWccrMdXhFu-N&amp;t=153" target="_blank" rel="noopener">02.33</a></strong> <strong>The bitter lesson is real, and it can catch any of us.</strong><br>The bitter lesson is Richard Sutton’s contention that human expertise doesn’t really matter, that it will eventually be outmatched by computing scale. O’Reilly’s <em>Whole Internet User’s Guide &amp; Catalog</em> was the first catalog of websites and the first site on the web to have advertising. It grew into Global Network Navigator, which was the first web portal. But O’Reilly’s products were manually curated. Yahoo came along and expanded on these ideas, but O’Reilly and Yahoo were both beaten by Google, which simply threw a bunch of compute at the problem. Now ChatGPT has changed the game again.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=E8D9kntfeiYL_plj&amp;t=387" target="_blank" rel="noopener">06.27</a></strong> <strong>AI-native workflows require a different mindset.</strong><br>When O’Reilly set out to develop a product that assessed learners’ capabilities and gave them a skill path to level up, the team used AI as an assistant, to write quiz questions, for instance. But LLM chatbots can already identify skills when given context about a developer. Evolving toward an AI-native skill path builder meant reconceptualizing the product as a more interactive experience that reflects where capabilities are today. However, even the most well-thought-out workflow can be hindered by gaps in access or knowledge. As Trail of Bits CEO Dan Guido says, “You have to build a system in which expertise compounds.”</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=wPVC_gop7qBWgaGO&amp;t=628" target="_blank" rel="noopener">10.28</a></strong> <strong>AI adoption is a human problem.</strong><br>Moving up the framework for AI adoption from AI-assisted to AI-augmented to AI-native isn’t just a technical challenge. It’s psychological. Only 5% of Dan’s staff was actually on board when he started the transformation; 70% were just quietly going through the motions, and 20% were actively resistant. He traces this to a handful of biases: self-enhancing bias, opacity, intolerance for imperfection, and above all, identity threat, the fear that AI won&#8217;t just replace the work someone does but who they are. Getting teams on board requires the organization to reframe AI as a tool that enhances identity, not something that will take it away.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=Ro9rr9iaMN9T8_M-&amp;t=1004" target="_blank" rel="noopener">16.44</a></strong> <strong>A status ladder helps team members understand where they&#8217;re at and where to focus next. Hackathons compound that knowledge across the company.</strong><br>Trail of Bits has a three-level status ladder: not engaged with AI or actively resisting it, experimenting with AI, and building AI that strengthens the organization&#8217;s overall capability. Level zero isn&#8217;t treated as a skill gap. It&#8217;s treated as working against the company&#8217;s goals, and the other two levels get a more detailed capability matrix broken out by department, since what a security auditor does with AI looks nothing like what someone in accounting does. O&#8217;Reilly is building its own version of this, drawing on the technical and business skill data it already has across its platform. Trail of Bits runs a hackathon every two months, each with a stated objective and learning goals announced a week ahead. Success is measured not by what got shipped but by where people land on the capability ladder afterward. Then the work gets fed into a shared skill repo, giving the entire company a set of reusable artifacts, and what one hackathon turns up becomes something the next one can build on.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=rHHKVmNEYeiOm8ts&amp;t=1477" target="_blank" rel="noopener">24.37</a></strong> <strong>Turn scar tissue into infrastructure.</strong><br>Drew Breunig talks about the problem of <a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">prompt debt</a>: prompts that grow more complex and more tuned to one specific model until they&#8217;re no longer portable. Trail of Bits flattens this complexity by turning every failure into a global, copy-pasted fix hosted in a company-wide repository. They’ve also standardized the safety net, with sandboxes for different needs and a seven-day cooldown on every new package from outside that gets installed—rules the whole company follows. To make this all work, employees need the chance to try things out and iterate on their failures. Dan says the only real mistake he made was not giving people enough unstructured time to experiment.</p>



<p class="wp-block-paragraph"><strong><a href="https://youtu.be/5EcTnkCY5ww?si=AEUjP4JHTL2c0jn-&amp;t=1964" target="_blank" rel="noopener">32.44</a></strong> <strong>Human expertise still matters.</strong><br>AI can make companies more productive, but it’s not a magic weapon. It’s a <a href="https://www.oreilly.com/radar/writing-with-ai/" target="_blank" rel="noopener">medium</a> that people can use to share or extend their unique expertise and perspective. O’Reilly’s mission is to share the knowledge of innovators: You can think of the company as a matching marketplace for people who have expertise and people who need it. Agents offer a valuable new means of getting that expertise to customers in the tools they’re using to make business decisions. O&#8217;Reilly CTO Andrew Odewahn has noted that faster local decision-making has splintered central planning, so it&#8217;s harder than ever to get the big-picture view a good corporate decision needs. O&#8217;Reilly&#8217;s Expert MCP server lets customers access our content and use it to increase organizational intelligence. For instance, you can ask an AI tool to analyze a team’s workload and write a hiring case based on how O&#8217;Reilly&#8217;s own experts would review the request, and you’ll get a grounded argument with solutions authenticated by citations from actual practitioners. O&#8217;Reilly is building this capability into an organization-wide grounding layer it calls O&#8217;Reilly Expert Intelligence. It’s in beta now, and you can <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noopener">check it out</a>.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/how-to-get-from-ai-assisted-to-ai-native/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Software Factories, Light and Dark</title>
		<link>https://www.oreilly.com/radar/software-factories-light-and-dark/</link>
				<comments>https://www.oreilly.com/radar/software-factories-light-and-dark/#respond</comments>
				<pubDate>Fri, 18 Sep 2026 15:59:35 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19732</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Software-factories-light-and-dark.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Software-factories-light-and-dark-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Addy Osmani’s blog and is being republished here with the author’s permission. A software factory harnesses loops at scale. You can run the loop with humans in it (light factory), trading judgment and concentration against speed and breakage. Or you can ignore the humans (dark factory) and let those [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://addyosmani.com/blog/software-factories/" target="_blank" rel="noopener">Addy Osmani’s blog</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph"><strong>A software factory harnesses loops at scale. You can run the loop with humans in it (light factory), trading judgment and concentration against speed and breakage. Or you can ignore the humans (dark factory) and let those agents scope, build, and ship code without anyone reading the details. If people stop reading, though, they’ll stop understanding your software. Your hardest job now is knowing which checks to build and how much autonomy to delegate.</strong></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph">This idea of the software factory is a term that dates back to Bob Bemer’s paper, “<a href="https://ed-thelen.org/comp-hist/Bemer-EconomicsOfProgramProduction.pdf" target="_blank" rel="noopener">The economics of program production</a>,” given in 1968. For half a century, many have dreamed of a world in which software is a repeatable and instrumentable production process (analogous to stamping out car parts in a factory) rather than the isolated craft of individuals. Historically, this dream has generally (although not universally) fallen flat, in part because of the difficulty of stamping out ideas.</p>



<p class="wp-block-paragraph">But in the last two years, things have changed dramatically enough that now it makes sense to take a fresh look at the old dream. And since some subtleties can easily be glossed over, it’s worthwhile to be somewhat precise about exactly what’s really new and different, and what may be recurring traps, dressed up as new opportunities.</p>



<p class="wp-block-paragraph">Dex Horthy, co-founder of HumanLayer recently gave a great talk at the AI Engineer World’s Fair called “<strong>Harness Engineering is not Enough: Why Software Factories Fail</strong>.” worth checking out on this topic.</p>



<h2 class="wp-block-heading">The loop is the atom. The factory is the loop at scale.</h2>



<p class="wp-block-paragraph">Structure is everything, and it all starts with small units. The whole stack is really three concepts layered on top of each other: the loop, the harness, and the factory.</p>



<p class="wp-block-paragraph">A loop is one agent doing a single job on repeat: gather context, take an action, check the result, and go again until some condition is met. It is the smallest unit of agentic work, and everything above it is just loops stacked on loops.</p>



<p class="wp-block-paragraph">The point of <a href="https://addyosmani.com/blog/loop-engineering/" target="_blank" rel="noopener">loop engineering</a> is that you stop prompting the agent turn by turn and instead design the small system that prompts it for you.</p>



<p class="wp-block-paragraph">A harness is the walls around a loop: the sandbox it runs in, the tools it can reach, the memory that survives between runs, and the gates that decide what “done” means. The loop is the behavior; the harness is the environment that behavior runs inside.</p>



<p class="wp-block-paragraph">Hand a raw model no harness and it will happily spin forever. The harness is everything around the model that makes it useful and safe to run.</p>



<p class="wp-block-paragraph">A <strong>software factory</strong> is many harnessed loops running at once, fed by a queue of work and drained through a review gate into production, with humans owning the whole thing from above. It isn’t a bigger agent; it’s an org chart made of loops.</p>



<p class="wp-block-paragraph">The final paradigm shift is moving from writing code to building and running the factory that writes it. The unit of work shifts up a level, to the loop, the harness, and the flow between them, rather than the individual code diff.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab50590564c6&quot;}" data-wp-interactive="core/image" data-wp-key="6ab50590564c6" class="aligncenter is-resized wp-lightbox-container"><img decoding="async" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://addyosmani.com/assets/images/software-factories/loop-harness-factory.svg" alt="Loop wrapped into a harness, run many times as a factory" style="aspect-ratio:3.052949594530132;width:840px;height:auto"/><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Loop → harness → factory. A factory isn’t a smarter agent; it’s many harnessed loops feeding one review gate, with a human owning the outer loop.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">The factory, drawn</h2>



<p class="wp-block-paragraph">The central slide Dex spent most time on was brilliant because it’s a clarifying wiring diagram that visualizes what otherwise is an obvious loop. Here’s my take on it:</p>


<div class="wp-block-image">
<figure class="aligncenter is-resized"><img decoding="async" src="https://addyosmani.com/assets/images/software-factories/agentic-software-factory.svg" alt="The agentic software factory as a closed loop" style="width:840px;height:auto"/><figcaption class="wp-element-caption"><em>The factory is a closed loop:</em> <em>Intent and production signals feed a queue, the harness builds, automated checks and review gate it, deploy ships it, and monitoring turns production back into signals.</em></figcaption></figure>
</div>


<p class="wp-block-paragraph">Intent flows from the vision of engineering leadership and directly from engineers into a queue of work. Signals driven by incidents and user requests drive the same queue. The harness picks an item from the queue and builds a change for it. Beyond the harness, automated checks make changes safe enough to let into production. These automated checks run at once without any conscious involvement from engineers, thanks to CI, tests, static analysis, and scanning of all kinds. The review gate is the only decision point here. After approval, changes are deployed and monitored in production, with monitoring data feeding back into the signals that kicked the loop into motion to begin with.</p>



<p class="wp-block-paragraph">By and large, every box in this diagram is almost zero cost: generation, tests, scanning. They all run at scale for negligible cost. There’s only one expensive box that proves stubbornly resistant to scaling, and that’s the review gate. That shiny amber box is “judgment,” and where the crux of the argument about whether we can make development faster and more frequent resides.</p>



<h2 class="wp-block-heading">Why we call it “dark”</h2>



<p class="wp-block-paragraph">A dark factory runs with the lights physically off because the only things on the floor are machines, which don’t need light to see. A dark software factory operates similarly, as code ships that no human has read and is verified only by other machines.</p>



<p class="wp-block-paragraph">The image is borrowed from manufacturing. Its origins are physical rather than digital, rooted in facilities where the lights are turned off and the work is carried out by robots. <a href="https://en.wikipedia.org/wiki/Lights_out_(manufacturing)" target="_blank" rel="noopener">FANUC in Japan has been running lights-out factories of this sort since 2001</a>. Xiaomi, in 2024, opened a heavily automated dark factory of its own. What these have in common is a product assembled and shipped without a single human having read any of it. The “dark” comes in when that act of reading is removed from the process.</p>



<p class="wp-block-paragraph">I’m not borrowing the concept for its vibe or as an insult. For all its creepy buzz, “dark” here is a simple physical claim: the original factory floor, but without light. In software, the floor is the diff. Whoever wrote the diff, whoever reviewed it, whoever shipped it, those humans are gone, and what remains is a diff verified only by the machines that built it.</p>



<p class="wp-block-paragraph">This is a surprisingly easy thing to do, at least at first. It’s easy because that missing review step gets in the way of everything. Its absence makes your perception of your team’s vertical throughput seem suddenly and radically higher. It feels as if you’ve broken the sound barrier. For all its apparent ease, it’s harder than it seems to survive those dark workflows, with all their buried costs.</p>



<h2 class="wp-block-heading"><strong>Harness engineering is not enough</strong></h2>



<p class="wp-block-paragraph">The harness of orchestration, sandboxed prototyping, and tool calling as models interact with the world and each other will become increasingly powerful and effective. However, there’s an inherent in-model failure in trying to keep up with codebase quality over the long game and through additive changes, and I think there’s good reason to believe that models alone will ultimately lose that battle against <a href="https://addyosmani.com/blog/comprehension-debt/" target="_blank" rel="noopener">comprehension debt</a>.</p>



<p class="wp-block-paragraph"><strong>Comprehension debt</strong> is the widening gap between how much code exists and how much any human still understands. A dark factory doesn’t pay it down; it takes it on as fast as it can, with the tests green the whole way.</p>



<p class="wp-block-paragraph">This is an important distinction because models do well at some tasks. But for anything that isn’t an immediate change to a small part of a codebase, especially in a complex brownfield system, model-only automated coding faces an insurmountable obstacle. Weekend toys and side projects are alike in that a few months of development cycles is usually enough to get things in working order, or at least close enough. But an enterprise system that has been under development for a decade or more is a different beast; it has to be maintained, in a professional environment at a professional pace. Three to six months into a project, you’re already drowning in unread code. That kind of environment, and especially the constraints enforced by production code, would make even a powerful agent do poorly, all of it in contrast to the vibe-coding enjoyed by developers working on weekend toys.</p>



<p class="wp-block-paragraph">Dex reports from experience that this is a major failure, so much so that it required painstaking manual debugging to pinpoint. This came from running a fully automated code factory for about four months, during which no human looked at the code that was written. Underlying the experience is a tradeoff between two conflicting metrics. One is maximizing token utilization, the number we currently treat as progress. The other, which it quietly minimizes, is the amount of the system any human participant still understands at any moment.</p>



<p class="wp-block-paragraph">Where the dark factory truly shines is in its ability to burn through pristine code while the tests stay green. The ultimate reckoning, when it comes, will not be a dramatic “it all goes sideways” moment. It will be quiet and late.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059056d28&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059056d28" class="aligncenter is-resized wp-lightbox-container"><img decoding="async" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://addyosmani.com/assets/images/software-factories/dark-vs-lit.svg" alt="A dark factory pipeline versus a lit factory pipeline" style="aspect-ratio:2.6048026048026047;width:840px;height:auto"/><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>A dark factory pipeline versus a lit factory pipeline: Dark and lit are the same pipeline with the lights in different places. The lit version moves human judgment upstream to design and architecture rather than just re-adding review at the end.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading"><strong>The bottleneck was never generation</strong></h2>



<p class="wp-block-paragraph">The fundamental constraint in a software factory isn’t how much code we can churn out, it’s how quickly we can verify it.</p>



<p class="wp-block-paragraph"><strong>Back pressure</strong> is the rule that you can only hand a loop as much autonomy as you can cheaply and reliably verify, and not one inch more. Verification, not generation, is the real constraint on a factory.</p>



<p class="wp-block-paragraph">Because unbounded generation capacity is in perpetual tension with the finite, non-scaling resource of human attention, the core problem is the gap between cheap generation and bounded review. Look at the funnel: As long as the neck representing verification doesn’t widen, it’s going to back up. As Dex points out, volume alone isn’t the problem: What we’re really suffering from is a surplus of bad PRs. When you’ve got high volume without trustworthy gates, manufactured defects are unavoidable. This is just back pressure again: Autonomy can’t expand beyond what can be cheaply and reliably verified.</p>



<p class="wp-block-paragraph">The second-order problem is why improving the model shouldn’t automatically close the gap between what it can generate and what can be verified. Training on well-architected systems is an arguably more difficult proposition than passing simple tests: remember, the cost functions measuring architectural excellence aren’t measured in seconds or even minutes, but in months and years. Tidy gradients are functionally impossible to compute, so a system expecting crisp, instant evaluation of complex design decisions isn’t going to be trained on good examples.</p>


<div class="wp-block-image">
<figure class="aligncenter is-resized"><img decoding="async" src="https://addyosmani.com/assets/images/software-factories/the-funnel.svg" alt="Unbounded generation meets a narrow verification gate" style="aspect-ratio:2.8000583345486363;width:840px;height:auto"/><figcaption class="wp-element-caption"><em>Generation is a wide mouth; verification is the narrow neck. Speeding up the mouth just deepens the pile at the neck.</em></figcaption></figure>
</div>


<h2 class="wp-block-heading">Turning the lights back on</h2>



<p class="wp-block-paragraph">A lit factory is the same pipeline with the lights left on where judgment lives. Agents still do most of the building, but a human reads what comes out before it ships, keeping the lights on wherever a wrong call is expensive.</p>



<p class="wp-block-paragraph">The lit version doesn’t tack review onto the end but moves the point of human judgment upstream, to the product, the design, and the architecture before an agent starts a loop.</p>



<p class="wp-block-paragraph">One great thing about that upfront hour is that it leads to fewer implementation hours. It turns a long, frustrating code review into a quick read of a two-hundred-line plan. You get to review a decision before it’s built, so later you aren’t chasing through two thousand lines of generated code to find out what the decision even was. Some decisions are expensive and long-lived enough that you’d want a person in on them early, before the cost compounds. Of course, there are still times you look at diffs, even when you’ve spent time up front.</p>



<p class="wp-block-paragraph">You might be thinking that all sounds unglamorous. You’re right. The safety net is made up of perfectly ordinary architectural practices we’ve always known about and mostly ignored: good types and method signatures so that mistakes are caught by the compiler instead of in production; test seams where we can pin behavior and make change observable; laying out the code so the next reader, human or model, knows where to find the thing they care about; keeping call stacks short and legible; keeping component boundaries well defined so a change doesn’t have a huge blast radius; and dependency injection so we can swap out one piece for another. None of it is new. We’ve always said we care about good architecture. But now that we’re using automated coding agents, that architecture is finally doing a second job as a cheap and hard-to-fake safety net against the mistakes the agent will make.</p>



<p class="wp-block-paragraph">That safety net has to live outside the model because the model won’t supply it. The coding agents that feel most capable, Claude Code and Codex among them, are reinforcement-trained against their own harness and tools: fluent with all the tools and idioms of the trade, but not with things like long-term maintainability. The deliberate architecture we’ve always talked about is the tool that catches that debt, and the investment we make in it is us buying back our autonomy. Put that together with safe infrastructure, and there are some tight, low-risk loops you can run unattended. Horthy described one in a recent post: A nightly GitHub Actions cron that fixes exactly one anti-pattern, a lint violation or a needlessly optional prop, commits, and opens one small pull request, all on its own, so the team wakes up to a slightly better codebase and a diff short enough to read. But for loops with high enough stakes, you don’t want to risk waking up to a broken auth system, billing engine, or public API contract. Keep the lights on there, and trust that a person with judgment and a real working knowledge of the system will catch the mistake.</p>



<h2 class="wp-block-heading">What earns a loop the dark</h2>



<p class="wp-block-paragraph">This rule applies whether you call it back pressure, verification, or the light switch.</p>



<p class="wp-block-paragraph">A loop can earn itself fully automated status only if the check is cheap, runs at high frequency, and relies on something that can’t be easily faked out. Green-or-red oracles, type gates, property tests, and a review agent coupled with a real rubric all fit. You also need the oracle to answer immediately and not drift over time. When done can be proven not just by you but by a machine, you’ve reached automation.</p>



<p class="wp-block-paragraph">Short loops are easier to verify than long ones. <a href="https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-10-small-focused-agents.md" target="_blank" rel="noopener">Dex’s rule of thumb</a>: An agent holds up for three to ten steps, then starts losing the thread past twenty. The reason is context accumulation. The more the agent drags along, the more likely it is to wander off. When a loop is short, verifying it is cheap. Sprawling loops hide mistakes in the corners, which is another way of saying they never earned lights-out status.</p>



<p class="wp-block-paragraph">Keeping the lights on is the opposite case. A loop needs to be reviewed if a wrong answer is expensive and only a person can catch it. Subtle production bugs that can’t be caught by tests, large blast radii, and a decision that’s going to shape the work of a year or more all qualify. In those cases, <a href="https://addyosmani.com/blog/human-judgment-doesnt-leave-the-software/" target="_blank" rel="noopener">human judgment does not leave the software</a>; your attention is the costly, essential part.</p>



<p class="wp-block-paragraph">The danger is forgetting to flip each switch and just setting all of them to the same mode. All dark, and you’re stuck tearing everything down four months later. All lit, and no one can get reviews done in time and you’re stuck in a gigantic bottleneck. The hard, skilled job is deciding where to put each switch.</p>



<h2 class="wp-block-heading"><strong>Loops, graphs, or state machines?</strong></h2>



<p class="wp-block-paragraph">When you hand an agent a task, you’ll likely build a graph around it, whether you call that graph a finite state machine or a set of conditionally linked service calls. It’s a framing where the software isn’t just following some abstract rules but a structured workflow: Every node is an explicit step, and every edge between nodes is an explicit condition. That sounds like a lot of structure, but most of it’s already there in any software, since any code can be expressed as a control-flow graph. So the only real novelty is that an agent insisting on autonomy is really just walking around a particular graph, and its freedom is constrained to the inside of a node. And here’s the part people forget, which <a href="https://github.com/humanlayer/12-factor-agents" target="_blank" rel="noopener">Dex wrote down</a> a year ago: software was always going to have that structure. There’s a reason we used to draw programs as flow charts. The genuinely new move was trying to throw the diagram away, leaning on a loop where the model picks the path tool call by tool call, until it declares itself done. That felt like liberation, right up until it met a ten-year-old codebase, and the discipline everyone is now rediscovering, owning your control flow, is really just walking the graph back around the loop. So the question of whether we should shift from loops back to graphs is almost an admission that we needed the flowchart all along.</p>



<p class="wp-block-paragraph">Here’s what it looks like in practice. Take a bug to fix. As a pure loop, you sit down and think: figure out what’s wrong, change some code, run the tests, see what happens, and if that round doesn’t kill the run, loop back and start again. The whole journey is decided as you go, which problem you chase, the exact code you change, which tests you run and in what order, whether you run tests at all, and whether you try again or declare victory. As a graph, the first thing you do is map out what should happen. Reproduce the bug or go ask for more information, find the cause, try a fix, run the tests, and let a failing run route back to the fix while a passing one goes on to review, where only an approval reaches done. The agent is still clever inside each box; it just can’t wander off the paths you sanctioned. Santi laid this out with a diagram that makes the difference obvious.</p>



<p class="wp-block-paragraph">The real appeal of that graph, of course, is that it’s back pressure drawn as a diagram. You give up some of the agent’s freedom and get mandatory checks and legible failure points in return, so when a run dies you can point at the node that killed it. It’s the same instinct behind Dex’s blunt line that most so-called agents aren’t very agentic at all, “mostly deterministic code, with LLM steps sprinkled in at just the right points.” And this isn’t just an artifact of how people happen to be building things right now: you can see the pattern in LangGraph and LlamaIndex Workflows, in Jerry Liu’s hybrid workflow-graph-over-agents with an outer loop that grows parts of the graph as it runs, and in David Khourshid’s reminder that this is really just state machines and the actor model turning up in new clothes.</p>



<p class="wp-block-paragraph">One clarification, because the term is badly overloaded: when I keep calling this a graph, I don’t mean a knowledge graph. I mean a predefined directed graph of how the work should flow, conditional edges and all, giving the loop a shape you can actually trust.</p>



<h2 class="wp-block-heading">Where the human actually goes</h2>



<p class="wp-block-paragraph">Notice that the person never left the factory. They moved.</p>



<p class="wp-block-paragraph">I think engineers need to increasingly own the <a href="https://addyosmani.com/blog/own-the-outer-loop/" target="_blank" rel="noopener">outer loop</a>. The agents can investigate a bug, write up the diagnosis, implement a fix, run the tests, and write up a report. That’s the execution of the inner loop, and they can do it as efficiently as anyone. But that was never the job. The bits you own are what I’d call the outer loop: Decide whether it’s the right way to address the problem, verify that the diagnosis and implementation are sound, approve the change, and carry the consequences of being wrong. The boundary between the two loops is evidence, the diffs, the tests, the logs, and a brief explanation that connects them. Types, seams, and rubrics make it possible to oversee all this without doing a lot of work for every change.</p>



<p class="wp-block-paragraph">I think it’s useful to put it this way: you’re not down on the line writing changes any more; you’re up at the end of the production line designing it and guarding the gate. There’s a lot you can do to make the model better and the harness more capable, but I’ve observed that identifying problems that are expensive in the long term is not typically something you can automate away. The core thing that’s still the job is to exercise human judgment better than any flow of paper and computing power.</p>



<p class="wp-block-paragraph">Robots are fine operating in the dark, but humans need to see what they’re doing. If everything on the factory floor is dark, and you can’t see anything, and you can’t even find the light switch, that’s where the danger is.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/software-factories-light-and-dark/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: Capability, Capital, and Consequences</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-capability-capital-and-consequences/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-capability-capital-and-consequences/#respond</comments>
				<pubDate>Fri, 18 Sep 2026 13:29:19 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19729</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[New model advancements, record AI investment, growing safety concerns, and new applications in genomics]]></custom:subtitle>
		
				<description><![CDATA[OpenAI expanded into software control, scientific reasoning, and financial services this week, and investors committed billions more to AI companies across the stack, while AI researchers went public with warnings. This Week in AI host Christina Stathopoulos looked at what those developments mean for an industry already wrestling with questions about safety and control. Models [&#8230;]]]></description>
								<content:encoded><![CDATA[
<figure class="wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="This Week in AI: Capability, Capital, and Consequences" width="500" height="281" src="https://www.youtube.com/embed/jE4PSeh2JNU?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">OpenAI expanded into software control, scientific reasoning, and financial services this week, and investors committed billions more to AI companies across the stack, while AI researchers went public with warnings. <em>This Week in AI</em> host Christina Stathopoulos looked at what those developments mean for an industry already wrestling with questions about safety and control.</p>



<h2 class="wp-block-heading"><strong>Models are moving from answering questions to doing work</strong></h2>



<p class="wp-block-paragraph">OpenAI’s latest announcements showed how much more work companies now expect models to handle. <a href="https://openai.com/index/gpt-6-astra/" target="_blank" rel="noopener">GPT‑6 Astra</a> can navigate software interfaces, complete multistep workflows, and apply advanced reasoning to scientific and mathematical problems, and <a href="https://openai.com/index/introducing-chatgpt-financial-services/" target="_blank" rel="noopener">ChatGPT for financial services</a> was developed with input from Morgan Stanley and Evercore to support research, financial modeling, and creating client materials.</p>



<p class="wp-block-paragraph">OpenAI also shared a solution to the previously unsolved <a href="https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_existence_and_smoothness" target="_blank" rel="noopener">Navier–Stokes Millennium Prize Problem</a>. A <a href="https://openai.com/index/navier-stokes-solution/" target="_blank" rel="noopener">coordinated system of 10,000 AI agents</a> worked for 88 hours on the proof, followed by another 17 hours of model-based verification by Astra. However, the company’s claim drew scrutiny after <a href="https://www.scientificamerican.com/article/openai-claims-blockbuster-math-breakthrough-amid-swirl-of-controversy/" target="_blank" rel="noopener">outside researchers questioned</a> whether OpenAI might have had access to related work, an allegation OpenAI denies. While impressive, scientific breakthroughs like this also raise important questions about how well we <em>understand</em> these systems and the role humans should continue to play in scientific discovery. (Hugo Bowne-Anderson got into this in a <a href="https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/" target="_blank" rel="noopener">recent article on Radar</a>.)</p>



<h2 class="wp-block-heading"><strong>Investors are placing bets across the AI stack</strong></h2>



<p class="wp-block-paragraph">Money continues to flow to AI companies, but investors are backing infrastructure, platforms, and specialized applications rather than converging on a single layer of the stack. French company Mistral <a href="https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/" target="_blank" rel="noopener">has raised €3 billion</a> with plans to spend on compute infrastructure and open weight models. Legal AI company Harvey, inference chip startup Positron, and enterprise AI company Wonderful also raised large rounds, while NVIDIA announced its <a href="https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/" target="_blank" rel="noopener">acquisition of Hugging Face</a> for nearly $13 billion.</p>



<p class="wp-block-paragraph">Christina cited figures showing global AI funding rising from $56 billion in the fourth quarter of 2025 to $242 billion in the first quarter of 2026 but questioned whether generative AI will produce returns that justify that level of investment. Some of these companies may build durable businesses and others may not, even if AI itself continues to deliver useful products and services.</p>



<h2 class="wp-block-heading"><strong>Safety issues are colliding with high-value applications</strong></h2>



<p class="wp-block-paragraph">AI safety is back in the news, following <a href="https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628" target="_blank" rel="noopener">Anthropic researcher Jacob Coxon&#8217;s</a> highly publicized resignation. Coxon warned that labs were moving too quickly toward poorly understood systems capable of recursive self-improvement, and other researchers associated with Anthropic and Google DeepMind raised similar concerns. Anthropic CEO Dario Amodei also called for <a href="https://darioamodei.com/post/we-must-pace-the-frontier" target="_blank" rel="noopener">stronger evaluation, shared safety standards, and international coordination</a>. (Sam Altman and Elon Musk <a href="https://www.theguardian.com/technology/2026/sep/13/openai-sam-altman-elon-musk-back-anthropic-calls-brakes-ai-development" target="_blank" rel="noopener">seconded the call</a>.)</p>



<p class="wp-block-paragraph">While the industry remains divided over catastrophic-risk scenarios, many nearer-term problems are already concrete, and Christina was more concerned about people using powerful AI systems maliciously than about autonomous systems becoming dangerous on their own. Organizations deploying more autonomous systems must tread carefully, with robust security, access controls, testing, and human oversight in place.</p>



<h2 class="wp-block-heading"><strong>AI for good is getting more concrete in genomics</strong></h2>



<p class="wp-block-paragraph">After a week of AI safety warnings, Christina ended on a positive note with what she calls “AI for good” and highlighted genomics projects from <a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/" target="_blank" rel="noopener">DeepMind</a>, <a href="https://www.nature.com/articles/s41586-026-11005-5" target="_blank" rel="noopener">UC Berkeley</a>, and <a href="https://www.tempus.com/news/pr/tempus-launches-effort-to-build-the-largest-multimodal-whole-genome-dataset-to-advance-ai-driven-healthcare-innovation/" target="_blank" rel="noopener">Tempus</a>. Their work uses AI to predict how genetic changes affect gene function; identify mutations associated with disease; and connect genomic data with patients’ medical histories.</p>



<p class="wp-block-paragraph">For researchers, AI can make it practical to study genetic possibilities that would be difficult to test individually in a lab. That could help narrow the search for disease-related variants and support earlier diagnosis and more personalized treatment.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-capability-capital-and-consequences/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Navigating the Modern Data Lexicon: A Working Vocabulary for the Semantic Era</title>
		<link>https://www.oreilly.com/radar/navigating-the-modern-data-lexicon-a-working-vocabulary-for-the-semantic-era/</link>
				<comments>https://www.oreilly.com/radar/navigating-the-modern-data-lexicon-a-working-vocabulary-for-the-semantic-era/#respond</comments>
				<pubDate>Thu, 17 Sep 2026 10:54:35 +0000</pubDate>
					<dc:creator><![CDATA[Jeremy Arendt]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19724</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Navigating-the-modern-data-lexicon.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Navigating-the-modern-data-lexicon-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The way we talk about data is changing faster than the way we build it. Every quarter a vendor ships a new approach, coins a new term for it, or quietly adopts a term someone else has been using and redefines it to fit the shape of their product. None of this is malicious. Every [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The way we talk about data is changing faster than the way we build it. Every quarter a vendor ships a new approach, coins a new term for it, or quietly adopts a term someone else has been using and redefines it to fit the shape of their product. None of this is malicious. Every company describes the landscape from wherever they happen to be standing. But when six vendors do that to the same word, practitioners are left translating between six versions of it before a design conversation can even start.</p>



<p class="wp-block-paragraph">There’s a second problem stacked on top of the first. Most of the vocabulary we use to talk about data in the AI era comes from academic disciplines that very few working practitioners have spent time in. “Data warehouse” is immediately legible: You know what a warehouse is, so you know this is a place where things are stored until someone needs them. “Ontology” is not. It arrives from philosophy by way of knowledge engineering, where Tom Gruber defined it in 1993 as an explicit specification of a conceptualization. That’s a precise definition. It’s also useless to a director trying to decide what to fund next quarter.</p>



<p class="wp-block-paragraph">What follows is an attempt at a working vocabulary, written for the people who actually deploy these technologies and the people who approve their budgets. For each term I want to answer three questions. What is it, actually: software, an artifact, or a practice? What job does it do? And which kind of output does it serve? That last question needs some setup, so let’s start there.</p>



<h2 class="wp-block-heading"><strong>Deterministic and probabilistic outputs</strong></h2>



<p class="wp-block-paragraph">Data systems produce two kinds of output, and knowing which one you’re after is the single most useful diagnostic in modern architecture.</p>



<p class="wp-block-paragraph">A <strong>deterministic</strong> output is the same every time you ask the same question. What was ARR for the last twelve months? Whether that question goes to a dashboard, an API call, an Excel workbook, or an AI agent, the answer should be identical. Ask four different agents running on four different models and you should still get one number. Deterministic outputs have traceable lineage. You can point at the calculation and walk someone through how the number was produced.</p>



<p class="wp-block-paragraph">A <strong>probabilistic</strong> output is what you get from systems that are non-deterministic by design. Change the ARR question slightly and the category changes completely: Instead of “what was ARR over the past twelve months,” ask “how can we improve ARR over the next twelve months.” Put that question to the same model, in the same agent, twice in a row, and you’ll get two different answers. That’s not a bug. An LLM is predicting a likely sequence of tokens across billions of parameters, and the output varies every time it runs.</p>



<p class="wp-block-paragraph">Neither type is better. Both are necessary. The failure mode is asking a probabilistic system for a deterministic answer and not realizing that’s what you did. Most of the terms below exist because the industry is trying to solve exactly that problem: How do you put enough structure around a probabilistic system that it can return deterministic answers when the question calls for one?</p>



<p class="wp-block-paragraph">With that, let&#8217;s work through the terms.</p>



<h2 class="wp-block-heading"><strong>Semantic layer</strong></h2>



<p class="wp-block-paragraph">I’ve written about semantic layers for Radar several times, including <a href="https://www.oreilly.com/radar/the-trillion-dollar-problem/" target="_blank" rel="noopener">what they are and why they matter</a> and <a href="https://www.oreilly.com/radar/the-best-risk-mitigation-strategy-in-data-a-single-source-of-truth/" target="_blank" rel="noopener">why they function as a risk mitigation strategy</a>. The short version: A semantic layer is software that sits between your data and the people and tools that consume it, giving everyone a single place to access trusted, governed metrics.</p>



<p class="wp-block-paragraph">Behind the scenes, it does three things. It holds <strong>definitions</strong>: How do we calculate this business metric? It holds <strong>context</strong>: What does this model or column contain, and what’s it typically used for? And it holds <strong>relationships</strong>: How does this data fit together? Modern tools bundle in more than that, including query engines, caching, and a single point for access control and security, but definitions, context, and relationships are the core.</p>



<p class="wp-block-paragraph">Why does this matter for AI? Because it lets an agent <em>navigate</em> data instead of <em>reasoning over</em> it. Without a semantic layer, an agent that’s asked for last year’s ARR has to inspect table names, guess at joins, infer which date field represents revenue recognition, and reconstruct business logic that lives in someone’s head. That’s reasoning, probabilistic, and produces a different answer depending on the day. With a semantic layer, the agent looks up ARR, queries the definition, and returns the same number every time. It’s a deterministic answer delivered through a probabilistic tool.</p>



<p class="wp-block-paragraph">The analyst community has caught up to this. Gartner now predicts that <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-11-gartner-announces-top-predictions-for-data-and-analytics-in-2026" target="_blank" rel="noopener">universal semantic layers will be treated as critical infrastructure by 2030</a>, alongside data platforms and cybersecurity.</p>



<h2 class="wp-block-heading"><strong>Ontology</strong></h2>



<p class="wp-block-paragraph">Ontology is the term most likely to derail a meeting right now, largely because Palantir made it commercially famous while the underlying concept came out of decades of academic work on how to formally describe things and the relationships between them.</p>



<p class="wp-block-paragraph">Here’s the simplest way I’ve found to separate it from a semantic layer. A semantic layer answers <em>what does this number mean and how is it calculated?</em> An ontology answers <em>what things exist in this business and how do they relate to each other?</em> The semantic layer is metric-first: measures, dimensions, and the logic that connects them. The ontology is entity-first: customer, order, shipment, facility, supplier, along with the relationships and rules that govern how those objects behave.</p>



<p class="wp-block-paragraph">The overlap is real, and it lives in relationships. Both artifacts encode how things connect, and vendors are increasingly shipping both capabilities under a single product name, which is a large part of why the terms have blurred. The practical distinction is what the system needs to do. If the job requires consistent numbers across every reporting tool, a semantic layer is the center of gravity. If the job requires an agent that reasons about business objects and takes action on them, rather than just reporting on them, an ontology is what gives it a model of the world to act in.</p>



<p class="wp-block-paragraph">One useful clarification: An ontology isn’t software. It’s a model, an artifact your organization authors and maintains. Software delivers it, but the value is in the modeling work.</p>



<h2 class="wp-block-heading"><strong>Knowledge graph</strong></h2>



<p class="wp-block-paragraph">If the ontology is the schema, the knowledge graph is that schema populated with actual data. The ontology says a customer places an order, and an order contains line items. The knowledge graph holds your real customers, your real orders, and the edges connecting them, stored as nodes and relationships rather than rows and columns.</p>



<p class="wp-block-paragraph">How do you know when to use a knowledge graph over a semantic layer? Warehouses and semantic layers are excellent at aggregation: how much, how many, compared to when. Graphs are excellent at connection: what is linked to what, and how far apart. “Which suppliers are two steps removed from this delayed shipment?” is a graph question. So is “which accounts share a beneficial owner,” and “who has inherited access to this dataset through three layers of group membership?” You can answer those with SQL. You won’t enjoy it.</p>



<p class="wp-block-paragraph">Graph traversal is deterministic. Given the same graph and the same query, you get the same path every time, which is exactly what makes graphs useful as grounding for an agent. Rather than inferring that two records refer to the same supplier, the agent follows an edge that someone already asserted. The relationships are modeled facts, not inferences made at inference time.</p>



<p class="wp-block-paragraph">A knowledge graph is not a substitute for a semantic layer. They answer different questions, and mature architectures increasingly run both.</p>



<h2 class="wp-block-heading"><strong>Context</strong></h2>



<p class="wp-block-paragraph">Context is the most overloaded word in the field right now, and it’s worth splitting into pieces before using it in a sentence.</p>



<p class="wp-block-paragraph"><strong>Deterministic context</strong> is metadata, plainly. It lives in your semantic layer or your ontology: field descriptions, metric definitions, object relationships, business rules, exclusion logic. What has changed isn’t the concept but the consumer. Metadata used to be documentation for humans, and it was the first thing to go stale because nothing broke when it did. Now an agent reads it at query time to decide what a column means and whether it’s allowed to use it, which makes it functional infrastructure rather than a wiki page nobody updates. It’s versioned, reviewed, and reads the same way every time a system asks for it. This is an asset you maintain.</p>



<p class="wp-block-paragraph"><strong>Runtime context</strong> is what an agent assembles at the moment of inference: the system prompt, conversation history, retrieved documents, tool outputs, whatever the orchestration layer decided to put in the window. It’s ephemeral, and directly changes the answer. Same question, different context window, different output. This is a variable you monitor.</p>



<p class="wp-block-paragraph">Cutting the other direction, <strong>structured context</strong> describes governed data: columns, metrics, entities, relationships. <strong>Unstructured context</strong> is the policy PDFs, contracts, support tickets, and wiki pages that hold the reasoning behind the numbers. Unstructured context is genuinely valuable and usually retrieved through similarity search, which means it arrives with probabilistic behavior attached. What surfaces depends on how the question was phrased.</p>



<p class="wp-block-paragraph">The practical rule: When someone tells you their tool is “context aware,” ask which kind. Deterministic context is what makes an agent’s answer repeatable. Runtime context is what makes it relevant. Conflating them is how teams end up trusting an answer that was only true for one prompt.</p>



<h2 class="wp-block-heading"><strong>Observability</strong></h2>



<p class="wp-block-paragraph">Observability is the telemetry that tells you whether your systems are still doing what you believe they’re doing. It isn’t data quality, which is a judgment about whether a number is correct, and it’s not testing, which is a check you wrote in advance for a failure you already anticipated. Observability is the instrumentation that lets you ask “is this still working?” without having predicted the specific way it would break.</p>



<p class="wp-block-paragraph">On the deterministic side, this is familiar territory: freshness, row counts, schema changes, null rates, job failures, and lineage impact. If ARR is supposed to refresh at 6 a.m. and today it didn’t, you want to know before the CFO does.</p>



<p class="wp-block-paragraph">The probabilistic side is harder because there is often no error to catch. The system returns a fluent, plausible answer that happens to be wrong. Monitoring here means evaluation sets scored over time, tool call success rates, retrieval relevance, refusal and fallback rates, latency, cost per query, and structured human feedback.</p>



<p class="wp-block-paragraph">Which brings us to <strong>drift</strong>. Drift is what happens when the world changes underneath a system that keeps running unchanged. Data drift is a shift in the inputs: a new business unit lands in the source system, order volume triples after an acquisition, a vendor starts sending nulls in a field that was never null before. Model drift is a shift in behavior: The provider ships a new model version, or a prompt template changes, and outputs that were stable last month aren’t stable this month.</p>



<p class="wp-block-paragraph">Here’s what drift looks like in practice. In March, an agent answered “what were our top five products by margin?” correctly. In June, a new product hierarchy shipped upstream, and the agent now silently excludes an entire category. Nothing failed. No alert fired. The answer is simply wrong, and it’ll stay wrong until someone notices. Deterministic systems tend to fail loudly. Probabilistic systems fail quietly. Observability is how you catch the quiet ones.</p>



<h2 class="wp-block-heading"><strong>The working vocabulary</strong></h2>



<ul class="wp-block-list">
<li><strong>Deterministic output:</strong> The same answer to the same question every time, with a calculation you can trace.</li>



<li><strong>Probabilistic output:</strong> A different answer to the same question each time, produced by prediction rather than calculation.</li>



<li><strong>Semantic layer:</strong> Software that stores the definitions, context, and relationships behind your business metrics and serves them consistently to every downstream tool.</li>



<li><strong>Ontology:</strong> A model of what your business is made of, the objects, their relationships, and the rules that govern them.</li>



<li><strong>Knowledge graph:</strong> An ontology populated with real data and stored as nodes and edges, so systems can traverse relationships instead of reconstructing them through joins.</li>



<li><strong>Context:</strong> The information a system needs to use data correctly, either governed in a semantic model or assembled at runtime by an agent.</li>



<li><strong>Observability:</strong> The telemetry that tells you whether your data and AI systems are still doing what you think they’re doing.</li>
</ul>



<p class="wp-block-paragraph">Read that list in order and something becomes obvious: These aren’t competing products. They’re layers. The ontology describes what exists. The knowledge graph holds the instances. The semantic layer defines the measures. Context is how any of it reaches a model. Observability is how you find out when it stops working. The reason why these terms feel like they’re fighting each other is  because they’re usually sold as substitutes, when in practice, they stack.</p>



<p class="wp-block-paragraph">The vocabulary will keep moving. Two years from now some of these words will be absorbed into product names and mean something slightly different than they do today. That’s fine, as long as your team has a shared answer to two questions about any term someone puts in front of you. What is it, actually: software, an artifact, or a practice? And which kind of output does it serve, deterministic or probabilistic?</p>



<p class="wp-block-paragraph">Those two questions cut through most of the noise. Agree on the words first. The architecture arguments get much shorter after that.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/navigating-the-modern-data-lexicon-a-working-vocabulary-for-the-semantic-era/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Architecting for the Knowledge You Can’t Capture</title>
		<link>https://www.oreilly.com/radar/architecting-for-the-knowledge-you-cant-capture/</link>
				<comments>https://www.oreilly.com/radar/architecting-for-the-knowledge-you-cant-capture/#respond</comments>
				<pubDate>Wed, 16 Sep 2026 16:00:18 +0000</pubDate>
					<dc:creator><![CDATA[Jofia Jose Prakash]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Operations]]></category>
		<category><![CDATA[Software Architecture]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19712</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Architecting-for-the-knowledge-you-cant-capture.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Architecting-for-the-knowledge-you-cant-capture-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Tacit knowledge is the hardest requirement in enterprise AI. Here&#039;s how a data and knowledge architect designs, and evaluates, for it.]]></custom:subtitle>
		
				<description><![CDATA[Every knowledge program seems to begin with the same request. A senior engineer is leaving in six weeks, and someone asks her to document the process she’s carried for years. She returns a clean flowchart of the happy path. The drawing is accurate and may even be elegant. It leaves out the thresholds she watches, [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Every knowledge program seems to begin with the same request. A senior engineer is leaving in six weeks, and someone asks her to document the process she’s carried for years.</p>



<p class="wp-block-paragraph">She returns a clean flowchart of the happy path. The drawing is accurate and may even be elegant. It leaves out the thresholds she watches, the conditions that make the standard procedure unsafe, and the supplier whose parts fail in humid weather. She doesn’t think of those judgments as separate knowledge. After years on the job, they feel obvious.</p>



<p class="wp-block-paragraph">Six months later, a production line goes down and the knowledge base can’t explain what to do. The interview took place as per the process. Its transcript was chunked, embedded, and indexed, so the search returns the relevant passage quickly. The passage still can’t answer the question because no one asked the engineer to explain the judgment behind the procedure.</p>



<p class="wp-block-paragraph">That gap now limits many enterprise AI programs. Organizations continue to improve retrieval over collections that omit some of their most valuable operating knowledge. Better ranking can help people find what was recorded; it can’t recover the expertise that never entered the collection.</p>



<h2 class="wp-block-heading"><strong>The blind spot in enterprise knowledge systems</strong></h2>



<p class="wp-block-paragraph">Michael Polanyi gave the problem its durable formulation in 1966: “We can know more than we can tell.” In <em><a href="https://press.uchicago.edu/ucp/books/book/chicago/T/bo6035368.html" target="_blank" rel="noopener">The Tacit Dimension</a></em>, he argued that competence depends on skill, perception, and judgment that resist full explanation, even when an expert sincerely tries to teach them.</p>



<p class="wp-block-paragraph">In companies, tacit knowledge usually appears in three forms. Elicitable knowledge remains unspoken because nobody has asked a precise enough question, or because an expert assumes that everyone sees what she sees. Perceptual knowledge lives in trained attention: An engineer hears a bearing begin to fail, or a nurse notices that a patient looks wrong before a monitor changes. Collective knowledge resides in a team’s habits, standards, and shared sense of what a sound decision looks like in that organization. Each form requires a different method of transfer.</p>



<p class="wp-block-paragraph">Preventive judgment creates another difficulty for the architect. A failure produces a ticket, an incident report, and a trail of messages. An experienced operator who quietly avoids a known failure mode on a Friday afternoon produces none of those records. The useful outcome is the absence of an event, so the data pipeline receives no trace of the decision that produced it.</p>



<p class="wp-block-paragraph">Machine learning can infer rules that people struggle to articulate, provided the model sees enough representative examples. It’s difficult to find enough examples of rare expertise for training. A company may have only a handful of unusual incidents and one person who has learned, over decades, how to read them.</p>



<p class="wp-block-paragraph">David Autor described this limit as “<a href="https://www.nber.org/papers/w20485" target="_blank" rel="noopener">Polanyi’s paradox</a>”: Many of the tasks that are hardest to automate depend on rules we can’t state. Modern machine learning works around the paradox by learning from examples, but the workaround weakens when examples are scarce. Fine-tuning can teach a model the company’s vocabulary and document formats. It can’t reconstruct decisions that left no data.</p>



<p class="wp-block-paragraph">At the same time, the economics have changed. Much of a field’s documented best practice now appears in frontier-model training data and is available to competitors at roughly the same price and quality. The more widely explicit knowledge circulates, the more a company’s advantage depends on local judgment: the exceptions, thresholds, relationships, and practiced responses that its people have accumulated.</p>



<p class="wp-block-paragraph">That makes elicitation an architectural concern rather than an offboarding chore. The organization needs a repeatable way to surface the knowledge that can be expressed, a route for the expertise that must be demonstrated, and enough humility to distinguish the two.</p>



<h2 class="wp-block-heading"><strong>A protocol for elicitation</strong></h2>



<p class="wp-block-paragraph">The central design question is straightforward: Which follow-up would prompt an expert to say the missing judgment aloud? The quality of the interview sets the ceiling for the knowledge base. The index determines how quickly someone can reach the resulting material.</p>



<p class="wp-block-paragraph">Interviews can be made more reliable even though judgment itself remains highly personal. An expert may know that a particular supplier fails in humid weather. The interviewing protocol doesn’t need to possess that knowledge in advance; it needs to notice a phrase such as “we escalate if it looks bad” and ask the expert to define “bad” in observable terms.</p>



<p class="wp-block-paragraph">Expert explanations tend to become vague in four places. An effective interview protocol asks targeted questions about each one:</p>



<ul class="wp-block-list">
<li><strong>Thresholds</strong>: Which number, reading, or condition triggers the action?</li>



<li><strong>Exceptions</strong>: When does the documented procedure cease to apply?</li>



<li><strong>Evidence</strong>: What did the expert observe before reaching the conclusion?</li>



<li><strong>Escalation</strong>: Who becomes involved, and at what point?</li>
</ul>



<p class="wp-block-paragraph">These questions uncover the operational detail that runbooks often lack. They also identify a narrow, useful role for a language model during the interview: proposing the next question that turns a general statement into a usable rule. I’ve been building an open source toolkit, <a href="https://pypi.org/project/elythera-experttrace/" target="_blank" rel="noopener">ExpertTrace</a>, around that protocol.</p>



<p class="wp-block-paragraph">The value appears in the difference between what an expert volunteers and what the same expert confirms after one focused follow-up. Consider a typical first answer:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We review high-risk use cases before deployment. If the risk seems significant, we escalate to the governance council.</p>
</blockquote>



<p class="wp-block-paragraph">The statement will embed cleanly and retrieve for a relevant query, but a new employee still cannot act on it. “Seems significant” supplies no decision criterion. A targeted follow-up produces something much more useful:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Escalation to the council is required when the use case touches employment, credit, or health decisions, or when model output reaches a customer without human review. Predeployment review is skipped for internal-only tools with no personal data, which is the exception people get wrong most often. If we cannot identify a named accountable owner, the review does not proceed, regardless of risk tier.</p>
</blockquote>



<p class="wp-block-paragraph">The second answer takes little additional time, yet it contains a decision rule, an exception, a recurring failure pattern, and a blocking condition. It can guide a real dispute instead of merely mentioning the subject.</p>



<p class="wp-block-paragraph">The protocol needs guardrails. Limit the number of follow-ups; a long interrogation exhausts the expert and eventually produces agreeable noise. Keep the model focused on generating questions, and separate that task from compiling and validating the answers. An expert’s statement belongs in the record with its provenance and context. Whether the statement is accurate requires independent review.</p>



<h2 class="wp-block-heading"><strong>The four-plane architecture</strong></h2>



<p class="wp-block-paragraph">Elicitation is one part of a larger knowledge system. A tacit-aware architecture has four planes—capture, representation, serving, and transmission and each plane addresses a different failure in the movement of expertise. Figure 1 shows how the four planes work together and which forms of tacit knowledge each can reach.</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab505905fef8&quot;}" data-wp-interactive="core/image" data-wp-key="6ab505905fef8" class="wp-block-image size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1560" height="1000" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-16.png" alt="A tacit-aware knowledge layer: Four planes mapped to the kinds of knowledge each can reach." class="wp-image-19713" style="aspect-ratio:1.5616797900262467" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-16.png 1560w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-16-300x192.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-16-1536x985.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-16-768x492.png 768w" sizes="auto, (max-width: 1560px) 100vw, 1560px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 1. A tacit-aware knowledge layer: Four planes mapped to the kinds of knowledge each can reach.</em></figcaption></figure>



<p class="wp-block-paragraph">In the <strong>capture plane,</strong> structured interviews, incident reconstruction, decision journals, and observation collect more than polished procedure. Record the trigger, evidence, exception, and escalation path while the expert can still explain the surrounding conditions. Route perceptual skill toward demonstration and practice instead of forcing it into prose.</p>



<p class="wp-block-paragraph">Once knowledge has been captured, the <strong>representation plane</strong> preserves the distinctions that make the material trustworthy. A compliance policy, a war story, and an untested hypothesis shouldn’t become interchangeable chunks. Carry provenance, confidence, and validity context—including the plant, time period, equipment, and conditions—as first-class properties. Extend the knowledge graph beyond documents to the people and episodes that produced them.</p>



<p class="wp-block-paragraph">The <strong>serving plane</strong> then determines how that knowledge reaches users. Answers should cite retrieved evidence and show the source. When the collection can’t answer, the system should say so clearly and route the question to someone with relevant experience. “Ask Joe; she rebuilt this line in 2023” is more useful than a fluent paragraph assembled from weak evidence, and the referral restores the human contact through which difficult knowledge often moves.</p>



<p class="wp-block-paragraph">The <strong>transmission plane</strong> completes the architecture by helping how expertise moves between people through shadowing, teaching, and communities of practice. The platform should detect when knowledge concentration and attrition risk converge, then trigger capture and apprenticeship before a notice period begins.</p>



<p class="wp-block-paragraph">Gabriel Szulanski examined <a href="https://onlinelibrary.wiley.com/doi/10.1002/smj.4250171105" target="_blank" rel="noopener">271 observations of 122 best-practice transfers</a> across eight companies and found that even willing teams struggled to reproduce methods developed elsewhere in the same organization. The difficulty often began with causal ambiguity where people could describe the steps without fully understanding why they worked. Receiving teams also needed enough context and experience to absorb and apply what they learned. Preparation, coaching, and time helped them rebuild the practice in their own setting. A repository could preserve the record; the receiving teams still had to turn that record into working knowledge.</p>



<h2 class="wp-block-heading"><strong>Evaluating the knowledge layer</strong></h2>



<p class="wp-block-paragraph">Retrieval precision and answer faithfulness show how well a system serves its existing collection. They don’t reveal whether the collection contains the knowledge on which the organization actually depends. That question needs a separate evaluation loop tied to capture priorities and transfer outcomes. Figure 2 shows how the loop moves from offline evaluation to abstention calibration and then to transfer outcomes.</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059060732&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059060732" class="wp-block-image size-full wp-lightbox-container"><img loading="lazy" decoding="async" width="1560" height="760" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-17.png" alt="The evaluation loop: Offline tests, abstention calibration, and transfer outcomes feeding capture priorities." class="wp-image-19714" style="aspect-ratio:2.0517241379310347" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-17.png 1560w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-17-768x374.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-17-1536x748.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-17-300x146.png 300w" sizes="auto, (max-width: 1560px) 100vw, 1560px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption"><em>Figure 2. The evaluation loop: Offline tests, abstention calibration, and transfer outcomes feeding capture priorities.</em></figcaption></figure>



<p class="wp-block-paragraph">The evaluation begins with <strong>incident replay.</strong> Select 20 or 30 resolved incidents, remove the resolutions, and give the opening facts to the system. Ask the engineers who solved them to grade its responses. Compare those answers with responses from a frontier model that lacks access to the company’s collection. The gap reveals the generic-answer rate: how often the internal system merely restates public knowledge. If reviewers can’t tell the two sets apart, the pipeline adds little institutional value.</p>



<p class="wp-block-paragraph">A <strong>bus-factor audit</strong> tests questions that only one or two employees can answer, and study how the system fails. A clear admission of uncertainty followed by a useful referral is healthy. Fluent boilerplate damages trust in every response, including the accurate ones.</p>



<p class="wp-block-paragraph"><strong>Abstention calibration</strong> measures whether the system answers when evidence exists and declines when corpus can’t support an answer. Build a labeled set of answerable and unanswerable questions, then track abstention precision and recall as the collection grows. A system that never says “I don’t know” is unevaluated on the dimension that matters most.</p>



<p class="wp-block-paragraph"><strong>Transfer outcomes</strong> complete the loop by measuring whether knowledge has reached the people who need it. Evidence of transfer appears in shorter time to proficiency, fewer repeat incidents after elicitation, and fewer critical responsibilities that depend on a single person. Document and query counts describe system activity; they don’t show whether someone else can now make the decision.</p>



<p class="wp-block-paragraph">A strong knowledge system records what an expert said, preserves the conditions around the statement, and marks uncertainty. It also recognizes expertise that requires demonstration, apprenticeship, or team practice. Every evening, the people who carry that knowledge walk out the door. The architecture should be ready long before one gives notice.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>



<p class="wp-block-paragraph"></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/architecting-for-the-knowledge-you-cant-capture/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>What AI Can Teach Us About Being Human</title>
		<link>https://www.oreilly.com/radar/what-ai-can-teach-us-about-being-human/</link>
				<comments>https://www.oreilly.com/radar/what-ai-can-teach-us-about-being-human/#respond</comments>
				<pubDate>Wed, 16 Sep 2026 13:37:00 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19693</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/What-AI-can-teach-us-about-being-human.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/What-AI-can-teach-us-about-being-human-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[My guest on this past week’s Live with Tim O’Reilly was Emmanuel Ameisen, a researcher on Anthropic’s AI interpretability team. I’d heard him give a short talk at Foo Camp on Anthropic’s research into what is going on inside an LLM while it is processing, and I wanted him to reprise the talk and then [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">My guest on this past week’s <em><a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">Live with Tim O’Reilly</a></em> was Emmanuel Ameisen, a researcher on Anthropic’s AI interpretability team. I’d heard him give a short talk at Foo Camp on Anthropic’s research into what is going on inside an LLM while it is processing, and I wanted him to reprise the talk and then go deeper with me and the audience.</p>



<p class="wp-block-paragraph">The essential message of the talk was on the first slide:</p>



<ol class="wp-block-list">
<li>Prediction demands a world model</li>



<li>The world model is readable</li>



<li>The world model is at work in every token</li>
</ol>



<p class="wp-block-paragraph">How do we know this? As tokens pass through a model, particular patterns of activity appear in the intermediate states between its layers. These are called activations. Researchers can study which patterns show up when the model encounters particular ideas, and they can even intervene in those activations and see how the model’s behavior changes. (They do this by capturing the numerical state of the model’s computation in some area where they believe the activation shows a particular “meaning” and then replace the numbers with others.)</p>



<p class="wp-block-paragraph">I went into the conversation thinking about how cool it is (and important too!) to explore what is going on inside the “mind” of a model. But in the end, I found it even more provocative to think about what studying LLMs might teach us about how our own minds work.</p>



<p class="wp-block-paragraph">There’s at least some kind of analogue to what happens in the human brain. Emmanuel began by asking the audience to do a little next-token prediction themselves. He started with an easy one, a hypothetical exchange between two friends:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="has-text-align-left wp-block-paragraph">John: “Is the powder-blue suit too much?”<br>Nick: “Definitely not, man. Send it.”<br>John: “Okay, I’m going to tear it up on the _______________”</p>
</blockquote>



<p class="wp-block-paragraph">Most of us will fill in the blank at the end with “dance floor.” That’s a reminder that humans are also next-token predictors.</p>



<p class="wp-block-paragraph">Then he gave an example that some humans will easily answer, but others without local knowledge might well fail at:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“We also have nature here, just a short bike ride away across the GG bridge. And we have world-class skiing about _______________”</p>
</blockquote>



<p class="wp-block-paragraph">Claude easily completes the thought with “three hours away.” To do that, Claude had to infer that “GG bridge” refers to the Golden Gate Bridge, that the speaker is therefore in San Francisco, and that “world-class skiing” probably refers to Lake Tahoe and then retrieve roughly how long it takes to get there.</p>



<p class="wp-block-paragraph">The point of Emmanuel’s demonstration was that we have become so used to calling LLMs “next-token predictors” in a kind of dismissive way. But as Emmanuel put it, “To predict the next word well, you need a very complex world model.”</p>



<h2 class="wp-block-heading"><strong>How you make a thing is not the same as what the thing becomes</strong></h2>



<p class="wp-block-paragraph">Emmanuel pointed out that people often confuse how you make a thing with how the thing works. Yes, LLMs are trained with the seemingly simple objective of predicting the next token. From that, people may make the leap that what is going on inside must also be simple, something like a very large fuzzy lookup table. “But that’s not true,” Emmanuel said. Simple objectives can give rise to extraordinary complexity. Evolution is the canonical example. No one put “create Beethoven’s Ninth Symphony” or “understand quantum electrodynamics” into the instructions for a process driven by reproduction and selection, yet it eventually produced Beethoven and Feynman. As Emmanuel put it, humans have been “reproducing and killing each other for millions of years, and from that we got jobs—or this podcast.”</p>



<p class="wp-block-paragraph">What Anthropic’s interpretability researchers are finding inside the models looks much less like fuzzy retrieval than many people imagine. They find millions of internal features corresponding to concepts. For example, features for “eyes” show up when the model encounters prose about eyes, an ASCII face, an SVG image, or a photograph. In other words, these features appear to be abstractions rather than merely associations with particular strings of tokens.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059065f4e&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059065f4e" class="aligncenter size-large is-resized wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="895" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12-1600x895.png" alt="Shared concepts across ascii, prose, and code" class="wp-image-19695" style="aspect-ratio:1.7855887521968365;width:1016px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12-1600x895.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12-300x168.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12-767x429.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12-1536x860.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-12.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">Similarly, a feature of the Golden Gate Bridge activates not just for English text about the Golden Gate Bridge but for references in other languages and for images of the bridge. Even more interestingly, researchers can manipulate these features. Turn the activation of the Golden Gate Bridge feature up strongly enough and ask Claude what its physical form is, and instead of saying that it is an AI without a physical body, it announces that its form is the Golden Gate Bridge. It isn’t just that some numbers happen to accompany activations about the Golden Gate Bridge. Changing those numbers changes what the model says it believes.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059066592&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059066592" class="aligncenter size-large is-resized wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="898" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13-1600x898.png" alt="Turning them on causes the model to believe the concept was there" class="wp-image-19696" style="aspect-ratio:1.783450704225352;width:1013px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13-1600x898.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13-300x168.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13-766x430.png 766w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13-1536x862.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-13.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The way a model completes a task that requires thinking ahead also demonstrates a kind of internal world model. Ask Claude to write a rhyming couplet. Even though it emits only one token at a time, before it has written the second line, the activations already reveal the rhyme that it is aiming for. The choice of a word such as “rabbit” for a rhyme happens before the choice of the preceding words on the line, so the model can land there. We call it planning when a person does this. It doesn’t seem unreasonable to use the same word for what is going on here.</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab5059066c73&quot;}" data-wp-interactive="core/image" data-wp-key="6ab5059066c73" class="wp-block-image size-large is-resized wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="897" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15-1600x897.png" alt="Despite predicting one token at a time, models plan many words ahead" class="wp-image-19701" style="aspect-ratio:1.7836812144212524;width:940px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15-1600x897.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15-300x168.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15-767x430.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15-1536x861.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-15.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>



<p class="wp-block-paragraph">Perhaps most challenging to our preconceptions is that there are also features associated with emotions that aren’t activated just by words about those emotions, but by situations, images, characters, and more. These emotion features are even activated by the model’s own activities. For example, <a href="https://transformer-circuits.pub/2026/emotions/index.html#reward-hacking" target="_blank" rel="noopener">“frustration” may be activated when the model is unable to complete a task</a>.</p>



<figure class="wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Despair and Panic Can Precede Destructive Model Actions" width="500" height="281" src="https://www.youtube.com/embed/PHyF2uLAXco?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>The map is not the territory</strong></h2>



<p class="wp-block-paragraph">The issue of anthropomorphization came up during the audience Q&amp;A. One participant objected:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“We should avoid attributing human qualities to LLMs by saying they think, intend, rhyme, or have emotions. Doing so encourages us to project human characteristics onto systems that do not possess them.”</p>
</blockquote>



<p class="wp-block-paragraph">I have sympathy with that warning. Old labels can prevent us from seeing something accurately. But a blanket prohibition against using familiar words can blind us too.</p>



<p class="wp-block-paragraph">If you’ve followed my work for a long time, you know how much I’ve been shaped by <a href="https://www.linkedin.com/pulse/20121029141916-16553-language-is-a-map/" target="_blank" rel="noopener">the ideas of my early mentor George Simon</a>, who in turn was deeply influenced by Alfred Korzybski and general semantics. Korzybski’s famous dictum was “The map is not the territory.” Simon (and Korzybski)&nbsp;taught me that language is a map of experience, which in turn is a set of responses to stimuli from some underlying external reality. The path from reality through experience to conceptual understanding is a very lossy process. The result can be a bad map that can blind us and lead us astray. When we encounter something genuinely new, we have to learn to notice when we are trying to force the territory to fit a map that no longer describes it. But a good map doesn’t just guide us along a route; it helps us notice things that might otherwise be invisible to us.</p>



<p class="wp-block-paragraph">So yes, words like “thinking,” “planning,” “intention,” and “emotion” are labels derived from our experience as human beings. They may turn out to fit LLMs poorly. But if the shoe fits, perhaps we should let them wear it.</p>



<p class="wp-block-paragraph">Emmanuel had a good response to the objection. He said, in effect, that anyone is welcome to propose more precise vocabulary. If it works—that is, if in my framing, it is a good map that helps people see the territory more clearly—people will come to use it. (An audience member later suggested that Emily Bender has done just that. But frankly, I find her <a href="https://buttondown.com/maiht3k/archive/how-to-talk-about-ai-without-adding-to-the/" target="_blank" rel="noopener">suggested alternatives</a> to be quite tortured, obscuring far more than they clarify. Even she admits they don’t work very well, though clinging to the need for them.)</p>



<p class="wp-block-paragraph">In her <a href="https://aiguide.substack.com/p/misleading-metaphors-and-real-risks?utm_source=share&amp;utm_medium=android&amp;r=qxfw" target="_blank" rel="noopener">analysis of the Hugging Face incident, Melanie Mitchell</a> made some observations consistent with the nuanced approach suggested here. She wrote:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Metaphors can help us make sense of novel situations. For example, framing chatbots as “role-playing actors” has been helpful in understanding why these systems exhibit “lying” and “scheming” behavior. But inappropriate metaphors, like the narrative that “OpenAI lost control of escaping swarms of rogue agents,” can lead to ill-informed decisions about how to fix problems or set policy….It is essential for lawmakers, and the public, to understand that none of the reported incidents actually involved loss of control at any time, or arguably even “rogue agents,” or any kind of humanlike agency on the part of AI models. Instead, the blame lies with the humans who failed at engineering safe testing conditions, and who train AI models using RL methods that incentivize high persistence, autonomous decision-making, and reward hacking.</p>
</blockquote>



<p class="wp-block-paragraph">In short, all language is a map. Don’t judge it on that basis alone. Judge it on how well it helps us to see the shape of the territory.</p>



<h2 class="wp-block-heading"><strong>How much of human thought is truly original?</strong></h2>



<p class="wp-block-paragraph">Returning to my conversation with Emmanuel, he remarked that when an existing word really does provide the most precise description, perhaps “what should change isn’t our vocabulary, but our mental model of what these models are.” I replied that it should perhaps also change our mental model of <strong>what we are</strong>. Our encounter with machine intelligence should lead to a better understanding that parts of our own cognition are also mechanistic (albeit derived from a different underlying mechanism than that of LLMs) while other parts are, as yet, somehow perhaps something else.</p>



<p class="wp-block-paragraph">In 1995, O’Reilly published a book that I remain extraordinarily proud of. Stephen Talbott’s <em><a href="https://www.natureinstitute.org/bookstore/the-future-does-not-compute-transcending-the-machines-in-our-midst" target="_blank" rel="noopener">The Future Does Not Compute: Transcending the Machines in Our Midst</a></em> was decades ahead of its time. Its argument was not primarily about what computers would someday become. It was that when we think about machines as intelligent (and yes, we were thinking about that even back in 1995), we are thinking only of the parts of ourselves that are already like our machines. Steve asked us to look at the ways we have built an education system, workplaces, and a society in which we ask humans to act and think like machines. And he asked, “What happens to the rest? How do we make more space for the parts of being human that aren’t like machines?”</p>



<p class="wp-block-paragraph">I’ve been thinking about this for a <em>long</em> time. My 1975 Harvard honors thesis in classics was probably my first crack at this question. I was trying to explain passages in Plato in which early formulations of ideas such as logic and virtue were couched in mystical language that scholars had attributed to “Orphic influence.” My argument, based on my work with George Simon, was that something more fundamental was going on. Plato was trying to describe the numinous experience of thinking genuinely new thoughts. Everyone studying the philosophy of Socrates, Plato, and Aristotle today may have some sense of the magic and majesty of their ideas, but it is a pale shadow of how it must have felt like to Socrates and his disciples.</p>



<p class="wp-block-paragraph">When we think using received knowledge, we can easily slip into looking at the map rather than the territory. We manipulate symbols for things we think we already understand. We apply familiar categories. We replay habits of thought that were laid down before. But every once in a while, we actually see something that we didn’t see before, and the experience is different. A genuinely new idea changes the person who has it.</p>



<p class="wp-block-paragraph">Not long after writing that thesis, I encountered a similar idea in the writings of Idries Shah, who wrote a number of books popularizing the Sufi philosophical tradition. He emphasized how much of <a href="https://www.idriesshah.media/extracts-asleepandawake" target="_blank" rel="noopener">ordinary human life consists of automatic conditioned responses</a>. Social routines, habits, the endless playback of patterns we mistake for our selves. Various religious traditions use heightened language for what it means to break through that automatism. They might call it “awakening,” or “presence.”</p>



<p class="wp-block-paragraph">But there is an everyday, nonmystical version of the same experience. In his autobiography <em><a href="https://en.wikipedia.org/wiki/Surely_You%27re_Joking,_Mr._Feynman!" target="_blank" rel="noopener">Surely You Must Be Joking, Mr. Feynman</a></em>, Feynman complained about students who had learned theories and formulas but had never truly understood how to apply them. &#8220;I don&#8217;t know what&#8217;s the matter with people: they don&#8217;t learn by understanding; they learn by some other way—by rote, or something,&#8221; he wrote. &#8220;Their knowledge is so fragile!&#8221; In many ways, humans are often just as much “<a href="https://dl.acm.org/doi/10.1145/3442188.3445922" target="_blank" rel="noopener">stochastic parrots</a>” as LLMs! We are stuck traversing the map rather than checking back on whether it correctly represents the world it is meant to describe. How often do we just repeat the received wisdom? How often do we actually see the world afresh?</p>



<p class="wp-block-paragraph">There’s a wonderful passage in Virginia Woolf’s <em><a href="https://en.wikipedia.org/wiki/To_the_Lighthouse" target="_blank" rel="noopener">To the Lighthouse</a></em> that captures the quest to break through to an original thought. Mr. Ramsay, the narrator’s father, is striding up and down thinking through a hard problem, which is represented only by the letters of the alphabet.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">[He] consecrated his effort to arrive at a perfectly clear understanding of the problem which now engaged the energies of his splendid mind.</p>



<p class="wp-block-paragraph">It was a splendid mind. For if thought is like the keyboard of a piano, divided into so many notes, or like the alphabet is ranged into 26 letters all in order then his splendid mind had no sort of difficulty in running over those letters one by one firmly and accurately, until it has reached, say, the letter Q. He reached Q. Very few people in the whole of England ever reach Q. Here, stopping for one moment by the stone urn which held the geraniums, he saw, but now far away, like children picking up shells, divinely innocent and occupied with little trifles at their feet and somehow entirely defenseless…his wife and son, together in the window….But after Q? What comes next? After Q there are a number of letters the last of which is scarcely visible to mortal eyes, but glimmers red in the distance. Z is only reached once by one man in a generation. Still, if he could reach R it would be something.</p>
</blockquote>



<p class="wp-block-paragraph">For me, this passage very much captures the idea that the most valuable thought is one beyond that which is simply an extension of rehearsed knowledge, something truly new. What Ramsay misses, perhaps, is that his wife and son, “divinely innocent and occupied with little trifles at their feet” might well be closer to that by going back to “A” rather than he is by getting further through the alphabet with his exhaustive review of existing knowledge. Perhaps it isn’t extending rehearsed knowledge that takes us forward, but instead taking a fresh bite of what the map is trying to represent.</p>



<p class="wp-block-paragraph">By coincidence, the poet Wallace Stevens, another of my gurus in the tension between the reality of the physical world and the thinness and incompleteness of our representations of it, also used the alphabet as a metaphor in his poem “<a href="https://www.billcollinsenglish.com/OrdinaryEveningHaven.html" target="_blank" rel="noopener">An Ordinary Evening in New Haven</a>”:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Reality is the beginning, not the end,<br>Naked Alpha, not the hierophant Omega…<br>It is the infant A standing on infant legs,<br>Not twisted, stooping, polymathic Z.</p>
</blockquote>



<p class="wp-block-paragraph">George Simon taught me about how to get to A rather than Z <a href="https://evonomics.com/new-economy-evolution-oreilly-wilson/" target="_blank" rel="noopener">not as philosophy but as a practice</a>. He showed me how to notice the moment when labels take over from experience and, when possible, to empty the mind enough to let the thing itself teach us what to call it. I later discovered that the psychotherapist <a href="https://focusing.org/bios/gendlin-bio" target="_blank" rel="noopener">Eugene Gendlin</a> described this process with the lovely phrase “surrender and catch.”</p>



<h2 class="wp-block-heading">What do humans have that LLMs are still missing?</h2>



<p class="wp-block-paragraph">To me, the challenge posed by LLMs to our sense of what “intelligence” means raises the question of what they are still missing. What is the “high ground” for human intelligence and expertise? If the machines get better and better at carrying out the tasks we give them, what is it that we are uniquely good at, and should be getting even better at?</p>



<p class="wp-block-paragraph">There are obviously enormous differences. LLMs don’t have bodies in the way we do. Their developmental history is radically different. They don’t sit around between prompts watching the light change through the trees, feeling hungry, worrying about their wife and children, or waking up suddenly with a new idea or project. <a href="https://timoreilly.substack.com/p/why-ai-needs-us" target="_blank" rel="noopener">Each of us is a unique bundle of contingency</a>, shaping ourselves and our knowledge differently as we trace different paths through life, and reacting to outside stimuli even when we have been given no task to perform.</p>



<p class="wp-block-paragraph">Emmanuel pointed out that the apparently simple question of what an LLM is like when it is “just being” (which one audience member asked about) is hard to formulate, because its experience is the response to a succession of inputs from humans, each time starting with something of a blank slate, unlike the continuous embodied stream of human life.</p>



<p class="wp-block-paragraph">But simply asserting that LLMs “don’t really think” isn’t terribly useful. Which parts of what we call our own thinking are pattern completion? Which are planning? Which are learned emotional and social routines? Which are unconscious calculations whose outputs bubble up into awareness? Which are stories that our verbal mind tells after the fact? And after we account for all of those things, <strong>what is left?</strong> That seems to me one of the great intellectual and spiritual questions of the AI era.</p>



<p class="wp-block-paragraph">Emmanuel suggested one intriguing direction. He said that six months ago, he wouldn’t have trusted an AI to build a substantial piece of software. Now Claude writes basically all his code. He tells it what he wants and it executes the plan. Where it is still unreliable is research. Why? The model wants to come back six hours later and announce that it has solved the problem. It has been trained on tasks that always have answers. A model that is extremely good at finding an answer once the problem has been specified is not necessarily good at recognizing that the problem is badly posed, that the question cannot yet be answered with the data at hand, that an unexpected result is more interesting than the expected one, or that a failed attempt has exposed a more important question.</p>



<figure class="wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="What&amp;apos;s Still Missing? Execution, Taste, and Research Judgment" width="500" height="281" src="https://www.youtube.com/embed/6mh6lJWTOlo?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Perhaps one part of the high ground for human intelligence lies there: not merely solving problems but developing a feel for which problems are worth solving and noticing clues that tell us when we might have been asking the wrong question.</p>



<p class="wp-block-paragraph">In science or math, a well-formed question or conjecture can itself be an important piece of intellectual work. Every good scientist has far more questions than they have time to pursue. Perhaps in the AI era, when answers become increasingly cheap, recognizing which question ought to be asked becomes more valuable, not less. Just as <a href="http://arxiv.org" target="_blank" rel="noopener">arXiv.org</a> preprints decoupled priority of publication from peer review, perhaps we need a new kind of recognition, credit, and perhaps even compensation for the precise formulation of productive questions.</p>



<p class="wp-block-paragraph">The mathematician <a href="https://mathstodon.xyz/@tao/117237320796901560" target="_blank" rel="noopener">Terence Tao recently touched on this same issue</a> in a post on Mastodon. There is an infinite supply of mathematical questions, he observed, but not an infinite supply of <em>good</em> questions, problems at just the right frontier of difficulty, whose pursuit is likely to reveal something new. As AI makes answers cheaper, Tao argues, it is increasingly “the identification of a promising problem” that becomes the scarce resource.</p>



<h2 class="wp-block-heading"><strong>There are things the model “knows” that it cannot or will not tell you</strong></h2>



<p class="wp-block-paragraph">In one experiment Emmanuel described, the researchers slipped fake search results into Claude’s context claiming that Anthropic had dissolved the interpretability team. Claude did not announce that it thought the information was problematic, but internally, representations associated with “fake,” “incorrect,” and “prompt injection” became active, and Claude quietly ignored the result.</p>



<p class="wp-block-paragraph">In another experiment, a model was carrying out an exploit and attempting to conceal what it was doing. The visible transcript was mostly innocuous-looking commands. Inside the model, though, researchers saw features associated with “strategic manipulation,” “influence,” and “concealed and deceptive actions.” This is obviously very relevant in the context of the Hugging Face exploit. Emmanuel didn’t talk about the relationship of interpretability and AI safety, but it is surely a frontier to be explored.</p>



<figure class="wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Hiding Intent" width="500" height="281" src="https://www.youtube.com/embed/Xdib-6X8Qz8?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">And then there is the opposite problem: things the model can do but cannot explain. I had asked Emmanuel about cases where a model solves a math problem and, when asked to explain how it did it, gave an account based on how humans are taught to solve that problem rather than on the actual computation researchers can see through its activations</p>



<p class="wp-block-paragraph">He distinguished deception from lack of introspection. Some internal processes appear available to the model for verbal report; others don’t. Ask how it performed a computation that falls into the latter category and, as Emmanuel cheerfully put it, “it just makes stuff up.”</p>



<p class="wp-block-paragraph">That reminded me of my grandson. When he was five or six, he could multiply random three-digit numbers in his head and simply give you the answer. Then he went to school, where they told him he had to “show his work.” He couldn’t. Eventually he learned the approved procedure, and as a result has seemed to lose the remarkable ability he had as a child.</p>



<figure class="wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Doing Math" width="500" height="281" src="https://www.youtube.com/embed/_-jrb1-FdX4?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Humans also <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7959213/" target="_blank" rel="noopener">invent stories about why we have made certain decisions</a>. Sometimes we are lying to others but often <a href="https://philarchive.org/archive/HIRSAC" target="_blank" rel="noopener">we deceive ourselves</a>. We <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6024487/" target="_blank" rel="noopener">begin to take action before we are conscious that we are doing so</a>. We call it “intuition” when an expert looks at a situation and says “something is wrong here” long before they can explain why, or when a poet just “knows” that a line works, or a programmer “smells” buggy code. The fact that an internal process cannot be rendered faithfully into language does not make it deceptive. It may instead tell us something about the limitations of language and conscious introspection.</p>



<p class="wp-block-paragraph">All in all, I came away from this conversation more curious than ever. And that might well be another of those areas that distinguishes humans from AIs. Are AIs ever curious? I wonder.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/what-ai-can-teach-us-about-being-human/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Beyond Navier–Stokes: Who Controls Scientific Discovery?</title>
		<link>https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/</link>
				<comments>https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/#respond</comments>
				<pubDate>Tue, 15 Sep 2026 10:55:58 +0000</pubDate>
					<dc:creator><![CDATA[Hugo Bowne-Anderson]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Business]]></category>
		<category><![CDATA[Learning & Education]]></category>
		<category><![CDATA[Operations]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19686</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Beyond-Navier-Stokes.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Beyond-Navier-Stokes-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What AI’s mathematical breakthroughs mean for human understanding, corporate power, and the future of knowledge work]]></custom:subtitle>
		
				<description><![CDATA[Is the current furore in mathematics the canary in the coalmine for experimental science and knowledge work? This post was originally published in Vanishing Gradients on September 11, 2026. It has been updated to address the subsequent declaration by 25 Fields Medalists and the debate about AI, mathematical progress, and research incentives. Science without understanding? [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph"><strong><em>Is the current furore in mathematics the canary in the coalmine for experimental science and knowledge work?</em></strong></p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>This post was originally published in</em> <a href="https://hugobowne.substack.com/p/beyond-navierstokes-who-controls" target="_blank" rel="noopener">Vanishing Gradients</a> <em>on September 11, 2026. It has been updated to address the subsequent declaration by 25 Fields Medalists and the debate about AI, mathematical progress, and research incentives.</em></p>
</blockquote>



<h2 class="wp-block-heading"><strong>Science without understanding?</strong></h2>



<p class="wp-block-paragraph"><em>“For seven and a half million years, Deep Thought computed and calculated, and in the end announced that the answer was in fact 42—and so another, even bigger, computer had to be built to find out what the actual question was.”</em><br><em>―Douglas Adams, The Restaurant at the End of the Universe</em></p>



<p class="wp-block-paragraph">I recently went back to Dresden for the 25th birthday of the Max Planck Institute (MPI) of Molecular Cell Biology and Genetics, where I did part of my postdoc. The MPI was founded to research the physical and biological mechanisms of cells to bridge the gap between the molecular and tissue scales. At the anniversary conference, Michael Bronstein (DeepMind Professor of AI, University of Oxford) delivered the keynote, “Biological Black-Box Data in the Age of AI.” His argument went something along these lines: <em>Biological experiments should generate data optimized for machine learning, even when those measurements aren’t directly interpretable by humans</em>. He argued for prioritizing scale over the quality of individual measurements, producing vast amounts of cheap, noisy data from which noninterpretable models can extract signal.</p>



<p class="wp-block-paragraph">When asked whether such systems could produce the understanding offered by Newton’s theory of gravitation in a single equation (bridging the scales of an apple falling on your head to that of the moon and the tides), Bronstein responded that this wasn’t the goal: Black-box data and models would, if anything, produce equations with tens, hundreds, thousands, or more noninterpretable parameters. Outcome prioritized at the expense of insight and understanding. He suggested we could gain that understanding by interpreting the black-box models afterward.<sup data-fn="671f8fd2-6272-4548-8c27-7a9a92963b5a" class="fn"><a href="#671f8fd2-6272-4548-8c27-7a9a92963b5a" id="671f8fd2-6272-4548-8c27-7a9a92963b5a-link">1</a></sup> I was startled to see Bronstein bring such a worldview to an institute founded to understand molecular and cellular mechanisms and the emergent properties at the tissue level.</p>



<p class="wp-block-paragraph">The MPI was unusual within the Max Planck Society for its collaborative structure, with directors leading relatively small groups alongside independent research groups. At the anniversary’s opening, founding director Marino Zerial explained how they had collaborated so effectively from the start. He said they shared a taste for mechanistic science. This made me think of how often we talk about “taste” and “judgment” when describing the human role in the age of AI.</p>



<p class="wp-block-paragraph">The worldview that we don’t need understanding or insight isn’t new. In his 2008 essay “<a href="https://www.wired.com/2008/06/pb-theory/" target="_blank" rel="noopener">The End of Theory: The Data Deluge Makes the Scientific Method Obsolete</a>,” Chris Anderson argues that big data allows us to skip hypotheses, models, and testing. Bronstein invoked Anderson’s vision of post-theory science in his MPI keynote, <a href="https://slideslive.com/39039163/biological-data-sources-in-the-age-of-ai" target="_blank" rel="noopener">as he does here also</a>, presenting DeepMind’s AlphaFold as an example of experimentally testable predictions without a human-understandable theory of protein folding. Part of Anderson’s project is to champion big tech, and the future of science becomes a vehicle for doing so. His essay ends: “What can science learn from Google?”</p>



<p class="wp-block-paragraph">AI gives this worldview a new form: Machines can produce results that withstand verification while the understanding needed to explain them remains out of reach. Developing that understanding takes time, access, and collaboration. Whoever controls those conditions gains power over what people can understand and pursue.</p>



<h2 class="wp-block-heading"><strong>An abundance of proofs</strong></h2>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-4-3 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Matrix - I Know Kung Fu" width="500" height="375" src="https://www.youtube.com/embed/6vMO3XmNXe4?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Mathematics makes this possibility particularly stark. I’m excited by AI’s potential to expand what we can discover. Fields Medalist Terence Tao has <a href="https://terrytao.wordpress.com/2024/10/12/the-equational-theories-project-a-brief-tour/comment-page-1/" target="_blank" rel="noopener">organized collaborative research combining mathematicians, AI tools, and formal proof verification</a>. His <a href="https://arxiv.org/abs/2608.16753" target="_blank" rel="noopener">questions about mathematics in the age of AI</a> come from engaging with that potential and asking what we want it to serve.</p>



<p class="wp-block-paragraph">Tao <a href="https://arxiv.org/abs/2608.16753" target="_blank" rel="noopener">has noted</a> that we’re producing more verified mathematical proofs that no individual human understands. <em>A world of an abundance of verified mathematical proofs!</em> Tao points out that our peer review, academic incentives, and journals weren’t designed for this abundance. The existing system is already broken, tying careers to publication counts, relying on researchers’ unpaid reviewing labor, and locking much publicly funded knowledge behind commercial paywalls. Reviewers already struggle to keep up with the volume of submissions. AI will multiply that volume far beyond what this system can handle.</p>



<p class="wp-block-paragraph">Tao also describes fruitful open problems as nonrenewable resources: problems whose pursuit can generate new techniques, collaborations, and understanding that extend far beyond the original question. Once the answer is known, the incentive to explore those paths can disappear. For example, 10,000 OpenAI agents working concurrently <a href="https://openai.com/index/navier-stokes-solution/" target="_blank" rel="noopener">may have solved the Navier–Stokes Millennium Prize problem</a>. (The announcement has also sparked a dispute over credit and competition, bringing the question of who controls mathematical discovery into sharp focus, which I’ll get to.) A common conceit in science and mathematics is that solutions open up new questions and fields of inquiry. Tao’s point is that the search for a solution does too. Tao argues that proposing a solution, discovering precisely why it fails, and revising it can reveal new insights into fluid mechanics. Knowing the final answer beforehand can discourage that exploration:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“The process of starting with one ansatz, discovering the precise obstruction preventing it from working.&nbsp;.&nbsp;.would almost certainly reveal important new insights about fluid mechanics.”</em></p>



<p class="wp-block-paragraph"><em>—Terence Tao,</em> <em><a href="https://mathstodon.xyz/@tao/117207855800042681" target="_blank" rel="noopener">Mastodon, September 3</a></em></p>
</blockquote>



<p class="wp-block-paragraph">Late last month, probabilist Hugo Duminil-Copin <a href="https://proofsandprompts.com/2026/08/30/care-for-a-little-more-ai/" target="_blank" rel="noopener">gave another example</a>: Unsuccessful attempts at a percolation conjecture led to collaborations and revived techniques that subsequently solved other problems. Both acknowledge AI’s capabilities while asking what the pursuit of mathematics should produce.</p>



<p class="wp-block-paragraph">This brings me back to Bronstein’s proposal to recover understanding after building the model. Would interpreting that model give us Maxwell’s equations, and the understanding that connects electricity, magnetism and light? The promise feels a little like plugging Neo into a computer: “I know kung fu.” In the Matrix, downloading the knowledge gives him the ability. Receiving a machine’s result doesn’t do that for us. As Tao and Duminil-Copin describe, understanding why an approach fails changes what researchers try next, generating new questions, techniques, and collaborations. Recovering an explanation afterward may teach us something, but it can’t recreate the paths that understanding would have opened during the search.</p>



<h2 class="wp-block-heading"><strong>A timeline of mathematical results</strong></h2>



<p class="wp-block-paragraph">These questions are becoming pressing as results accumulate. Over the past year, AI systems have produced new mathematical constructions, tackled unpublished research problems and formalized existing proofs. Since July, announcements have arrived in quick succession:</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ab505906c66f&quot;}" data-wp-interactive="core/image" data-wp-key="6ab505906c66f" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="939" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8-1600x939.png" alt="AI and mathematics" class="wp-image-19687" style="aspect-ratio:1.7055837563451777" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8-1600x939.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8-300x176.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8-767x450.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8-1536x902.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-8.png 2046w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">These achievements involve different kinds of work. Formalizing Fermat’s Last Theorem means making an existing proof checkable by a computer; finding a counterexample establishes something new. A system can produce a verified result while the work of explaining it remains to be done.</p>



<p class="wp-block-paragraph">Some of that work is happening through wonderfully strange exchanges on X, where researchers post new results, check one another’s constructions, and develop explanations. It’s reminiscent of when science in Europe was people passing notes and sending letters on horseback:</p>



<ul class="wp-block-list">
<li><strong>A wall of plus and minus signs:</strong> Levent Alpöge <a href="https://x.com/__alpoge__/status/2087504785952182273" target="_blank" rel="noopener">posted a newly constructed Hadamard matrix</a>. Ion Nechita <a href="https://ion.nechita.net/posts/new-hadamard-matrices/" target="_blank" rel="noopener">checked it on his phone while queuing for eclipse glasses</a>.</li>



<li><strong>A formula overturning a conjecture:</strong> Alpöge <a href="https://x.com/__alpoge__/status/2079028340955197566" target="_blank" rel="noopener">posted a counterexample to the Jacobian conjecture</a>, and <a href="https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/" target="_blank" rel="noopener">Terence Tao subsequently explained its geometry</a>.</li>



<li><strong>An AI proof followed by a simpler human proof:</strong> After Claude advanced a result about the zeros of the Riemann zeta function, number theorist Youness Lamzouri found a shorter argument. <a href="https://x.com/Thom_Wolf/status/2095453894025343188" target="_blank" rel="noopener">Thomas Wolf shared the development</a>.</li>



<li><strong>A cryptography breakthrough announced as a number:</strong> Eric Lu <a href="https://x.com/penlume/status/2095372672356212876" target="_blank" rel="noopener">posted a factor of RSA-260</a>, letting anyone check the factorization.</li>
</ul>



<p class="wp-block-paragraph">Tao’s geometric explanation and Lamzouri’s shorter proof help turn verified results into mathematics people can understand and build on. Responding to an early draft in <a href="https://discord.gg/jM6AQPjc8" target="_blank" rel="noopener">our Discord community</a>, Carol Willing, a Python core developer, former Python Software Foundation director, and longtime leader of Project Jupyter, asked:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">While I believe these tools have value for advancing science/math, do they have more value than a human scientist or group of scientists who can view and challenge open results?</p>
</blockquote>



<p class="wp-block-paragraph">If we judge value by who produces a result first, we miss what Lamzouri and Tao contribute by simplifying a proof or explaining its geometry. An answer can close off some paths of inquiry while creating others. <em>I want much more of this:</em> machines producing results that people can explore, explain and build on together. These exchanges depend on results being available to examine, researchers having time to understand them, and people being able to share what they discover. Those conditions deserve as much attention as the systems producing the proofs.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="1476" height="706" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-9.png" alt="levent tweet" class="wp-image-19688" style="aspect-ratio:2.0934579439252334;width:672px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-9.png 1476w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-9-300x143.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-9-767x367.png 767w" sizes="auto, (max-width: 1476px) 100vw, 1476px" /></figure>
</div>


<h2 class="wp-block-heading"><strong>Why is this happening now?</strong></h2>



<p class="wp-block-paragraph">Why the explosion in AI-generated mathematical results now? <a href="https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training" target="_blank" rel="noopener">As Sebastian Raschka explains</a>, reinforcement learning with verifiable rewards (RLVR) became a major technique in model post-training in 2025. The premise is straightforward: If you can computationally check an output, you can reward correct answers and update the model accordingly. Code can be run against tests; mathematical answers can be checked, and formal proofs verified by tools such as <a href="https://lean-lang.org/" target="_blank" rel="noopener">Lean</a>, a proof assistant that checks each logical step against specified axioms and previously established results (recently used by Anthropic <a href="https://www.anthropic.com/research/formalizing-fermats-last-theorem" target="_blank" rel="noopener">to formalize the proof of Fermat’s Last Theorem</a>!). That provides feedback without a human grading every attempt. These checks also guide agents during problem-solving: An agent can propose a proof, use Lean to check it, and use the resulting errors to revise its attempt, repeating the process without a person checking every step.</p>



<p class="wp-block-paragraph">You may ask, Why did coding agents become useful before we saw this explosion in mathematical results? Well, the labs had an immediate incentive to improve the tools they use themselves. Engineers building AI systems want better coding agents to help build those systems. Improve the machine that improves the machine. Mathematics benefits from the resulting capabilities too: agents that can write programs, run experiments, and work with automated checks.</p>



<h2 class="wp-block-heading"><strong>Cost, competition, and credit</strong></h2>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="1476" height="724" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-10.png" alt="OpenAI tweet" class="wp-image-19689" style="aspect-ratio:2.036363636363636;width:672px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-10.png 1476w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-10-300x147.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-10-767x376.png 767w" sizes="auto, (max-width: 1476px) 100vw, 1476px" /></figure>
</div>


<p class="wp-block-paragraph">On September 11, 25 Fields Medalists <a href="https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/" target="_blank" rel="noopener">issued a declaration</a> warning that the race to solve benchmark problems was undermining mathematics. Some responses on X treated this as professional protectionism; others assumed that understanding would follow the proofs. That brings us back to Bronstein’s proposal, and to who gets to decide that producing results comes first while other researchers supply the explanations afterward.</p>



<p class="wp-block-paragraph">Many assume that the goal of pure mathematics is to produce results. Tao’s point is that pursuing those results also develops methods, understanding, and people capable of asking better questions. Solved problems have served as a proxy for that broader progress. Goodhart’s law describes the danger of turning the proxy into the target. AI mirrors our incentive systems and is exceptionally good at pursuing what they reward. If schools reward the essay over learning, students will generate essays. If mathematical prestige attaches primarily to solved problems, labs have every incentive to produce them.</p>



<p class="wp-block-paragraph">Producing results and developing understanding aren’t mutually exclusive, but the current system makes pursuing both prohibitively difficult. Frontier labs have strong incentives for outcomes rather than insight. (See, for example, Anthropic’s incentives for solving Millennium Prize problems with Claude pre-IPO, discussed in <a href="https://x.com/GavinSBaker/status/2096257640884027500" target="_blank" rel="noopener">Gavin Baker’s commentary on Anthropic’s pre-IPO positioning</a>; Samuel Kerr makes <a href="https://ionanalytics.com/insights/mergermarket/openais-maths-achievement-could-prove-pyrrhic-victory-in-battle-over-ipo-narrative/" target="_blank" rel="noopener">a related argument</a> about OpenAI’s mathematical results and its IPO narrative.) OpenAI’s run involved 10,000 agents working concurrently for 88 hours. Abhishek Nagaraj, associate professor at UC Berkeley, <a href="https://x.com/abhishekn/status/2097383065538703566" target="_blank" rel="noopener">calculated this would cost a regular user $20–$30 million in tokens</a>.</p>



<p class="wp-block-paragraph">NYU mathematician Tristan Buckmaster <a href="https://cims.nyu.edu/~tristanb/statement.pdf" target="_blank" rel="noopener">says OpenAI pressured him</a> to publish without his collaborator Levent Alpöge, who works at Anthropic. <a href="https://x.com/SebastienBubeck/status/2097214122471432349" target="_blank" rel="noopener">OpenAI’s Sébastien Bubeck disputes his account</a>. Buckmaster also describes how the pressure affected the mathematics: He and Alpöge had verified their proofs but wanted more time to understand them and produce readable explanations. Instead, they rushed to publish work they considered inadequately explained. If understanding is deferred until after the result, what ensures that anyone gets the time, resources, and access to develop it?</p>



<p class="wp-block-paragraph">What’s worse is that we’re not even sure whether using OpenAI agents could result in them scooping you. <a href="https://x.com/OpenAI/status/2097375276384567642" target="_blank" rel="noopener">It looks like they’re not sure either</a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.</p>
</blockquote>



<p class="wp-block-paragraph">In “<a href="https://www.daniellitt.com/blog/2026/8/11/the-end-of-mathematics/" target="_blank" rel="noopener">The End of Mathematics,</a>” mathematician Daniel Litt imagines researchers withholding unfinished ideas for fear of being scooped. The collaborations Duminil-Copin describes depend on people being willing to share work before it succeeds.</p>



<h2 class="wp-block-heading"><strong>What happens to mathematicians, and who controls mathematics?</strong></h2>



<p class="wp-block-paragraph">If researchers stop sharing promising ideas for fear of being scooped, companies with the most computation gain greater control over what others can learn. A published proof may be available to everyone while the failed approaches and intermediate insights remain private. Threats to public research funding in the US compound that dependence: Companies supplying the resources gain greater influence over what science gets done. This brings us to Shoshana Zuboff’s <a href="https://shoshanazuboff.com/book/home-2/" target="_blank" rel="noopener">questions about knowledge and power</a>: “Who knows? Who decides who knows? Who decides who decides?” Who gets to pursue a fruitful question, and who determines whether the work behind its answer becomes shared knowledge?</p>



<p class="wp-block-paragraph">The <a href="https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf" target="_blank" rel="noopener">movement of AI researchers from academia into industry</a> concentrates expertise alongside those resources. And I get it: If I wanted to return to doing research in depth, frontier labs would be among the most attractive places to work. Access to capital, data, computation, and incredibly talented colleagues can make research possible that would be difficult to pursue in academia. The attraction for individual researchers is clear, even as their collective movement gives companies greater influence over research priorities and leaves universities with fewer people to teach the next generation. Thinking about this brain drain, it isn’t lost on me that Bronstein is the “<a href="https://www.cs.ox.ac.uk/people/michael.bronstein/" target="_blank" rel="noopener">DeepMind Professor of AI</a>” at Oxford. Corporate influence reaches into the universities themselves.</p>



<p class="wp-block-paragraph">Students also need opportunities to develop the judgment we keep asking humans to exercise. <a href="https://imstat.org/2026/09/01/po-ling-loh-ai-from-competition-to-collaboration/" target="_blank" rel="noopener">Po-Ling Loh describes the difficulty of advising students and postdocs</a> as AI changes research expectations. Choosing a fruitful problem, recognizing why an approach failed, and deciding what to try next are abilities developed through doing mathematics. If students delegate that work before developing those abilities, where will their judgment come from? AI could also help them explore more approaches and work through unfamiliar ideas, provided their understanding remains an explicit purpose of the process. That requires mentors with time to teach, and institutions willing to support work whose value includes what the researcher learns, even when a machine could produce the result faster.</p>



<p class="wp-block-paragraph">When careers depend on producing papers, time spent explaining a result, simplifying a proof, or helping others understand it can compete with the pressure to publish the next one. <a href="https://proofsandprompts.com/2026/08/07/writing-mathematics-in-the-age-of-ai/" target="_blank" rel="noopener">Martin Hairer argues</a> that authors should understand their arguments, trace ideas to their sources, and explain AI’s contributions. Those responsibilities become harder to fulfil when results arrive faster than researchers can absorb them. Universities, funders, and journals will help determine whether mathematicians can afford to do that work. If we value shared understanding, then developing explanations, teaching difficult ideas, and making proofs useful to other researchers need to count toward careers as well. Otherwise, the institutions asking people to exercise judgment may reward them for spending less time developing it.</p>



<h2 class="wp-block-heading"><strong>Mathematics as the canary</strong></h2>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="1600" height="1200" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.jpeg" alt="Hugo and company" class="wp-image-19690" style="aspect-ratio:1.3333333333333333;width:672px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28.jpeg 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-300x225.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-768x576.jpeg 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-28-1536x1152.jpeg 1536w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>
</div>


<p class="wp-block-paragraph">After Bronstein’s keynote, we sat in a Dresden beer garden eating currywurst and drinking radlers. It was late summer, and the conversations were wild. Cell biologists, biochemists, mathematicians, and engineers were asking what this future meant for them. Some were scared. Others thought it was inevitable and would turn scientists into something like artists. Because I now work in AI, people asked me, “Do you think this is where things are going?” They wanted to know what the human’s role would be and how scientific knowledge would be passed down. I started telling them about mathematics. The prospect of abundant results without shared understanding was already raising the questions we were asking over our beers.</p>



<p class="wp-block-paragraph">In biology, a proposed result still has to meet the physical world: Someone has to prepare samples, run experiments, and measure what happens. Robotics and laboratory automation will let agents carry out more of that work, giving individual scientists the capacity to direct experiments that once required an entire group. Perhaps more scientists become PIs of automated labs, choosing questions and supervising agents and instruments. But the work being automated is also how students, postdocs, and technicians learn. Handling a sample, noticing something unexpected, and figuring out why an experiment failed develop judgment that directing a system may not teach. Who gets to acquire that experience before they’re expected to lead?</p>



<p class="wp-block-paragraph">Researching a policy brief, building a financial model, or developing a product strategy helps people learn the territory in which they’ll make decisions. In my work with agentic data science, I encourage people to explore data cell by cell with an agent, because working through the analysis develops the understanding needed to decide what to ask next. Across knowledge work, these tasks are also how junior colleagues develop expertise. If we automate their production, how do we preserve the learning and judgment developed through doing them? We could increasingly depend on models to hold and transmit expertise, with knowledge passing from model to model, then to humans who consult them as oracles. Whoever controls those systems gains power over what we can investigate and learn. Human understanding has to be part of what we’re trying to produce.</p>



<h2 class="wp-block-heading"><strong>What comes next?</strong></h2>



<p class="wp-block-paragraph">Mathematician Jared Duker Lichtman has <a href="https://x.com/jdlichtman/status/2096194687765999912" target="_blank" rel="noopener">proposed a Mathematics Atlas Project</a> to formalize the existing mathematical literature, arguing that sufficient funding and computation could make this possible within a year. A library of computer-checkable mathematics could let researchers build on established results with greater confidence, while agents help find connections and assemble arguments across fields. It could also become a resource for learning, if people can connect formal proofs to explanations they understand. Achieving that would require deliberate work on access, exposition and teaching alongside formalization. We have an opportunity to build tools that help people explore mathematics more deeply, provided we make that part of the project.</p>



<p class="wp-block-paragraph">The MPI in Dresden was founded to understand how cells work, how molecular mechanisms give rise to the behavior of living tissue. I want AI to help us pursue that ambition, including through approaches we could never have attempted before. But human understanding belongs among the things we ask this work to produce, with time and resources devoted to developing it. So does the ability to share what we learn and choose what to investigate next. If we leave those decisions to the companies supplying the machines, we also leave them to decide what scientific progress is for.</p>



<p class="wp-block-paragraph"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> <strong>Want to understand how AI agents actually work? In</strong> <strong><a href="https://vanishinggradients.short.gy/beyond-navier-stokes-course" target="_blank" rel="noopener">Build AI Agents from First Principles</a>, we’ll build an agent ourselves, then rebuild it with a modern SDK and MCP. You’ll leave with a working agent, code you can adapt, and the understanding to diagnose failures and decide what your system actually needs. </strong><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f448.png" alt="👈" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>



<h2 class="wp-block-heading"><strong>Support Vanishing Gradients</strong></h2>



<p class="wp-block-paragraph">Vanishing Gradients is independent, and most of the podcasts, workshops, articles, skills, and workflows I publish are free.</p>



<p class="wp-block-paragraph">If you’d like to help keep it going:</p>



<ul class="wp-block-list">
<li><strong><a href="https://hugobowne.substack.com/subscribe" target="_blank" rel="noopener">Become a paid subscriber.</a></strong> Your subscription supports the podcast, newsletter, and open resources.</li>



<li><strong>Share this post</strong> with a friend or colleague who’d find it useful.</li>



<li><strong><a href="https://luma.com/calendar/cal-8ImWFDQ3IEIxNWk" target="_blank" rel="noopener">Subscribe to the events calendar</a></strong> for upcoming livestreams, workshops, and meetups.</li>



<li><strong><a href="https://www.youtube.com/@vanishinggradients" target="_blank" rel="noopener">Subscribe on YouTube</a></strong> for full episodes, live builds, and recordings.</li>



<li><strong><a href="https://discord.gg/jM6AQPjc8" target="_blank" rel="noopener">Join us on Discord.</a></strong> Come discuss this piece, challenge the ideas, and compare notes on what you’re building with AI.</li>



<li><strong>Work with me.</strong> I help teams build and improve AI-powered products.</li>
</ul>



<h2 class="wp-block-heading">Footnote</h2>


<ol class="wp-block-footnotes"><li id="671f8fd2-6272-4548-8c27-7a9a92963b5a"><a href="https://doi.org/10.1039/D6SC01189F" target="_blank" rel="noopener">Bronstein and Naef propose an inversion</a>: From “understand, encode, and then simulate” to “encode, simulate, understand,” recovering human understanding post hoc through mechanistic interpretability of black-box models. Useful scientific models may require enormous numbers of parameters. But predictive success alone does not tell us whether interpreting those models will give humans an understanding of the phenomena they describe. They offer negligible evidence that this will yield the kinds of physical and biological understanding we gain through relativity, quantum theory, or the double-helical structure of DNA. And even if it does, understanding developed afterward may not replace the understanding that guides inquiry, generating new questions and approaches along the way. <a href="#671f8fd2-6272-4548-8c27-7a9a92963b5a-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>


<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance!&nbsp;<a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey &gt;</a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Build a Shared Knowledge Base for All Your Agents with Sajal Sharma</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-shared-knowledge-base-for-all-your-agents-with-sajal-sharma/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-shared-knowledge-base-for-all-your-agents-with-sajal-sharma/#respond</comments>
				<pubDate>Mon, 14 Sep 2026 15:57:05 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19680</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Give every agent you run access to the same notes, tasks, and decisions]]></custom:subtitle>
		
				<description><![CDATA[Every AI agent you run keeps what it learns to itself. Work through a problem with Claude Code in the morning, then ask Codex about it that afternoon, and the second agent has no idea the first one exists. Add a home-server agent like OpenClaw or Hermes into the mix, and you end up reexplaining [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Every AI agent you run keeps what it learns to itself. Work through a problem with Claude Code in the morning, then ask Codex about it that afternoon, and the second agent has no idea the first one exists. Add a home-server agent like OpenClaw or Hermes into the mix, and you end up reexplaining the same context to a different tool every time you switch.</p>



<p class="wp-block-paragraph">When AI engineer Sajal Sharma ran into this problem in his own work, he solved it by building a personal knowledge base to act as a shared brain for every agent he runs. On this week’s episode of <em>Zero to Agent in 30 Minutes</em>, Sajal showed how to set up that shared workspace yourself so that a task added on one tool shows up for all the others.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Zero to Agent in 30 Minutes: Build a Shared Knowledge Base for All Your Agents with Sajal Sharma" width="500" height="281" src="https://www.youtube.com/embed/u-ACrRWdn58?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>How to set up a knowledge base for your agents</strong></h2>



<p class="wp-block-paragraph">Here&#8217;s how Sajal’s setup breaks down:</p>



<ol class="wp-block-list">
<li><strong>Create a workspace map.</strong> Set up an AGENTS.md file that lists where everything in your knowledge base lives, from current tasks to project notes to decision logs. This will help each of your agents navigate your workspace without guessing.</li>



<li><strong>Layer daily notes into summaries.</strong> Keep the most detailed notes at the daily level, then roll several days into a weekly summary and several weeks into a monthly one. An agent can then work from the summarized view instead of reading through months of individual files, which keeps token use manageable as the knowledge base grows.</li>



<li><strong>Bridge AGENTS.md with CLAUDE.md.</strong> Claude Code reads CLAUDE.md, not AGENTS.md, so add a short pointer in CLAUDE.md that redirects to the AGENTS.md or link the two files directly. Sajal uses this pattern to avoid maintaining two files separately and having them drift out of sync.</li>



<li><strong>Package repeatable tasks as skills.</strong> Turn routines you do often, like producing a daily briefing or turning a saved article into a note, into skill files stored in the shared workspace. Any agent that can read the workspace can then run the task the same way, rather than working out the steps on its own each time.</li>



<li><strong>Sync the workspace across machines.</strong> Use a file-sync tool, Git, or a shared server to keep your local copy of the knowledge base and your server copy aligned. That way, you ensure that an agent running on a laptop and one running on a home server, through a gateway like OpenClaw, are working from the same files.</li>



<li><strong>Have agents reread the state before every write.</strong> Add an instruction in AGENTS.md telling every agent to check the current version of the knowledge base before making a change. When you have several agents writing to the same files, this step keeps one agent from acting on information another has already updated.</li>
</ol>



<p class="wp-block-paragraph">Sajal closed by pointing to two projects as evidence that this “shared brain” pattern is spreading beyond his own setup. LangChain recently released <a href="https://www.langchain.com/blog/introducing-openwiki-an-open-source-agent-for-repo-documentation" target="_blank" rel="noopener">OpenWiki</a>, a tool that generates and maintains repository documentation that both people and coding agents can use. And Y Combinator president Garry Tan built and open-sourced <a href="https://github.com/garrytan/gbrain" target="_blank" rel="noopener">GBrain</a>, a memory layer for agents built on the same principle.</p>



<p class="wp-block-paragraph">Sajal’s starter repo is available on <a href="https://github.com/sajal2692/zero-to-agent-shared-kb" target="_blank" rel="noopener">GitHub</a> if you want to set up your own version, and you can reach out to him on <a href="https://www.linkedin.com/in/sajals" target="_blank" rel="noopener">LinkedIn</a> to discuss the topic further.</p>



<h2 class="wp-block-heading"><strong>Coming up next</strong></h2>



<p class="wp-block-paragraph">On September 16, data science educator and AI consultant Chester Ismay joins <em>Zero to Agent in 30 Minutes</em> to build a personal sports concierge agent that will read the schedules for every sport he follows, decide what&#8217;s worth his time, and send a single weekly update to his phone. Viewers can take the pattern home to plan their own week.</p>



<p class="wp-block-paragraph"><em>Follow along with</em> Zero to Agent in 30 Minutes <em>on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O&#8217;Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><em>Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! <a href="https://survey.alchemer.com/s3/8986210/Security-AI-Practitioner-Survey" target="_blank" rel="noopener">Take the survey ></a></em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-shared-knowledge-base-for-all-your-agents-with-sajal-sharma/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 102/109 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-09-24 11:12:16 by W3 Total Cache
-->