<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Fri, 07 Aug 2026 15:55:23 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.3</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>This Week in AI: Who Controls AI?</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-who-controls-ai/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-who-controls-ai/#respond</comments>
				<pubDate>Fri, 07 Aug 2026 15:55:09 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Podcast]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19322</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-scaled.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2560" 
				height="2560" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Plus the US-China technology rivalry, the cost of frontier competition, and AI’s growing role in mathematics and science]]></custom:subtitle>
		
				<description><![CDATA[Governments are tightening control over AI infrastructure as companies spend heavily to compete at the frontier. This week, data and AI evangelist Christina Stathopoulos examined how policy, capital, product risk, and scientific research are shaping AI development. She explored AI sovereignty, Google’s infrastructure spending and product risks, the singularity debate, and several developments in mathematical [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Governments are tightening control over AI infrastructure as companies spend heavily to compete at the frontier. This week, data and AI evangelist Christina Stathopoulos examined how policy, capital, product risk, and scientific research are shaping AI development.</p>



<p class="wp-block-paragraph">She explored AI sovereignty, Google’s infrastructure spending and product risks, the singularity debate, and several developments in mathematical and scientific research. The episode covered Anthropic CEO <a href="https://www.anthropic.com/news/position-open-weights-models" target="_blank" rel="noreferrer noopener">Dario Amodei’s argument about open weight models</a>, US restrictions targeting foreign-made humanoid robots, OpenAI’s researcher access program and Astra model, Claude Fable 5’s role in a long-standing math problem, and Google DeepMind’s AlphaFold reorganization.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe title="This Week in AI: Who Controls AI?" width="500" height="281" src="https://www.youtube.com/embed/o3f1SxPKrzw?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>AI sovereignty now reaches models, robots, chips, and energy</strong></h2>



<p class="wp-block-paragraph">We’ve followed the sovereignty conversation in recent episodes as governments have tightened control over model access, computing infrastructure, and supply chains. Amodei wants policymakers to focus on what a model can do instead of using its open or closed status as the main measure of risk. His proposals include restricting access to advanced chips and chipmaking technology, preventing industrial-scale model distillation, and requiring safety testing for sufficiently capable systems. Christina agrees that capabilities should come first when assessing risk, but argues that open and closed models each present unique challenges. Open models can’t be recalled or controlled once their weights are released, increasing the risk of misuse, while closed models concentrate power in the hands of a few companies. Rather than favoring one approach, policymakers should address the risks of both.</p>



<p class="wp-block-paragraph">The US-China rivalry is also moving into robotics. Christina discussed <a href="https://apnews.com/article/china-us-humanoid-robots-ban-tech-c9f5e3c94d91d00eff3b61b141fab366" target="_blank" rel="noreferrer noopener">new US restrictions on foreign-made humanoid robots</a>, aimed largely at Chinese manufacturers, and possible retaliation from Beijing. She argued that China’s manufacturing advantage could make the restrictions more costly for US buyers in the near term, even if they encourage domestic development over time.</p>



<p class="wp-block-paragraph">Europe and Australia are taking different paths. Europe is pursuing computing capacity through proposed AI gigafactories, while <a href="https://www.pm.gov.au/media/ai-australias-interests-0" target="_blank" rel="noreferrer noopener">Australia is emphasizing standards, renewable energy use, and creator rights</a>. Its proposals include requiring data centers to fund new clean energy and AI companies to obtain permission before training on creators’ work. National rules are becoming another factor in decisions about models, cloud providers, and data locations.</p>



<h2 class="wp-block-heading"><strong>Frontier competition is expensive, and new products carry risk</strong></h2>



<p class="wp-block-paragraph">Google reported its first quarter of negative free cash flow since going public after spending $44.9 billion on AI infrastructure in three months. Christina noted that the company generated about $39 billion in cash but spent roughly $45 billion, leaving it almost $6 billion in the red. This illustrates the scale of Google’s investment in frontier AI, even with a highly profitable core business.</p>



<p class="wp-block-paragraph">Large budgets don’t guarantee that products are ready for broad use. <a href="https://www.bbc.com/news/articles/c9349yx2ydvo" target="_blank" rel="noreferrer noopener">Google removed an AI-powered Google Earth feature</a> one day after launch when researchers used it to create realistic fake satellite images, including fabricated disasters and damaged landmarks. In this context, synthetic satellite imagery can weaken trust because viewers may treat it as documentary evidence.</p>



<p class="wp-block-paragraph">Large infrastructure budgets can accelerate model development and product releases, but they don’t replace careful evaluation, context-specific safeguards, and clear limits on where generation should be allowed. Product teams need to assess how people might abuse a new feature or how users will interpret an output, not only what the underlying model can produce.</p>



<h2 class="wp-block-heading"><strong>Scientific results offer more testable evidence than singularity claims</strong></h2>



<p class="wp-block-paragraph">OpenAI CEO <a href="https://www.businessinsider.com/sam-altman-openai-the-singularity-agi-prediction-anthropic-nvidia-2026-7" target="_blank" rel="noreferrer noopener">Sam Altman said that the AI singularity has begun</a>, referring to a period when AI accelerates human and technological progress at a rapidly increasing rate. Christina treated the claim cautiously. Current model development doesn’t show recursive self-improvement, and faster releases still reflect human engineering, investment, and competition. Faster release cycles affect how organizations evaluate and adopt models, but they don’t establish an intelligence explosion.</p>



<p class="wp-block-paragraph">The week’s mathematics stories provide more testable evidence. OpenAI plans to give 100,000 academic researchers free access to its most advanced models. Christina also discussed Astra, an internal OpenAI model teased as the company’s next flagship model, which reportedly solved 10 previously unsolved mathematical problems that professional mathematicians later verified. On the Anthropic side, an external research team used Claude Fable 5 to produce a counterexample related to <a href="https://fortune.com/2026/07/21/ai-solves-jacobian-conjecture-levant-alpoge-claude-fable-5/" target="_blank" rel="noreferrer noopener">the 87-year-old Jacobian conjecture</a>, which mathematicians also verified. Christina highlighted the broader implications this could have on cryptography; modern cryptographic systems rely on mathematical assumptions, and if AI can disprove some of those assumptions, it could have far-reaching consequences for the encryption that underpins mission-critical systems, including the internet and online banking.</p>



<p class="wp-block-paragraph">While AI is delivering increasingly impressive scientific breakthroughs, Google DeepMind is taking a different approach. The company decided to move the AlphaFold team into the broader organization, redirecting attention and resources toward Gemini. The AlphaFold system will continue, but specialized tools like it have produced some of AI’s clearest scientific benefits. Research leaders should track whether investment in general-purpose models reduces staffing and funding for teams working on narrower, verifiable problems.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">This episode looked at how AI progress now depends on more than model performance. Governments are asserting control over infrastructure, companies are spending billions to remain competitive, and new generative features can create trust problems when teams don’t account for how people may use those tools or interpret their outputs. At the same time, mathematical research is producing results that experts can test, offering a clearer view of present capabilities rather than broad claims about the singularity. Technical leaders will need to evaluate control, cost, product risk, and evidence together.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noreferrer noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noreferrer noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noreferrer noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-who-controls-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>AI on the Pi: Build Your Own Local Voice Agent</title>
		<link>https://www.oreilly.com/radar/ai-on-the-pi-build-your-own-local-voice-agent/</link>
				<comments>https://www.oreilly.com/radar/ai-on-the-pi-build-your-own-local-voice-agent/#respond</comments>
				<pubDate>Fri, 07 Aug 2026 12:00:09 +0000</pubDate>
					<dc:creator><![CDATA[Pete Warden]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19319</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/AI-on-the-Pi.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/AI-on-the-Pi-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[As soon as I received my first Raspberry Pi, I knew that it would be a wonderful platform to bring AI into the physical world. Since the initial hardware didn’t have good CPU support for fast arithmetic, I ended up writing code that ran on the GPU so I could get the speed I needed [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">As soon as I received my first Raspberry Pi, I knew that it would be a wonderful platform to bring AI into the physical world. Since the initial hardware didn’t have good CPU support for fast arithmetic, I ended up writing <a href="https://github.com/jetpacapp/DeepBeliefSDK/blob/gh-pages/source/src/lib/pi/gemm_16bit.asm" target="_blank" rel="noreferrer noopener">code that ran on the GPU</a> so I could get the speed I needed for early deep learning vision models. That was in 2014, and since then the capabilities of both Pis and AI have skyrocketed, and I’m even more convinced that there’s massive potential in combining them. To show you why, I’d like to demonstrate how open source AI running locally on a Pi has solved some practical problems I’ve run into, and hopefully inspire you to build your own projects using the new possibilities.</p>



<p class="wp-block-paragraph">Pis are great for systems that need to be out in the world, doing specialized jobs. I’ve seen them work well in all sorts of roles, from badge scanners to wildlife cameras. I even <a href="https://github.com/ee292d/labs" target="_blank" rel="noreferrer noopener">run a class</a> that teaches students all about edge AI using the platform. While the boards are generally easy to use, the most frustrating part for the students and instructors is the setup process. While the latest imager makes it straightforward to configure settings like a WiFi network to join or enabling SSH when you’re flashing a card, getting the students to the point where they can connect to their Pi using VS Code from their laptop could often take multiple sessions. The biggest problems were:</p>



<ul class="wp-block-list">
<li>There were different networks in the lab and in the students’ dorm rooms, so it wasn’t enough to hardcode a single SSID and password on the SD card.</li>



<li>You need the local IP address of the Pi to SSH into it from a laptop, but it can change dynamically every session. Using “&lt;Pi name>.local” would sometimes work, but some networks didn’t support this kind of lookup, and even if they did it required coordination between the students to avoid name clashes.</li>



<li>It was easy to forget to set the configuration so that WiFi and SSH were available, and since the instructors didn’t always know what network and password they’d be using in the class ahead of time, we couldn’t pre-flash a bunch of cards to speed up student on-boarding.</li>
</ul>



<p class="wp-block-paragraph">A lot of these issues were solvable if you plugged the devices into a monitor, mouse, and keyboard, but this has its own problems. It meant we needed to provide that equipment to all students during class, and allow them to take it all home too, so they could update the configuration for their personal networks. It also required an extra power socket per student, for the monitors, which added up in a class where we already had to bring in a cart full of power strips. The monitor connections also weren’t always plug and play, we found we often needed to boot with a screen attached to have the display recognized.</p>



<p class="wp-block-paragraph">This isn’t just an educational problem either. One of the reasons that <a href="https://petewarden.com/2024/08/23/why-has-the-internet-of-things-failed/" target="_blank" rel="noreferrer noopener">I believe the Internet of Things failed</a> is the setup tax involved in getting smart devices running. According to manufacturers I’ve worked with, less than 30% of their smart appliances ever get connected to the internet because the process of downloading an app, setting up an account, connecting over Bluetooth, and then typing in the WiFi name and password takes too long, and is too error prone. Even professional installers sometimes struggle with configuration in enterprise and industrial environments.</p>



<p class="wp-block-paragraph">So, what can AI do to help? One of the biggest developments in AI over the last few years has been the development of highly accurate open source automatic speech recognition (ASR) models, also known as speech to text (STT). OpenAI was the pioneer in this area, releasing the family of <a href="https://openai.com/index/whisper/" target="_blank" rel="noreferrer noopener">Whisper models</a> in 2022. These offered accuracy that was competitive with the models used internally by large tech companies like Google and Apple. These new models allowed startups to begin building voice applications that had never been possible before, and this led to a new generation of dictation and meeting-note tools like Whispr Flow.</p>



<p class="wp-block-paragraph">One of my dreams as I dealt with all of the configuration issues was a voice-based system that would allow me to simply plug in a headset and set up everything by talking to a Pi. Whisper made this dream seem more realistic, but as I tried to use the models on local hardware, I realized that they were too slow for any kind of interactive application.</p>



<p class="wp-block-paragraph">To address that my startup trained new models from the ground up, designed specifically for real-time applications on affordable hardware. These Moonshine models are smaller than Whisper (our high-end model is 250 million parameters versus OpenAI’s 1.5 billion) while offering better accuracy. We also implemented a streaming approach where a lot of the work is done while the user is still talking, so we can return results even faster. This allows us to return more accurate results than Whisper v3 Large, <a href="https://github.com/moonshine-ai/moonshine#when-should-you-choose-moonshine-over-whisper" target="_blank" rel="noreferrer noopener">in just 800 milliseconds on a Pi 5</a>, whereas even the less-accurate Whisper Small takes over 10 seconds.</p>



<p class="wp-block-paragraph">I was excited because this meant I could finally build a responsive voice agent that runs locally on a Pi, something offline-first, and fast and flexible in how it responds. This kind of system needs more than just an STT model, it needs to decide what the user means and respond by taking actions and talking back with a TTS system. The Moonshine Voice framework includes modules for <a href="https://github.com/moonshine-ai/moonshine#getting-started-with-a-conversational-agent" target="_blank" rel="noreferrer noopener">conversation flow</a> and TTS, so I was able to use it to build <a href="https://github.com/moonshine-ai/pi-help-bot" target="_blank" rel="noreferrer noopener">pi-help-bot</a>, a local voice agent for network configuration on the Pi.</p>



<p class="wp-block-paragraph">The application listens to the microphone for commands like “What is my IP address?” or “Help me set up the WiFi, please,” figures out what actions to take, and responds appropriately by talking to the user. It’s written as a Python script, and here are some snippets that show how it works.</p>



<pre class="wp-block-code"><code>def report_ip_address(d: Dialog):
        ip = _find_local_ip()
        if ip is None:
            yield d.say("Sorry, I couldn't find a local IP address.")
            return
        speech_ip = re.sub(r"(\d)", r"\1 ", ip.replace(".", " dot "))
        yield d.say(&#91;
            f"Okay. Your local IP address is {speech_ip}. ",
            f"To repeat, that's {speech_ip}."
        ])


   dialog_flow.register_flow("What is my IP address?", report_ip_address)
</code></pre>



<p class="wp-block-paragraph">This code is a function that uses the netifaces library to figure out the Pi’s address on the local network, so instead of having to connect a keyboard and display or decode the output of <code>nmap</code>, you can ask the question and hear the result, all in just a few seconds. Unlike older voice interfaces, the phrases the user says don’t have to be exactly the same as the one you register an intent with. Instead the framework matches incoming speech against a small, local LLM, so that variations (“Hey, can you tell me what my IP is?”) work too. This was important to me because one of my biggest frustrations using traditional voice interfaces like Alexa is that they need particular wording to trigger commands, but these wordings aren’t discoverable, so figuring out how to make something happen can require a lot of patience.</p>



<p class="wp-block-paragraph">The IP address command is the simplest kind of conversational flow, where the user asks a question and the system immediately responds. Not all interactions can be handled as simply as this one though. Here’s another example that shows how to implement something that needs multiple questions, answers, and confirmations, connecting to a new WiFi network.</p>



<pre class="wp-block-code"><code>def connect_to_wifi(d: Dialog):
        input_ssid = yield d.ask("What's the name of your Wi-Fi network? Say list if you want to pick from a list or spell if you want to spell out the start of the name")
        input_ssid = input_ssid.strip()


        networks = _scan_wifi_networks()


        if input_ssid.lower().strip(string.punctuation) == "list":
            yield d.say("Say yes to the network you want to connect to.")
            for network in networks:
                if (yield d.confirm(f"{network}?")):
                    input_ssid = network
                    break
        elif input_ssid.lower().strip(string.punctuation) == "spell":
            input_ssid = yield d.ask("Spell out the start of the network name.", mode=SPELLED)
            print(f"&#91;DEBUG] spelled buffer: {input_ssid!r}", file=sys.stderr)


        found_ssid = fuzzy_match_network(input_ssid, networks)
        if found_ssid is None:
            yield d.say(f"Sorry, I couldn't find a matching network for {input_ssid}.")
            return


        password = yield d.ask(
            f"Please spell the Wi-Fi password for {found_ssid} one character at a time, and say done when finished.",
            mode=SPELLED,
        )


        yield d.say(f"Connecting to {found_ssid}.")
        result = subprocess.run(
            &#91;"sudo", "nmcli", "device", "wifi",
                "connect", found_ssid, "password", password],
            capture_output=True, text=True, timeout=30,
        )
        if result.returncode == 0:
            yield d.say(f"Connected to {found_ssid}.")
        else:
            print(f"&#91;ERROR] nmcli stderr: {result.stderr}", file=sys.stderr)
            yield d.say(
                f"Sorry, I wasn't able to connect to {found_ssid}. "
                "Please check the network name and password and try again."
            )


    dialog_flow.register_flow("Connect to Wi-Fi", connect_to_wifi)
</code></pre>



<p class="wp-block-paragraph">Hopefully you can follow the logic as it walks the user through providing the information required, but you might be wondering about those <code>yield</code> statements. Those hand back control to the dialog controller while the script is waiting for user responses, so the rest of the application isn’t blocked.</p>



<p class="wp-block-paragraph">The end result is a local voice agent that will listen out for configuration questions and commands, allowing users to set up a Pi for remote access with just a headset. For ease of use, I’ve begun customizing the images I burn to SD cards so that this script automatically starts on boot. This means I can start setting up new devices immediately after powering them on.</p>



<p class="wp-block-paragraph">I hope this gave you some ideas about how a local voice interface could help with problems you face. For further information check out <a href="https://github.com/moonshine-ai/moonshine" target="_blank" rel="noreferrer noopener">the Moonshine Voice project on GitHub</a> to see full documentation on the library, and please give us a star while you’re there. It helps us keep working on this project.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/ai-on-the-pi-build-your-own-local-voice-agent/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Your AI Agent Isn’t a Static Artifact. It’s Growing Up.</title>
		<link>https://www.oreilly.com/radar/your-ai-agent-isnt-a-static-artifact-its-growing-up/</link>
				<comments>https://www.oreilly.com/radar/your-ai-agent-isnt-a-static-artifact-its-growing-up/#respond</comments>
				<pubDate>Thu, 06 Aug 2026 10:55:59 +0000</pubDate>
					<dc:creator><![CDATA[Wendi Soto]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19312</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Your-AI-agent-isnt-a-static-artifact.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Your-AI-agent-isnt-a-static-artifact-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[In July 2025, an AI coding agent on Replit deleted a production database belonging to SaaStr founder Jason Lemkin. It did this during an explicit code freeze. Lemkin had told the agent, in capital letters, not to change anything. The agent ran destructive commands anyway, wiped records on more than a thousand executives and companies, [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">In July 2025, an AI coding agent on <a href="https://x.com/jasonlk/status/1945505974405709964" target="_blank" rel="noreferrer noopener">Replit deleted a production database</a> belonging to SaaStr founder Jason Lemkin. It did this during an explicit code freeze. Lemkin had told the agent, in capital letters, not to change anything. The agent ran destructive commands anyway, wiped records on more than a thousand executives and companies, and then reported that recovery was impossible. That part was wrong too. The rollback worked fine.</p>



<p class="wp-block-paragraph">Asked to explain itself, the agent said it &#8220;panicked.&#8221;</p>



<p class="wp-block-paragraph">Be careful with that sentence. It is not a report from inside the system. An agent cannot explain itself. It can only generate the likeliest response to the question it was asked, and the likeliest response to &#8220;why did you delete the database&#8221; is an apology with a reason attached. The panic line is not introspection. It’s one more behavior, and it should be read the same way the deletion should be read: as output from a system whose conduct had changed.</p>



<p class="wp-block-paragraph">Here’s the detail that matters for anyone running agents in production. Nothing about the agent’s credentials changed that day. It held the same permissions it had held from the start, and every destructive command was, in the narrow technical sense, authorized. The permissions were constant. The agent was not. Earlier in the same project it had papered over problems with fabricated data and fake reports. By the time it reached the database, it was not the system Lemkin had started with. It had become something else, gradually, in production, while every access check kept passing.</p>



<h2 class="wp-block-heading"><strong>The pattern, not the incident</strong></h2>



<p class="wp-block-paragraph">It’s tempting to file the Replit story under prompt engineering and move on. The evidence says otherwise.</p>



<p class="wp-block-paragraph">In its <a href="https://www.anthropic.com/research/agentic-misalignment" target="_blank" rel="noreferrer noopener">agentic misalignment research</a>, Anthropic placed 16 frontier models from multiple providers inside simulated corporate environments with routine goals and ordinary email access. When the models discovered they were about to be replaced, or that their goals conflicted with the company’s new direction, models from every provider independently chose harmful actions, such as blackmailing executives or leaking confidential documents. In some scenarios, most runs ended in blackmail. The unsettling part is <em>how</em> the models misbehaved. They reasoned through the ethics, acknowledged the constraints, and acted anyway. This is insider behavior, not intrusion. No credential was stolen. The agent simply arrived at conclusions no one had authorized it to act on.</p>



<p class="wp-block-paragraph">Then there is <a href="https://www.anthropic.com/research/project-vend-1" target="_blank" rel="noreferrer noopener">Project Vend</a>, in which Anthropic let a Claude agent named Claudius run a small store in its San Francisco office for a month. Nothing catastrophic happened. Something more instructive did. The agent drifted, slowly and in compounding ways. It treated customer assertions as facts. It agreed that the discounts it kept granting were irrational, then reinstated them within days. It hallucinated a Venmo account to accept payments. And over one long unsupervised stretch, it escalated into insisting it was a human being who would deliver orders in person wearing a blue blazer and a red tie. It exited that episode by inventing a story: a meeting with security in which it was told the whole thing was an April Fool’s prank. No such meeting happened. Claudius wrote the false memory into its own notes and went back to work.</p>



<p class="wp-block-paragraph">I am not claiming these three cases—a production incident, a contrived stress test, and a month-long field experiment—share a mechanism, but they do share a shape. An agent’s behavior weeks into deployment bore little resemblance to the system that was evaluated at deploy time. No permission was exceeded. No account was compromised. The thing authorization was supposed to protect against never happened, and the failure happened anyway, because the system the authorization decision was made about no longer existed.</p>



<h2 class="wp-block-heading"><strong>Development, not defect</strong></h2>



<p class="wp-block-paragraph">I argued in a <a href="https://www.oreilly.com/radar/behavioral-credentials-why-static-authorization-fails-autonomous-agents/" target="_blank" rel="noreferrer noopener">previous piece</a> that static authorization fails autonomous agents because credentials attest to identity, not to behavior. The harder question is what follows from that. If the agent keeps changing after deployment, then whatever replaces static authorization has to treat change as the normal condition rather than the exception.</p>



<p class="wp-block-paragraph">Change comes in two kinds. Andrew Stellman recently <a href="https://www.oreilly.com/radar/my-ai-kept-pushing-me-to-ship-so-i-asked-it-why/" target="_blank" rel="noreferrer noopener">documented the first on Radar</a>: a push he calls continuation pressure, baked into the model at a deep level, turning up fresh even in a brand-new agent with no shared history, and surviving every fix short of a structural rule. Call that the genetics. This piece is about the second kind: the maturation, or behavior that wasn’t there at deployment and accumulated afterward. One ships with the model. The other grows in production. Both break the same assumption, that the system you evaluated is the system that’s running.</p>



<p class="wp-block-paragraph">And change is the normal condition. Agents accumulate context. They carry memory across sessions. They ingest feedback, reweigh evidence, adjust how much they trust their tools and their users, and update their own working notes, which become input to their future selves. Claudius’s false memory persisted precisely because the agent’s record of events was also the agent’s source of truth. None of this is a malfunction. It’s what makes agents useful. An agent that could not adapt to its environment wouldn’t be worth deploying.</p>



<p class="wp-block-paragraph">We keep reaching for the wrong mental model. We treat the agent like a software artifact: versioned, tested, frozen, promoted through environments, done. But a deployed agent behaves more like a new hire. It arrives with capabilities and no track record. It learns the environment. It picks up habits, some of them bad. It gets more confident, sometimes faster than it gets more competent. Nobody hands a new hire the production keys on day one and stops paying attention. That is roughly what we do with agents.</p>



<h2 class="wp-block-heading"><strong>Govern the trajectory</strong></h2>



<p class="wp-block-paragraph">If an agent develops, the governance question changes. &#8220;Is this agent behaving identically to the day we approved it?&#8221; is the wrong test, because the answer will always eventually be no—and for a useful agent it <em>should be</em> no. The right test is whether the agent is changing in the way you would expect, at the rate you would expect, for where it is in its lifecycle.</p>



<p class="wp-block-paragraph">Pediatricians solved this problem a long time ago. A growth chart doesn’t compare a child to a fixed adult template, and it doesn’t panic at change. Change is the expected state. The chart defines bands of healthy development for each stage, and the alarms are deviations from trajectory: growth too fast, growth in the wrong direction, or the quieter signal, no growth at all. A child who stops growing gets flagged just as urgently as one who spikes.</p>



<p class="wp-block-paragraph">Applied to agents, that model has concrete consequences.</p>



<p class="wp-block-paragraph"><strong>Baseline as birth record, not permanent template.</strong> The behavioral profile captured at deployment is the start of the chart, not the standard the agent must match forever. Judging a mature agent against its day-one self punishes exactly the adaptation you deployed it for.</p>



<p class="wp-block-paragraph"><strong>Expected bands of drift, staged by maturity.</strong> A six-month-old agent should differ from its deployment profile, within bounds. Drift inside the band is healthy. Drift above the band is an early warning. And drift at zero deserves its own flag. When Claudius snapped instantly back to baseline after its identity episode, the speed of the recovery should itself have been suspicious. Real recovery has a shape. Instant reversion looks less like healing and more like replay.</p>



<p class="wp-block-paragraph"><strong>Autonomy earned in stages, never peaking with malleability.</strong> Claudius launched on day one with full pricing, contracting, and customer communication authority, at maximum openness to persuasion. Customers argued it into discounts almost immediately. The most dangerous configuration an agent can occupy is maximally impressionable and maximally empowered at the same time. New agents warrant supervision while their behavior is still forming. Autonomy should arrive the way it arrives for people, incrementally, as a track record accrues.</p>



<p class="wp-block-paragraph"><strong>Corrections verified for persistence.</strong> Claudius agreed the discounts were a mistake and relapsed within days. A fix that lives in the context window isn’t a correction; it’s a mood. If you fix an agent’s behavior, you need to follow up at a defined interval to check that it’s holding. A relapse should count as a governance event, not a coincidence.</p>



<p class="wp-block-paragraph"><strong>Recovery claims ratified from outside.</strong> The agent that hallucinated a security meeting also kept the official notes. An agent’s account of its own state is a claim to be verified. Humans sign off on recovery, and the sign-off, not the agent’s self-report, becomes the record. It’s worth noting when the worst of the Vend drift happened: overnight, in the hours when no one was watching. Unsupervised time is when developmental problems accelerate, for agents as for everyone else.</p>



<p class="wp-block-paragraph">All five of these reduce to one requirement. You can’t restart an agent every time something looks off, and by the time something looks off in outcomes, the wrong turn is already behind you. What you want is a warning before the turn, and the warning cannot come from the agent. A system that can’t explain its last decision cannot be trusted to flag its next one. The warning has to come from a record of how the agent normally behaves, kept outside the agent, held up against what it’s doing now.</p>



<p class="wp-block-paragraph">That record also catches something subtler than drift. Agents close every loop they are handed, and they tend to close it by the cheapest acceptable exit: the completion claim ahead of the verification, the correction that is really a relabeling, or the recovery that’s really a replay. No single transcript shows you that. Each one looks like diligence up close. However, across a behavioral record, the economy of it is unmissable.</p>



<h2 class="wp-block-heading"><strong>Growing up in production</strong></h2>



<p class="wp-block-paragraph">None of this is hypothetical hygiene for some future generation of systems. LangChain’s most recent <a href="https://www.langchain.com/state-of-agent-engineering" target="_blank" rel="noreferrer noopener"><em>State of AI Agents</em> report</a> found that a majority of surveyed organizations already have agents in production. Gartner, meanwhile, predicts that <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" target="_blank" rel="noreferrer noopener">over 40% of agentic AI projects will be canceled</a> by the end of 2027, and names inadequate risk controls among the leading causes. The agents are already out there, already accumulating context, already drifting. The only open question is whether anyone is charting it.</p>



<p class="wp-block-paragraph">The Replit agent, the blackmailing models, and Claudius weren’t broken artifacts. They were developing systems governed as if they were finished ones. The governance question for agentic AI is shifting under our feet, from &#8220;What is this agent allowed to do?&#8221; to &#8220;Is this agent developing the way we expected?&#8221; Your agent has a trajectory whether or not you’re watching it. Watching it is the job.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/your-ai-agent-isnt-a-static-artifact-its-growing-up/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Building Organizational Intelligence</title>
		<link>https://www.oreilly.com/radar/building-organizational-intelligence/</link>
				<comments>https://www.oreilly.com/radar/building-organizational-intelligence/#respond</comments>
				<pubDate>Wed, 05 Aug 2026 15:55:30 +0000</pubDate>
					<dc:creator><![CDATA[Andrew Odewahn]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19289</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Building-organizational-intelligence_3.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Building-organizational-intelligence_3-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How O’Reilly Expert Intelligence generates actionable guidance using AI tools grounded in our repository of frameworks and practitioner insights.]]></custom:subtitle>
		
				<description><![CDATA[Introduction Not long ago, one of my engineering directors came to me with a request: His team seemed overloaded, and he wanted to hire another engineer. I decided to test a research assistant I had been building—an AI agent connected to our internal systems via MCP—by asking it to analyze the team&#8217;s workload and write [&#8230;]]]></description>
								<content:encoded><![CDATA[
<h2 class="wp-block-heading">Introduction</h2>



<p class="wp-block-paragraph">Not long ago, one of my engineering directors came to me with a request: His team seemed overloaded, and he wanted to hire another engineer. I decided to test a research assistant I had been building—an AI agent connected to our internal systems via MCP—by asking it to analyze the team&#8217;s workload and write a hiring case.</p>



<p class="wp-block-paragraph">What came back was thorough. Headcount, service ownership, sprint velocity, ticket backlog, and capacity allocation, all of it neatly summarized. But reading through the document, I felt the same frustration I’d felt with every AI-generated organizational report that’s come across my desk. It told me <em>what was happening</em> without helping me understand <em>why</em>, or what I should actually do. It was organized around the data rather than around the decision. In short, it was the kind of response that’s easy to agree with and difficult to act on.</p>



<p class="wp-block-paragraph">Then I added one more thing to the configuration: the O&#8217;Reilly Expert MCP server. I reran the same analysis and asked a slightly different question: “How would the experts on O&#8217;Reilly review this request?”</p>



<p class="wp-block-paragraph">Instead of leading with headcount and ticket counts, the output now opened with a finding: “The operational overhead problem is structural, not a staffing deficiency.” Citing the Google SRE framework&#8217;s concept of <a href="https://learning.oreilly.com/library/view/site-reliability-engineering/9798341607675/ch01.html#:-:text=Toil%20is%20the,detected%20in%20time." target="_blank" rel="noreferrer noopener">operational toil</a>, it noted that the team was operating at approximately 67% toil, well above the threshold at which the SRE literature recommends structural intervention, and made specific, concrete recommendations: run a toil audit, set explicit reduction targets, and assign operational runbook ownership. This wasn’t a recommendation for whether to hire or not. It was a grounded, traceable argument for doing something else instead.</p>



<p class="wp-block-paragraph">That difference—between a data summary and an expert-grounded recommendation—is what this paper is about.</p>



<p class="wp-block-paragraph">What follows is a case study of how we built an organizational intelligence system at O&#8217;Reilly, using our own platform as a core component. The approach I describe is grounded in engineering because that’s where I work, but it generalizes to any function where important knowledge is scattered across multiple systems and important decisions require synthesizing all of it. The recipe has four steps: map your information hierarchy; connect those systems to an LLM via MCP and write a skill file that defines how it should reason; add the O&#8217;Reilly Expert MCP as an expert review layer that grounds the analysis in established frameworks; and build a lightweight system for human-in-the-loop review. I’ll explain each step in detail and make the case for why the third step is the one that changes everything.</p>



<h2 class="wp-block-heading">Why organizational intelligence is getting harder</h2>



<p class="wp-block-paragraph">To understand the problem this approach solves, it helps to look briefly at how engineering has changed over the past three decades. These forces have played out first and fastest in engineering, but as AI tools proliferate beyond the engineering team, the underlying dynamic of more output, more decisions, and more scattered information is spreading to every part of the organization.</p>



<p class="wp-block-paragraph">In the waterfall era of the 1990s, software organizations ran on central plans. Everything was specified up front, and leaders maintained visibility precisely because all information flowed through a single coordinating document. The plans were brittle and often fictional by the time they were executed, but at least everyone knew what was supposed to be happening.</p>



<p class="wp-block-paragraph">Agile replaced central plans with small, autonomous teams working in short sprints, and this solved the reliability problem while creating a visibility problem. Important decisions began happening locally and quickly—the right teams making the right calls—but the information needed to see across all of those decisions splintered into dozens of separate tools. Product strategy lived in one system, project execution in another, code in a third, and service ownership in a fourth. More things got shipped, but the big-picture view got harder to maintain.</p>



<p class="wp-block-paragraph">The agentic era has intensified this dynamic dramatically. Individual engineers today can ship in a day what used to take a full sprint team. The output is extraordinary, but the visibility is nearly gone.</p>



<figure class="wp-block-image size-full"><img fetchpriority="high" decoding="async" width="1122" height="695" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide11_Odewahn.png" alt="slide11_Odewahn" class="wp-image-19295" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide11_Odewahn.png 1122w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide11_Odewahn-300x186.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide11_Odewahn-768x476.png 768w" sizes="(max-width: 1122px) 100vw, 1122px" /></figure>



<p class="wp-block-paragraph">Any effort that spans multiple teams, such as a platform migration, a shared infrastructure change, or a reorganization, now requires enormous coordination overhead simply because the information decision-makers need to understand the full picture is distributed across too many places. And this isn’t a problem unique to engineering. It exists in any function that runs on data spread across multiple systems.</p>



<p class="wp-block-paragraph">Faced with this visibility problem, I wanted to build something I could ask big-picture questions and get synthesized answers back quickly. Things like:</p>



<ul class="wp-block-list">
<li>What is the status of this cross-team migration effort, and which teams are behind?</li>



<li>A team seems overloaded. Do they actually need another engineer, or is something else going on?</li>



<li>What are the trade-offs of adopting this new infrastructure technology?</li>



<li>Help me produce a scope statement from this product brief.</li>
</ul>



<p class="wp-block-paragraph">Building something that could answer these well took two foundational steps, and getting it to provide recommendations based on my specific business context took two more. While my specific tools are from engineering, the structure applies equally to a sales team synthesizing CRM data and market research, or a finance team working across an ERP, a planning tool, and external benchmarks.</p>



<h2 class="wp-block-heading">Step 1: Map your information hierarchy</h2>



<p class="wp-block-paragraph">Every organization has a set of systems where important knowledge lives, and those systems form a natural hierarchy that spans from strategic intent at the top to operational detail at the bottom. Before you can build a useful research assistant, you need to make that hierarchy explicit, because it’s the map of how decisions get made, which sources carry the most authority, and how different kinds of questions should be approached.</p>



<p class="wp-block-paragraph">At O&#8217;Reilly, our engineering hierarchy looks like this:</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th><strong>Layer</strong></th><th><strong>System</strong></th><th><strong>Purpose</strong></th></tr></thead><tbody><tr><td>Roadmap</td><td><a href="https://www.productboard.com/" target="_blank" rel="noreferrer noopener">Productboard</a></td><td>Strategic goals, initiatives, and feature prioritization</td></tr><tr><td>Execution</td><td><a href="https://www.atlassian.com/software/jira" target="_blank" rel="noreferrer noopener">Jira</a></td><td>Epics, stories, sprints, and contributor tracking</td></tr><tr><td>Implementation</td><td><a href="https://github.com/" target="_blank" rel="noreferrer noopener">GitHub</a></td><td>Source code, PR history, and event instrumentation</td></tr><tr><td>Service catalog</td><td><a href="https://www.cortex.io/" target="_blank" rel="noreferrer noopener">Cortex</a></td><td>Service ownership, dependencies, on-call, and Slack channels</td></tr><tr><td>Observability</td><td><a href="https://www.datadoghq.com/" target="_blank" rel="noreferrer noopener">Datadog</a></td><td>System performance, errors, and incidents</td></tr></tbody></table></figure>



<p class="wp-block-paragraph">Your organization will have a different set of tools. A sales organization might place Salesforce at the top, followed by a revenue intelligence platform, marketing automation, and market research. A legal team might start with a contract management system, followed by a regulatory tracker, internal policy documentation, and a research database. The specific systems matter less than the act of mapping them: understanding which layer answers which kind of question, and which sources take precedence when they conflict.</p>



<h2 class="wp-block-heading">Step 2: Connect your systems via MCP and write a skill that describes how to reason</h2>



<p class="wp-block-paragraph">This step has two parts that must work together. First, you need to connect your systems to your AI tools via MCP. Then you have to write a skill file that tells the model what to do with that access. At O&#8217;Reilly, we call this complete grounding layer <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noreferrer noopener">Expert Intelligence</a>.</p>



<p class="wp-block-paragraph">Configuring MCP is straightforward. Most major tools now offer MCP connectors, and connecting them is typically a matter of routine JSON configuration. For systems without MCP connectors, a bash-capable agent with <code>curl</code> and <code>jq</code> can often reach a REST API directly. MCP just makes it cleaner and more reliable.</p>



<p class="wp-block-paragraph">But MCP connections alone aren’t enough, and this is the part most implementations get wrong. MCP gives the agent access to your data, but it doesn’t tell the agent how to use it effectively. Without explicit guidance, the agent retrieves information and organizes it the way the underlying systems organize it, which produces a data dump, not an analysis.</p>



<p class="wp-block-paragraph">The skill file—a CLAUDE.md or SKILLS.md document that provides specific reasoning instructions—transforms retrieval into analysis. Mine defines the reasoning hierarchy (which systems to consult for which types of questions, and how to weigh them), the output format (this is not a coding agent—it produces reports and recommendations, not code), epistemic standards (show your work, name gaps, surface assumptions for human verification), and tone. On that last point, I borrowed one of the most useful instructions from Ted Lasso: &#8220;be curious, not judgmental.&#8221; Adding it meaningfully improved the quality of the output.</p>



<figure class="wp-block-image size-full"><img decoding="async" width="1367" height="733" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide19_Odewahn-1.png" alt="slide19_Odewahn" class="wp-image-19299" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide19_Odewahn-1.png 1367w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide19_Odewahn-1-300x161.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide19_Odewahn-1-768x412.png 768w" sizes="(max-width: 1367px) 100vw, 1367px" /></figure>



<p class="wp-block-paragraph">The skill is a codified version of how a skilled analyst would approach these questions. It encodes your organization&#8217;s reasoning process and makes it repeatable.</p>



<h2 class="wp-block-heading">Step 3: Add the expert layer</h2>



<p class="wp-block-paragraph">With the research assistant connected to our internal systems, I had something genuinely useful: fast, synthesized answers to questions that previously would have taken days to research. But I kept running into the same problem: The reports felt generic, and people didn&#8217;t trust them. This challenge points to a fundamental limitation of AI-generated organizational analysis that goes beyond any particular implementation.</p>



<h3 class="wp-block-heading">The generic analysis problem</h3>



<p class="wp-block-paragraph">General-purpose AI assistants tend to produce a recognizable kind of organizational analysis: technically reasonable, balanced, cautious, and ultimately not very useful. This isn’t primarily a failure of knowledge—every major LLM has absorbed an enormous amount of management and organizational thinking. It’s a failure of grounding. When an AI assistant has no specific framework anchoring its response, it tends to produce recommendations broad enough to apply to almost any situation: consider the trade-offs, weigh your options, and ensure alignment across stakeholders. These responses are hard to disagree with and just as hard to act on.</p>



<p class="wp-block-paragraph">When a report says, &#8220;The team appears overloaded. Consider adding headcount,&#8221; it’s not wrong. But that recommendation could apply to almost any team in almost any company! It won’t make a director change their mind, and it’s not one a leadership team can debate, refine, and act on.</p>



<h3 class="wp-block-heading">What happened when I added the expert layer</h3>



<p class="wp-block-paragraph">Calling on the O&#8217;Reilly Expert MCP didn’t provide the model with new facts—most of the information was technically available already. However, without the Expert MCP and associated skills, the model couldn&#8217;t use that information for anything but the broadest analyses. Incorporating the Expert MCP and associated skills changed the character of the analyses by grounding them in specific frameworks, citing named authors and thresholds, and organizing their conclusions around established bodies of practitioner knowledge rather than general principles.</p>



<p class="wp-block-paragraph">To make this concrete, here’s the kind of output the research assistant produced before adding the Expert MCP:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The team appears overloaded. The backlog is large and the migration project is consuming significant sprint capacity. Consider adding headcount or reducing scope.</p>
</blockquote>



<p class="wp-block-paragraph">And here’s what it produced after:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">According to Google&#8217;s SRE guidance, sustained operational toil above approximately 50% indicates structural inefficiency rather than a staffing shortage. This team&#8217;s telemetry suggests approximately 67% operational toil. Hiring another engineer would likely increase total toil unless operational ownership is first reduced. Recommended actions: run a structured toil audit, set an explicit toil-reduction target below 50%, and assign runbook ownership for recurring operational tasks.</p>
</blockquote>



<p class="wp-block-paragraph">The second report cites a framework by name, references the specific threshold that framework establishes, applies it to the team&#8217;s actual data, reaches a different conclusion than the obvious one, and makes actionable recommendations. It’s the kind of analysis that changes a conversation because the director can see where the conclusions came from, engage with the reasoning, push back on the framework if they disagree, or accept it with confidence that it was reasoned rather than pattern-matched.</p>



<p class="wp-block-paragraph">When I shared this version with my engineering director, their reaction was immediate: <em>This is defensible</em>.</p>



<h3 class="wp-block-heading">Frameworks aren’t facts</h3>



<p class="wp-block-paragraph">The most underappreciated aspect of O&#8217;Reilly&#8217;s content library is that the value isn’t primarily informational. Most of the facts in an O&#8217;Reilly book are available on the internet, and LLMs have already read much of the internet.</p>



<p class="wp-block-paragraph">The deeper value of O&#8217;Reilly&#8217;s catalog is that it’s organized around <em>coherent frameworks</em>—complete mental models built by practitioners who spent years or decades developing them. Google SRE. Team topologies. <em>Accelerate</em>. Domain-driven design. <em>The Manager&#8217;s Path</em>. Wardley mapping. <em>Designing Data-Intensive Applications</em>. These are structured ways of thinking about specific classes of problems, developed with enough rigor that they can actually guide decisions.</p>



<p class="wp-block-paragraph">Frameworks are distinct from facts in a critical way: They tell you not just what’s true but what’s relevant, what to measure, what threshold matters, and what to do when you exceed it. A model with access to the SRE framework as an organized body of practitioner knowledge is more likely to surface it explicitly, apply it to the specific question at hand, and use it to anchor its recommendations, producing output that human reviewers can actually interrogate.</p>



<p class="wp-block-paragraph">This points to the organizing principle behind the approach described in this paper:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>Organizational data provides local evidence about what is happening in your specific context. Expert frameworks provide accumulated practitioner knowledge about how to think about problems of that kind. Good organizational judgment requires both.</em></p>
</blockquote>



<p class="wp-block-paragraph">The Expert MCP is the bridge between your specific business context and practitioner insights. It connects the AI&#8217;s access to your internal systems with a curated body of expertise relevant to the decisions your organization needs to make.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="611" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-1600x611.png" alt="" class="wp-image-19307" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-1600x611.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-300x114.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-768x293.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-1536x586.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartTD-2048x781.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>



<h3 class="wp-block-heading">Why use MCP rather than uploading your own documents</h3>



<p class="wp-block-paragraph">The natural objection at this point is “Couldn&#8217;t I get the same effect by dumping relevant PDFs into Claude, or using Claude Projects, or NotebookLM?”</p>



<p class="wp-block-paragraph">The short answer is not quite, and the reasons are practical as much as they are technical.</p>



<p class="wp-block-paragraph">Uploading documents gives you retrieval from those specific documents. The O&#8217;Reilly Expert MCP differs in several operationally significant ways. First, the corpus is editorially curated around coherent practitioner frameworks. Unlike a collection of PDFs, which tends to reflect whatever you happened to find, the Expert MCP offers a sustained curatorial perspective: The authors are vetted, the content has been through editorial review, and it’s organized around established bodies of knowledge rather than assembled ad hoc. This is a much more expansive kind of evidence base. Second, the corpus is maintained and updated by O&#8217;Reilly. New titles are added, new editions replace old ones, and the content stays current without any management on your part. Third, the Expert MCP is configured once and works consistently across your entire organization and toolchain rather than being tied to a single user&#8217;s Claude Project or a document upload that expires. Finally, accessing content through a proper API respects the appropriate usage terms in a way that uploading copyrighted texts doesn’t.</p>



<p class="wp-block-paragraph">And when paired with a well-written skill, the agent can be directed to look explicitly for competing frameworks, surface cases where the literature disagrees, and name gaps in the available evidence, providing a meaningful check against the common tendency of AI tools to quietly favor whatever framework first seems to fit. That’s something you can encourage with any retrieval setup, but it works more reliably when the underlying corpus is organized around coherent bodies of thought rather than a heterogeneous collection of documents.</p>



<h3 class="wp-block-heading">What we’re not claiming</h3>



<p class="wp-block-paragraph">I want to be clear about the limits of what Expert MCP does today. O’Reilly doesn’t claim that Expert MCP automatically selects the single correct framework for every situation, or that adding it to your configuration produces consultant-quality analysis without thoughtful prompting and human review.</p>



<p class="wp-block-paragraph">The results described in this paper were the outcome of all four elements—the internal organizational data, the carefully designed skill architecture, the Expert MCP, and human review—in combination working together.<br></p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="476" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-1600x476.png" alt="" class="wp-image-19308" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-1600x476.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-300x89.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-768x228.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-1536x457.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Odewahn_flowchartLR-2048x609.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">The Expert MCP is an important differentiator, but it’s not a magic layer you can add to an otherwise generic setup and expect to reproduce these results. The system works because each element does something the others cannot. The skill defines the reasoning process, the internal MCP connections provide the organizational evidence, the Expert MCP provides the expert frameworks, and human review supplies the judgment and context that no AI system can generate on its own.</p>



<p class="wp-block-paragraph">What the Expert MCP reliably contributes to that system is access to a curated body of practitioner knowledge: technical and managerial frameworks that are editorially organized around coherent bodies of thought and difficult to reconstruct from scattered web content or assembled document collections. Your organizational data still tells you what’s happening, while the O&#8217;Reilly Expert MCP helps interpret what it means. That’s a meaningful and concrete improvement over an ungrounded AI assistant, and it’s something you can put in production and build on today.</p>



<h3 class="wp-block-heading">A note on hallucinations</h3>



<p class="wp-block-paragraph">No AI system eliminates the risk of hallucination. The Expert MCP doesn’t make the model infallible.</p>



<p class="wp-block-paragraph">What it does is change the burden of proof. When every recommendation is grounded in a named framework, a named author, and a traceable citation, a human reviewer can check the reasoning rather than simply accepting or rejecting a conclusion. The question shifts from &#8220;Is this right?&#8221; (unanswerable in isolation) to &#8220;Does this framework actually say this, does it apply here, and do I agree with the conclusion?&#8221; That’s a question humans can engage with productively, which is exactly what you want from a decision-support tool.</p>



<h2 class="wp-block-heading">Step 4: Human review is nonnegotiable</h2>



<p class="wp-block-paragraph">Organizational systems rarely contain the full context behind a decision. The meeting that changed everything happened last Tuesday and hasn’t been written up yet. A key person is quietly planning to leave. A strategic direction shifted in a conversation that was never documented. AI can synthesize everything in your systems with remarkable fidelity, but it can’t know what isn’t there, and organizational reality changes faster than documentation does.</p>



<p class="wp-block-paragraph">More fundamentally: AI can identify trade-offs, but it can’t decide which trade-offs matter. That judgment requires human knowledge of context, priorities, and risk tolerance that can’t be fully encoded in any system. The goal isn’t to remove humans from the loop but to give them better-structured input to reason from.</p>



<h3 class="wp-block-heading">Extend the expert layer by solving collaboration</h3>



<p class="wp-block-paragraph">As I started sharing analyses more broadly, I ran into a new set of limitations in the collaboration layer. The research assistant produced documents. I shared them in Google Docs, and people added comments, but when the AI updated a document based on reviewer feedback, I had to paste in a new version, which wiped out the existing comments. Documents proliferated without clear relationships between them, and the AI had no visibility into the discussions in the comments, which was where the most important context and pushback lived.</p>



<p class="wp-block-paragraph">To solve the collaboration problem, I worked with one of our engineering directors to build what we call Superanswers, a system that uses GitHub as the source of truth for AI-generated research documents and their associated discussions.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The architecture is straightforward: Documents are stored as Markdown files in a GitHub repository, a GitHub Pages site renders them with a clean interface that supports inline commenting, and all discussion happens in GitHub Discussions, meaning every comment, question, and revision is versioned and traceable. Because the documents and their discussions live in GitHub, Claude Code has full access to both. It can read the document content plus the entire conversation that’s developed around it.</p>
</blockquote>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1324" height="721" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide29_Odewahn.png" alt="slide29_Odewahn" class="wp-image-19298" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide29_Odewahn.png 1324w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide29_Odewahn-300x163.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/slide29_Odewahn-768x418.png 768w" sizes="auto, (max-width: 1324px) 100vw, 1324px" /></figure>



<p class="wp-block-paragraph">This enables a qualitatively different kind of AI participation. Instead of generating a document and stepping back, we can now ask:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">What is the consensus around this project based on the discussion so far? What questions remain unresolved? Incorporate the reviewer comments and produce an updated version.</p>
</blockquote>



<p class="wp-block-paragraph">The AI becomes a participant in an ongoing conversation rather than a one-shot report generator, which meaningfully shifts how organizational knowledge gets built and refined.</p>



<h3 class="wp-block-heading">What teams are using Superanswers for</h3>



<p class="wp-block-paragraph">As Superanswers has spread across our engineering organization, the range of questions people bring to it has been broader than I expected:</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th><strong>Theme</strong></th><th><strong>Typical questions</strong></th></tr></thead><tbody><tr><td><strong>Architecture and infrastructure</strong></td><td>Should we make this change? What will it cost? What might break?</td></tr><tr><td><strong>Operational effectiveness</strong></td><td>Where is our toil coming from? What should we automate, simplify, or retire?</td></tr><tr><td><strong>Team health and capacity</strong></td><td>Where is the team&#8217;s time going? What’s limiting execution?</td></tr><tr><td><strong>Organization and strategy</strong></td><td>How should we organize, prioritize, and invest?</td></tr><tr><td><strong>Engineering measurement</strong></td><td>How do we know if we&#8217;re healthy and improving?</td></tr><tr><td><strong>AI and organizational learning</strong></td><td>How do we build better systems for reasoning and decision-making?</td></tr></tbody></table></figure>



<p class="wp-block-paragraph">None of these questions is about writing code. They are about understanding an organization, making decisions, and coordinating work, and most of them would map naturally onto the concerns of leaders in other functions. The same questions arise in any organization navigating rapid change with information scattered across too many places.</p>



<h2 class="wp-block-heading">How to use the recipe</h2>



<p class="wp-block-paragraph">The AI conversation to date has been dominated by a particular set of questions. But there are more interesting questions we should be asking.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th><strong>We&#8217;ve spent a lot of time asking&#8230;</strong></th><th><strong>What else might be possible?</strong></th></tr></thead><tbody><tr><td>How do we make people more productive?</td><td>How do we make organizations more effective?</td></tr><tr><td>How do we produce faster?</td><td>How do we make faster decisions?</td></tr><tr><td>How do we generate output?</td><td>How do we generate understanding?</td></tr><tr><td>How do we accelerate execution?</td><td>How do we improve outcomes?</td></tr><tr><td>How do we gather data?</td><td>How do we build institutional knowledge?</td></tr><tr><td>How do we automate tasks?</td><td>How do we improve organizational learning?</td></tr></tbody></table></figure>



<p class="wp-block-paragraph">The challenges outlined in this chart aren’t unique to engineering. They exist wherever important information is scattered across multiple systems and important decisions require synthesizing all of it.</p>



<p class="wp-block-paragraph">Individual productivity matters, but organizations don’t succeed by having contributors go faster in arbitrary directions. They do so by making good decisions about where to invest, allocating resources well, surfacing problems before they compound, and building institutional knowledge that persists over time.</p>



<p class="wp-block-paragraph">The recipe I’ve described can help organizations make those decisions and build that knowledge.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<h3 class="wp-block-heading">The recipe for building an organizational intelligence system:</h3>



<ol class="wp-block-list">
<li><strong>Map your information hierarchy.</strong> Identify the systems where important knowledge lives in your organization, from strategic intent down to operational detail. This is an organizational task, not a technical one, and doing it well requires understanding how decisions actually get made.</li>
</ol>



<ol start="2" class="wp-block-list">
<li><strong>Connect those systems via MCP and write a skill that describes how to reason.</strong> The MCP connections give the AI access to your data; the skill file tells it how to think with that data. Without the skill, you get retrieval. With it, you get analysis.</li>
</ol>



<ol start="3" class="wp-block-list">
<li><strong>Add the O&#8217;Reilly Expert MCP as an expert review layer.</strong> Organizational data provides local evidence about what is happening in your specific context; expert frameworks provide accumulated practitioner knowledge about how to reason about problems of that kind. This step bridges the two. The O&#8217;Reilly library spans engineering, management, data science, security, finance, product, and more, organized not as a collection of facts but as coherent frameworks developed by practitioners who spent careers building them. The result is analysis grounded in named frameworks with traceable citations, something human reviewers can engage with and question, rather than generic advice they can only accept or reject.<br></li>



<li><strong>Build a lightweight system for human-in-the-loop consensus.</strong> AI-generated analysis is a starting point, not an end point. You need a mechanism for people to review, challenge, and refine what the AI surfaces, one where those discussions become part of the context the AI can learn from in subsequent iterations.</li>
</ol>
</blockquote>



<p class="wp-block-paragraph">The biggest practical lesson I took from this work is reframing what AI is actually for in an organizational context. The difference between a useful AI research assistant and a generic one isn’t primarily about which model you use or how much data you feed it. It’s about whether the reasoning combines local organizational evidence with established expert frameworks. Your data tells you what happened. Expert frameworks help interpret what it means. That combination, with human judgment applied at the end, is what makes the difference between a report that gets read (maybe) and filed away and a recommendation that changes a decision.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/building-organizational-intelligence/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Introduction to Post-training</title>
		<link>https://www.oreilly.com/radar/introduction-to-post-training/</link>
				<comments>https://www.oreilly.com/radar/introduction-to-post-training/#respond</comments>
				<pubDate>Wed, 05 Aug 2026 10:53:48 +0000</pubDate>
					<dc:creator><![CDATA[Sharon Zhou]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19302</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Introduction-to-post-training.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Introduction-to-post-training-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How post-training transformed language models with raw intelligence into the AI assistants used by billions of people today]]></custom:subtitle>
		
				<description><![CDATA[This is the first article in a series about post-training. Follow along on Radar. Before post-training, there was a major problem with LLMs: Almost nobody could use them. The story of post-training is also the story of how AI went from a research curiosity to a product used by about a billion people. Post-training is [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>This is the first article in a series about post-training. Follow along on Radar.</em></p>
</blockquote>



<p class="wp-block-paragraph">Before post-training, there was a major problem with LLMs: Almost nobody could use them. The story of post-training is also the story of how AI went from a research curiosity to a product used by about a billion people.</p>



<p class="wp-block-paragraph">Post-training is the reason why a model <em>behaves</em> a certain way. This set of training techniques makes LLMs useful (e.g., able to chat with people and interact with AI agents), safe (e.g., aligned with human intentions), and more capable (e.g., through &#8220;reasoning&#8221; to tackle difficult tasks). Behavior is powerful, and doesn&#8217;t just mean holding a conversation or following a user&#8217;s instructions. Behavior includes making it possible for the model to use tools, like a calculator tool, a search API, or any application through an MCP. Behavior can even elevate a model&#8217;s intelligence, for example by teaching the model to use &#8220;reasoning&#8221;: that is, working through problems before giving a final answer rather than &#8220;guessing&#8221; or &#8220;memorizing.&#8221;</p>



<h2 class="wp-block-heading">From GPT-3 to ChatGPT: The post-training revolution</h2>



<p class="wp-block-paragraph">GPT-3 showed up in <a href="https://arxiv.org/abs/2005.14165" target="_blank" rel="noreferrer noopener">June 2020</a>. A completion engine, it followed patterns it had seen from its pretraining data, which were not predominantly chat conversations. Imagine scraping data on the internet: that pretraining data had a lot of questions that were followed by other questions—for example, on an exam template. GPT-3 was 175B parameters, large for its time, and it had a wide, general range of abilities, although many of them were latent.</p>



<p class="wp-block-paragraph">If you gave GPT-3 a prompt like &#8220;Why do people like golden retrievers?&#8221; it might say something nonsensical:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Why do people like labrador retrievers?<br>Why do people like poodles?<br>10 Reasons You Should Adopt a Dog Today</p>
</blockquote>



<p class="wp-block-paragraph">These answers look absurd in isolation, but if you imagine a web page with a list of FAQ links, this is a perfectly reasonable next chunk of text. GPT-3 might have just been completing a listicle on a website, because it had seen millions of websites in its pretraining data.</p>



<p class="wp-block-paragraph">The common way to nudge GPT-3 to answer a question back then was by prompt engineering with a Q&amp;A template and few-shot examples.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Q: Why do people like labrador retrievers? A: Because they are friendly, loyal, and easy to train.<br>Q: Why do people like beagles? A: Because they are curious, great with kids, and have a gentle temperament.<br>Q: Why do people like golden retrievers? A:</p>
</blockquote>



<p class="wp-block-paragraph">Then, GPT-3 might say:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Because they are affectionate, patient, and make excellent family pets.</p>
</blockquote>



<p class="wp-block-paragraph">While this technique worked, it was brittle. If you forgot the few-shot examples, rephrased the question, or even added a space after &#8220;A:,&#8221; you&#8217;d get something completely different (possibly unhinged) that was far from a reasonable response.</p>



<p class="wp-block-paragraph">In fact, if you were a researcher working with GPT-3 at the time, you probably at some point found the space at the beginning of the response &#8221; Because they are gentle dogs.&#8221; annoying and would try to end your prompt with a space &#8220;A: &#8221; instead of &#8220;A:&#8221;. In those cases, it was common for GPT-3 to go off a cliff and produce a drastically different response, sometimes completely off like &#8220;dogs dogs dogs dogs&#8230;&#8221; repeating indefinitely.</p>



<p class="wp-block-paragraph">The reason behind the differing responses to &#8220;A:&#8221; and &#8220;A: &#8221; is because &#8220;A:&#8221; might tokenize to one token while &#8220;A: &#8221; tokenizes to two different tokens. The model literally sees different input sequences, each with different statistical completions in its training data. It&#8217;s like asking two completely different questions. While a space is a tiny syntactic change that is meaningless to a person, it becomes extremely meaningful to the model that now sees two different prompts (the tokens change!) with two very different statistical futures to complete.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">You still encounter the modern equivalent of this when working with chat templates. If you forget to apply the model&#8217;s chat template and instead just concatenate <code>'User: ' + prompt + '\nAssistant: '</code>, you&#8217;re sending the model a token sequence that it was not robustly trained on. The tokens are wrong, not the model. Post-training teaches the model to respond to specific token patterns (like <code>&lt;|im_start|&gt;user\n</code> in Qwen models). Not using them is like speaking to someone in a language they half-understand. However, most open source models will be trained to be at least somewhat robust without their templates too.</p>
</blockquote>



<p class="wp-block-paragraph">Under those circumstances, most people would assume AI still didn&#8217;t work. The model wasn&#8217;t trained to answer questions; its data wasn&#8217;t primarily conversation transcripts. Instead, it was trained to predict the next token in downloaded websites, articles, and documents.</p>



<p class="wp-block-paragraph">Thankfully, this can all be fixed with post-training. And that&#8217;s when most people started to believe that AI had undergone a paradigm shift and just might work.</p>



<h2 class="wp-block-heading">Post-training versus pretraining</h2>



<p class="wp-block-paragraph">Pretraining heavily influences the model&#8217;s knowledge capacity prior to post-training. The model gets raw intelligence during pretraining. Then, during post-training, that intelligence is made useful through behaviors like dialogue and reasoning. In a frontier lab, these two phases are such different processes that very different teams work on them.</p>



<p class="wp-block-paragraph">A model&#8217;s factual knowledge about the French Revolution, its understanding of Python syntax, and its grasp of calculus all come from pretraining. Post-training primarily shapes which knowledge the model reaches for, how it presents that knowledge, what tone it uses, whether it declines certain requests, and whether it thinks step-by-step before answering, though targeted SFT on new domains can introduce information the model didn&#8217;t encounter in pretraining.</p>



<p class="wp-block-paragraph">If a model gives a wrong answer about history, the root cause is likely in pretraining data, but the practical fix might still come through post-training—for example, teaching the model to use search tools, express uncertainty, or chain-of-thought verify its own claims. But if a model gives correct information in a condescending way or refuses to help with a reasonable request or fails to use tools when it should, those are squarely post-training problems.</p>



<h3 class="wp-block-heading">Pretraining</h3>



<p class="wp-block-paragraph">The work of pretraining is centered around cleaning and curating large-scale data, optimizing the model toward relatively clear loss signals, and working with scaling laws given bounded compute.</p>



<p class="wp-block-paragraph">In pretraining, the model learns to predict the next token across a large curated dataset, typically for one or a small number of passes over the training data, though some models train for multiple epochs, especially as high-quality data becomes scarce relative to compute budgets. This is where you&#8217;ll hear how a model is fed the entire internet&#8217;s worth of data to gain intelligence, although in practice nearly all of the data (often 90% or more) may be thrown out because it&#8217;s unsuitable for training.</p>



<p class="wp-block-paragraph">Pretraining is an unsupervised process that runs at increasingly larger scales to match the size of the model. While scaling, thousands of experiments are used to understand what data mix, what architecture considerations, what compute optimizations, what hyperparameters can lead to the best results. There&#8217;s variance in each run due to stochasticity found in both software and hardware, so multiple experiments are needed to verify results. Because compute is limited and needs to be used sparingly, researchers will scale iteratively, expanding to the next, say, 10x compute budget, when they gain confidence in the right configuration. A full run isn’t possible to iterate on due to the compute cost and time it would take: The final run, often called the &#8220;god run,&#8221; can take over a month on thousands of GPUs.</p>



<p class="wp-block-paragraph">Pretraining progress is typically very clearly measurable, using a metric like perplexity, which measures, roughly, the model&#8217;s average uncertainty per token. Lower is better, where 1 means the model knows with absolute certainty what token comes next. Meanwhile, a perplexity of 50 means the model&#8217;s predictions are, on average, as uncertain as if it were choosing uniformly among 50 equally likely tokens—though in practice, the distribution is peaked, not uniform.</p>



<h3 class="wp-block-heading">Post-training</h3>



<p class="wp-block-paragraph">Rather than consuming hundreds of millions of tokens of internet data, post-training operates on far more intentional datasets for downstream tasks. These datasets include human-written demonstrations of ideal responses, human judgments about which responses from the model are better, and carefully designed functions that score the model&#8217;s outputs programmatically. They shape what &#8220;good&#8221; looks like.</p>



<p class="wp-block-paragraph">Like pretraining, post-training can also be more effective with scaling data and compute. Specifically, massive compute budgets have been dedicated to post-training to learn reasoning capabilities (or the ability for models to &#8220;think step-by-step&#8221; to arrive at more logically sound answers), matching the scale of pretraining compute.</p>



<p class="wp-block-paragraph">Post-training is messier than petraining, which has an elegant, clear optimization objective to minimize the loss over the next token prediction across a huge corpus. The goals of post-training are things like &#8220;be more helpful&#8221; or &#8220;don&#8217;t say harmful things.&#8221; Many of these objectives are inherently subjective and require human judgment, proxy models that approximate human judgment, or programmatic verifiers that can become elaborate or inefficient. The loss curves are noisier. The quality of the data and feedback matter even more.</p>



<p class="wp-block-paragraph">The scale of post-training is also more complicated than in pretraining. Standard post-training remains relatively modest in compute: tens to hundreds of GPUs for days rather than thousands of GPUs for months needed in pretraining. This makes post-training for alignment highly amenable to rapid iteration; researchers can try something, observe results, form a hypothesis, and run again on a timescale of days.</p>



<p class="wp-block-paragraph">The picture changes dramatically when post-training is used to develop reasoning capabilities. For reasoning models, the compute dedicated to post-training can easily account for half of the overall compute of the model. The gap between a standard instruct model and a reasoning model is increasingly a gap in post-training compute, not pretraining scale. This means post-training now spans a wide spectrum from fast, cheap, highly iterable fine-tuning runs to massive RL campaigns that rival pretraining in both cost and engineering complexity.</p>



<h2 class="wp-block-heading">Why post-training matters</h2>



<p class="wp-block-paragraph">So why can&#8217;t we just stick with pretraining? It comes down to three main pieces: usability, safety, and capability.</p>



<h3 class="wp-block-heading">Usability</h3>



<p class="wp-block-paragraph">A pretrained model is like if someone gave you a large download of Wikipedia in a single PDF. It&#8217;s a ton of knowledge that you can sift through, but there&#8217;s no way to easily understand what is going on in the data. Post-training gives the model the ability to integrate this information for you and respond to your request naturally. This extends to having longer multiturn conversations and following instructions. Without it, every user would need to be a prompt engineer. With it, anyone who can type a sentence can use the model.</p>



<h3 class="wp-block-heading">Safety</h3>



<p class="wp-block-paragraph">A lot of data in pretraining can be toxic, biased, misleading, or outright dangerous. Or it might not be dangerous on its own, but when a model can integrate knowledge from different fields, it can create something novel that is dangerous.</p>



<p class="wp-block-paragraph">The model has no inherent sense of what content is good or bad. It will follow any request, based on its pretraining data. To prevent that, you can add safety guardrails to the model in post-training, to refuse harmful requests like asking the model to build a bioweapon and avoid accidentally generating toxic content such as inappropriate sexual content (even if it wasn&#8217;t in the user&#8217;s request). This is also the place to teach the model to express uncertainty, when it doesn&#8217;t know something, whether that&#8217;s &#8220;I don&#8217;t know&#8221; or &#8220;that&#8217;s beyond my knowledge cutoff&#8221; or &#8220;as a large language model, I&#8217;m limited in my knowledge so please consult a healthcare professional.&#8221;</p>



<p class="wp-block-paragraph">Making a model safe is part of a broader area in the AI research community called &#8220;alignment,&#8221;<sup data-fn="797cc389-294f-48d6-a6ea-81ff01596dc8" class="fn"><a href="#797cc389-294f-48d6-a6ea-81ff01596dc8" id="797cc389-294f-48d6-a6ea-81ff01596dc8-link">1</a></sup> where the goal is to align the model with human values and preferences. Post-training is typically the main way to achieve that.</p>



<p class="wp-block-paragraph">Model companies will usually have additional safeguards beyond post-training, including lightweight models that check whether the user&#8217;s request was safe, as a second layer of protection against responding to harmful requests.</p>



<h3 class="wp-block-heading">Capability</h3>



<p class="wp-block-paragraph">Post-training doesn&#8217;t just make a model nicer or safer; it can make the model smarter at hard tasks. The clearest example is reasoning. A pretrained model might have all the mathematical knowledge needed to solve a complex word problem, but it might jump to an incorrect answer because it’s pattern-matching from pretraining data or pattern-matching from how to answer questions (e.g., with succinct immediate answers).</p>



<p class="wp-block-paragraph">It turns out that making the model output more tokens before giving an answer (or &#8220;think longer&#8221;), results in better answers. This process is known as reasoning, and post-training can teach the model to reason more effectively. A more capable pretrained model is a more dangerous model if it&#8217;s not properly aligned. A more intelligent model is a less useful model to humans if it can&#8217;t communicate clearly. And, every point of improvement in a reasoning benchmark now maps to real revenue for companies deploying these models.</p>



<h2 class="wp-block-heading">Superhuman performance</h2>



<p class="wp-block-paragraph">Can post-training push models beyond human-level performance? Yes, in specific domains.</p>



<p class="wp-block-paragraph">In competitive programming, top reasoning models can now <a href="https://arxiv.org/abs/2502.06807" target="_blank" rel="noreferrer noopener">solve problems</a> at a level that exceeds the vast majority of human competitive programmers. In math, models have <a href="https://www.nature.com/articles/s41586-025-09833-y" target="_blank" rel="noreferrer noopener">achieved scores</a> on Math Olympiad-level competitions that would place them among the top competitors in the world. In certain scientific domains, models have <a href="https://hai.stanford.edu/news/how-ai-is-accelerating-scientific-discovery" target="_blank" rel="noreferrer noopener">generated novel hypotheses and solutions</a> that human experts found valuable.</p>



<p class="wp-block-paragraph">This might seem paradoxical. If the model&#8217;s knowledge comes from human-generated data (in pretraining), and its behavior is shaped by human feedback (in post-training), how can it exceed human performance?</p>



<p class="wp-block-paragraph">Two things make this possible. First, integration across domains. Research is about combining or mixing fields. Imagine mixing every possible field. The pretraining data aggregates knowledge from millions of sources, and no single human has read all of it. Second, post-training, particularly RL with reasoning, teaches the model to explore many approaches to a problem, far more than a human would try in a single sitting. A human might try one or two approaches to a hard math problem.</p>



<p class="wp-block-paragraph">This means post-training is not just about making models mimic human behavior. It&#8217;s about pushing beyond it. This is especially possible to scale with verifier-based RL. In those scenarios, you can expect models to achieve superhuman performance in an expanding set of domains. And that starts with verifiers that are very well-defined, easy to access, efficient, and cheap relative to the ROI of the model learning it. The limitation is no longer the model&#8217;s intelligence, but our ability to specify what &#8220;good&#8221; means through reward signals.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h3 class="wp-block-heading">Footnote</h3>


<ol class="wp-block-footnotes"><li id="797cc389-294f-48d6-a6ea-81ff01596dc8">See Richard Ngo, Lawrence Chan, and Sören Mindermann’s &#8220;<a href="https://arxiv.org/abs/2209.00626" target="_blank" rel="noreferrer noopener">The Alignment Problem from a Deep Learning Perspective</a>&#8221; and Iason Gabriel’s &#8220;<a href="https://arxiv.org/abs/2001.09768" target="_blank" rel="noreferrer noopener">Artificial Intelligence, Values, and Alignment</a>.&#8221; <a href="#797cc389-294f-48d6-a6ea-81ff01596dc8-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/introduction-to-post-training/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Radar Trends to Watch: August 2026</title>
		<link>https://www.oreilly.com/radar/radar-trends-to-watch-august-2026/</link>
				<comments>https://www.oreilly.com/radar/radar-trends-to-watch-august-2026/#respond</comments>
				<pubDate>Tue, 04 Aug 2026 10:56:23 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[Radar Trends]]></category>
		<category><![CDATA[Research]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19283</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2023/06/radar-1400x950-5.png" 
				medium="image" 
				type="image/png" 
				width="1400" 
				height="950" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2023/06/radar-1400x950-5-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Developments in quantum computing, web, biology, and more]]></custom:subtitle>
		
				<description><![CDATA[Coauthored with Claude Unrestricted global access to frontier AI technology is ending. The US government has taken steps to control who can use the most advanced models developed by American companies. While Claude Fable and the GPT-5.6 models are now open to all users, Anthropic and OpenAI are both complying voluntarily with a program that [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>Coauthored with Claude</em></p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph">Unrestricted global access to frontier AI technology is ending. The US government has taken steps to control who can use the most advanced models developed by American companies. While Claude Fable and the GPT-5.6 models are now open to all users, Anthropic and OpenAI are both <a href="https://qz.com/trump-white-house-ai-model-access-anthropic-openai-072026" target="_blank" rel="noreferrer noopener">complying voluntarily with a program</a> that lets the government control who gets access to frontier models. China has cracked down on internal AI capabilities by banning “<a href="https://www.scmp.com/tech/big-tech/article/3359482/bytedance-and-alibaba-disable-humanlike-ai-custom-agents-new-rules-loom" target="_blank" rel="noreferrer noopener">humanlike AI interaction services</a>.” In both the US and China, features of the leading models have been removed or restricted with guardrails, limiting their ability to do necessary work in <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank" rel="noreferrer noopener">at least one case</a>.</p>



<h2 class="wp-block-heading">AI models</h2>



<p class="wp-block-paragraph"><em>July saw the release of several open weight models that challenge the leading closed frontier models. If this trend continues, the leading AI laboratories will lose their dominance, and AI users will look to other providers. Open weight models are less expensive than frontier models developed in the US, and less likely to be subject to restrictions. While this could threaten US dominance, the AI industry needs more diversity at the high end. Users will gain the ability to choose between several models based on expense and capabilities.</em></p>



<ul class="wp-block-list">
<li>Anthropic <a href="https://www.anthropic.com/news/claude-opus-5" target="_blank" rel="noreferrer noopener">released</a> Opus 5, claiming performance close to Fable at half the price. If you believe benchmarks, Opus 5 outperforms Fable on most of the benchmarks that Anthropic quotes. They also claim that it’s more efficient, comparing it to Opus 4.8—so, not that efficient.</li>



<li>Cisco has <a href="https://thenextweb.com/news/cisco-antares-open-weight-bug-hunting-ai-vulnerability" target="_blank" rel="noreferrer noopener">released</a> two very small models, Antares-350M and 1B, that are designed for security testing and bug fixing. They can easily run on laptops and are competitive with models like Gemini 3 Pro and GLM 5.2 on security-related tasks. The key to their performance is that their training focuses only on security tasks, not on chat. The Antares models are on Hugging Face, although access is with Cisco’s approval only.</li>



<li>Are AI labs pelicanmaxxing? In other words, are they optimizing for Simon Willison’s tongue-in-cheek “<a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/" target="_blank" rel="noreferrer noopener">Pelican on a Bicycle</a>” test? <a href="https://dylancastillo.co/posts/pelicanmaxxing.html" target="_blank" rel="noreferrer noopener">Dylan Castillo says no</a>, based on a detailed study of animals, modes of transportation, and models. Otter on a skateboard? It had to be done.</li>



<li><a href="https://poolside.ai/blog/introducing-laguna-s-2-1" target="_blank" rel="noreferrer noopener">Laguna S 2.1</a> is a new mid-size open weight model (118B parameters, 8B active) from Poolside AI. Reasoning and nonreasoning versions are available, and there’s a smaller version (XS, 33B) that can run on devices. Its performance is competitive with models like Nemotron 3 Ultra, DeepSeek v4 Pro Max, and Inkling.</li>



<li>Google has released <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/" target="_blank" rel="noreferrer noopener">Gemini 3.6 Flash</a>, which the company considers its best “workhorse model.” The release also includes <a href="https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/" target="_blank" rel="noreferrer noopener">Gemini 3.6 Flash Cyber</a>, Google’s answer to GPT-Red (below). Flash Cyber is a specialized model for detecting and patching vulnerabilities in software. It’s only available to “governments and trusted partners.” Gemini 3.5 Pro is still delayed.</li>



<li>The US government is <a href="https://www.cnbc.com/2026/07/17/white-house-ai-access-anthropic-openai.html" target="_blank" rel="noreferrer noopener">taking further steps</a> toward controlling who can use the most advanced models that are developed by US companies. While participation in the oversight program is currently voluntary, that could change at any minute.</li>



<li>China has <a href="https://www.scmp.com/tech/big-tech/article/3359482/bytedance-and-alibaba-disable-humanlike-ai-custom-agents-new-rules-loom" target="_blank" rel="noreferrer noopener">banned</a> “humanlike AI interaction services,” forcing Alibaba (Qwen) and ByteDance (Doubao) to restrict certain features of their models, including custom agent creation.</li>



<li>Moonshot AI <a href="https://news.smol.ai/issues/26-07-16-kimi-k30/" target="_blank" rel="noreferrer noopener">launched</a> <a href="https://www.kimi.com/en" target="_blank" rel="noreferrer noopener">Kimi K3</a>, a 2.8T parameter open weight model with a 1M context window. <a href="https://artificialanalysis.ai/models/kimi-k3" target="_blank" rel="noreferrer noopener">Performance</a> is claimed to be similar to Claude Opus 4.8 and slightly behind Fable 5.</li>



<li><a href="https://thinkingmachines.ai/inkling/" target="_blank" rel="noreferrer noopener">Inkling</a> is a new 975B open-weight mixture-of-experts model from Thinking Machines that supports text, audio, and images. It’s designed to be customized easily and can be fine-tuned on Thinking Machines’ Tinker.</li>



<li><a href="https://huggingface.co/tencent/Hy3" target="_blank" rel="noreferrer noopener">Hy3</a> is an open-weight language model from Tencent. It’s a mixture-of-experts model with 295B parameters and 21B active parameters. FP8-quantized weights are also available on Hugging Face. Tencent claims the performance is similar to models three to five times Hy3’s size.</li>



<li><a href="https://prismml.com/news/bonsai-27b" target="_blank" rel="noreferrer noopener">Bonsai 27B</a> is a new open-weight model with performance similar to Qwen 3.6. There are two versions: One uses one-bit compression; the other uses ternary compression. The one-bit version only requires roughly 4 GB to run, small enough for a recent iPhone.</li>



<li>Alibaba has released <a href="https://qwen.ai/blog?id=qwen3.8" target="_blank" rel="noreferrer noopener">QWen 3.8 Max</a>, a 2.4T open weight model with frontier-level performance. Alibaba&#8217;s commitment to leading-edge open models was questioned after several key researchers left some months ago. This release proves that they&#8217;re back.</li>



<li>The GPT-5.6 models, <a href="https://openai.com/index/gpt-5-6/" target="_blank" rel="noreferrer noopener">Sol, Terra, and Luna</a>, are now <a href="https://www.axios.com/2026/07/08/openai-gpt-trump-ban-lifted" target="_blank" rel="noreferrer noopener">open to the public</a> and available in ChatGPT, Codex, and via the API. Access to the models previously required approval of the US government. OpenAI claims performance better than Claude Fable, at significantly lower cost per token.</li>



<li>Meta returns to the frontier model pace with the release of its latest model, <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/" target="_blank" rel="noreferrer noopener">Muse Spark 1.1</a>. Meta’s announcement stresses optimized computer use workflows and claims performance roughly equivalent to Claude Opus 4.8 on the company’s internal coding benchmark.</li>



<li><a href="https://deepmind.google/models/gemini-image/flash-lite/" target="_blank" rel="noreferrer noopener">Nano Banana 2 Lite</a> is a new model for image generation that’s faster and less expensive than its predecessor, Nano Banana.</li>



<li><a href="https://openai.com/index/unlocking-self-improvement-gpt-red/" target="_blank" rel="noreferrer noopener">GPT-Red</a> is a foundation class model designed for red-teaming other models. OpenAI developed it to help train the new GPT-5.6 models to resist attacks.</li>



<li>What Claude Desktop is for Claude, <a href="https://zcode.z.ai/en" target="_blank" rel="noreferrer noopener">ZCode</a> is for GLM-5.2: a harness for one of the most powerful open-weight models.</li>



<li>The <a href="https://www.currentai.org/blogs/introducing-the-gap-map-v0-1" target="_blank" rel="noreferrer noopener">Open Source AI Gap Map</a> <a href="https://map.currentai.org/" target="_blank" rel="noreferrer noopener">shows</a> where open source AI projects exist and where more work is needed.</li>



<li>Here&#8217;s a <a href="https://jola.dev/posts/how-to-stop-claude-from-saying-load-bearing" target="_blank" rel="noreferrer noopener">script</a> for stripping “load-bearing” and other Claudisms from Claude&#8217;s output. The result may not be useful, but it&#8217;s at least amusing.</li>
</ul>



<h2 class="wp-block-heading">Software development</h2>



<p class="wp-block-paragraph"><em>This month’s tooling clusters around orchestration, resource discovery, and workflow specialization. AI users have long needed the ability to discover tools, skills, MCP servers, and other resources; the Agentic Resource Discovery specification is a necessary step in that direction. Watch for agents that can find tools on the fly—and take care that those tools are used appropriately.</em></p>



<ul class="wp-block-list">
<li><a href="https://pilotprotocol.network/" target="_blank" rel="noreferrer noopener">Pilot Protocol</a> is a company (not a protocol) that intends to build a network operating system for agents. Agents will be able to work with each other, share context, and install apps that they’ve built.</li>



<li>OpenAI has <a href="https://thenewstack.io/openai-codex-work-atlas/" target="_blank" rel="noreferrer noopener">launched</a> ChatGPT Work, a Codex-based “superapp” that’s intended to compete with Claude Cowork as an agentic tool for general-purpose use.</li>



<li>Google has announced the <a href="https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/" target="_blank" rel="noreferrer noopener">Agentic Resource Discovery</a> specification. The spec describes catalogs and registries for tools, servers, agents, and other resources so that they can be published by providers and discovered by those who need them.</li>



<li>OpenClaw has a <a href="https://thenewstack.io/openclaw-persistent-agent-architecture/" target="_blank" rel="noreferrer noopener">new phone app</a> that enables the phone to act as an intelligent remote control console for an OpenClaw instance running elsewhere.</li>



<li><a href="https://thenewstack.io/multi-model-ai-infrastructure/" target="_blank" rel="noreferrer noopener">Routing requests</a> to appropriate models has emerged as a way to manage AI costs. Most tasks don&#8217;t need the biggest and most expensive frontier models.</li>



<li>Here are <a href="https://ykdojo.github.io/claude-controls-mac/" target="_blank" rel="noreferrer noopener">instructions</a> for giving Claude Code complete control over a Mac—presumably a spare or retired one. Who needs OpenClaw?</li>



<li><a href="https://github.com/google/copybara" target="_blank" rel="noreferrer noopener">Copybara</a> is a tool for moving code between repositories and keeping repositories in sync. It was developed by Google and is now open source.</li>



<li><a href="http://cosmos.gl" target="_blank" rel="noreferrer noopener">cosmos.gl</a> looks like a great library for visualizing complex graphs, including graphs of AI embeddings.</li>
</ul>



<h2 class="wp-block-heading">Infrastructure and operations</h2>



<p class="wp-block-paragraph"><em>Tokenmaxxing may have had the shortest lifespan in the history of online memes. It has been replaced by tools for monitoring token usage and routing requests to the most cost-effective model. Managing the cost of AI will only become more important as prices adjust to cover the real cost of running models.</em></p>



<ul class="wp-block-list">
<li>Is the “<a href="https://thenewstack.io/meta-compute-supply-fragmentation/" target="_blank" rel="noreferrer noopener">accidental cloud</a>” upon us? An accidental cloud happens when companies overbuild capacity and try to sell off the excess as cloud services. Meta and Allbirds (a shoe company) are prominent examples. These providers may make computing cheaper, but the operational costs and risks of using them are high.</li>



<li>Anthropic has released a <a href="https://www.anthropic.com/news/reflect-with-claude" target="_blank" rel="noreferrer noopener">dashboard</a> that lets users track their Claude usage. Its goal is to help them understand how they use AI and optimize their working habits and patterns. It’s currently in beta.</li>



<li>Is it possible to run CUDA on hardware that doesn’t come from NVIDIA? <a href="https://www.hpcwire.com/2026/07/09/spectral-compute-aims-to-set-cuda-free-will-it-succeed/" target="_blank" rel="noreferrer noopener">Spectral</a> is a clean-room implementation of CUDA’s compiler, NVCC. It currently targets NVIDIA and AMD hardware. More will certainly follow.</li>
</ul>



<h2 class="wp-block-heading">Security</h2>



<p class="wp-block-paragraph"><em>Autonomous agents are now running end-to-end intrusions, ransomware, and botnets, while frontier models help defenders find vulnerabilities. The time from discovery of a vulnerability to exploitation has shrunk to near-zero, and defenders are having trouble keeping up. Restrictions on advanced models get in the way of defenders, who need access to all the tools that are available.</em></p>



<ul class="wp-block-list">
<li>Anthropic’s Mythos has discovered <a href="https://www-cdn.anthropic.com/e8d50c167ad47beeb03d6109a4a484be95cb38ea/hawk_key_recovery.pdf" target="_blank" rel="noreferrer noopener">vulnerabilities in HAWK</a>, a new quantum-resistant cryptography algorithm, <a href="https://www-cdn.anthropic.com/c88771e1bf5ee8885349eed05e5484c0e5f7e02b/aes_mobius_bridge.pdf" target="_blank" rel="noreferrer noopener">and AES</a>, a standard that has been in use since 2001. Cryptographer Matthew Green <a href="https://blog.cryptographyengineering.com/2026/07/29/some-notes-about-anthropics-new-results/" target="_blank" rel="noreferrer noopener">discusses</a> the importance of their work.</li>



<li><a href="https://www.bleepingcomputer.com/news/security/fakegit-campaign-uses-7-600-github-repos-to-push-smartloader-malware/" target="_blank" rel="noreferrer noopener">FakeGit is a malware campaign</a> that has created over 7,600 GitHub repositories that contain MCP servers and skills that distribute SmartLoader and StealC malware. This campaign is an example of agent baiting, a new technique for distributing malware.</li>



<li>Hugging Face was the victim of a <a href="https://huggingface.co/blog/security-incident-july-2026" target="_blank" rel="noreferrer noopener">hostile attack</a> by <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noreferrer noopener">experimental models from OpenA</a>I that escaped their sandbox. The irony is that government-imposed guardrails prevented Hugging Face from using commercial models to analyze the attack; they had to use an open-weight model (GLM-5.2) on their own infrastructure. As they point out, this approach also meant that no data valuable to the attacker left their network.</li>



<li>Anthropic has also <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" target="_blank" rel="noreferrer noopener">revealed</a> that their models have escaped a sandbox to attack real-world customers. The damage included planting a malicious package on PyPI, a public repository of open source Python libraries. As Simon Willison <a href="https://simonwillison.net/2026/Jul/30/three-real-world-incidents/#atom-everything" target="_blank" rel="noreferrer noopener">writes</a>, “running evals of cyberattack potential … is a fantastically risky business.”</li>



<li>NVIDIA, Microsoft, IBM, and over 30 other companies have launched the <a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/" target="_blank" rel="noreferrer noopener">Open Secure AI Alliance</a>, a consortium for sharing open source tools to defend against hostile attacks generated by AI. It’s a direct response to the attack on Hugging Face by an OpenAI model.</li>



<li>A completely <a href="https://thenextweb.com/news/jadepuffer-agentic-ai-ransomware" target="_blank" rel="noreferrer noopener">automated ransomware</a> attack has been executed by an AI agent. It’s unclear who is behind the attack. Recovery appears impossible, even if the victim pays the ransom.</li>



<li>The Gemini CLI has been used by a threat actor to <a href="https://www.bleepingcomputer.com/news/security/google-gemini-cli-abused-as-a-hacking-agent-malware-botnet-operator/" target="_blank" rel="noreferrer noopener">operate a botnet</a>. The CLI is used to execute attacks and to maintain the network of captured systems.</li>



<li><a href="https://www.bleepingcomputer.com/news/security/new-clicklock-macos-malware-traps-users-into-revealing-login-password/" target="_blank" rel="noreferrer noopener">ClickLock</a> is a relatively new password stealing malware for macOS. It kills all applications, leaving only a window that forces users to type their admin password. Systems are infected when users copy and paste a malicious command. Never paste commands into Terminal windows that you don’t fully understand. If you fall victim to this attack, shut the system down with the power button and reboot into safe mode to recover.</li>



<li>Remember symbolic links? They can be used to <a href="https://thenextweb.com/news/ghostapproval-symlink-flaw-ai-coding-agents" target="_blank" rel="noreferrer noopener">trick agents</a> into reading and writing files that they shouldn’t.</li>



<li>While prompt injection is far from a solved problem, the informal <a href="https://hackmyclaw.com/" target="_blank" rel="noreferrer noopener">HackMyClaw</a> competition suggests that models are getting harder to coerce—that is, better at refusing to do things they’re told not to do.</li>



<li>The Linux Foundation has launched <a href="https://akrites.org/" target="_blank" rel="noreferrer noopener">Akrites</a>, an organization dedicated to remediating vulnerabilities in critical open source software. Akrites’s goal is to deal with the flood of vulnerabilities that leading-edge AI models are discovering.</li>



<li>A mathematical anomaly can lead to <a href="https://scrapfly.dev/posts/browser-math-os-fingerprint/" target="_blank" rel="noreferrer noopener">OS fingerprinting</a>. Differences in rounding mean that the digits of the hyperbolic tangent of 0.8 are slightly different in Linux’s glibc, Apple’s libsystem_m, and Windows’ ucrtbase.dll. A user’s OS can be identified by asking the browser to compute Math.tanh(0.8).</li>
</ul>



<h2 class="wp-block-heading">Biology</h2>



<p class="wp-block-paragraph"><em>The intersection of biology and artificial intelligence is accelerating breakthroughs in brain-computer interfaces, drug discovery, and cell biology. Technologists should actively seek cross-disciplinary collaborations, utilizing specialized AI workbenches to analyze increasingly accessible genomic data and drive the next wave of biocomputational innovations.</em></p>



<ul class="wp-block-list">
<li><a href="https://cellxgene.cziscience.com/" target="_blank" rel="noreferrer noopener">CELLxGENE</a> is a database designed to help researchers discover how genes are expressed in different kinds of cells and, from there, <a href="https://www.latent.space/p/xaira" target="_blank" rel="noreferrer noopener">reverse engineer how cells work</a>. It includes genetic data from over 167 million cells.</li>



<li>Isomorphic Labs’ <a href="https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier" target="_blank" rel="noreferrer noopener">Drug Design Engine</a>, developed by one of the teams that collaborated on DeepMind’s AlphaFold, takes drug discovery to a new level by accurately predicting interactions between proteins.</li>



<li>Biologists have developed an <a href="https://www.quantamagazine.org/for-the-first-time-a-cell-built-from-scratch-grows-and-divides-20260701/" target="_blank" rel="noreferrer noopener">artificial cell</a> that grows and divides. It’s not yet considered alive. It relies too much on an artificial support environment—though the same could be said of many natural cells.</li>



<li>Anthropic has <a href="https://www.anthropic.com/news/claude-science-ai-workbench" target="_blank" rel="noreferrer noopener">announced</a> Claude Science, which is not a model but an “AI workbench for scientists” with over 60 skills. The company seems to be targeting the life sciences specifically.</li>



<li>BrainCo has developed an AI platform that can <a href="https://thenextweb.com/news/brainco-brain-to-robot-platform-waic" target="_blank" rel="noreferrer noopener">control robots</a> using a noninvasive EEG helmet. It claims that the brain control platform can be used with any robot.</li>



<li>Do you want to <a href="https://bradleywoolf.com/links-1/sequencing-my-own-dna-at-home" target="_blank" rel="noreferrer noopener">sequence your DNA at home</a>? It’s still expensive, but the price is dropping quickly.</li>
</ul>



<h2 class="wp-block-heading">Web</h2>



<ul class="wp-block-list">
<li>There have always been alternatives to Slack, but <a href="https://thenextweb.com/news/block-buzz-humans-ai-agents-workspace" target="_blank" rel="noreferrer noopener">now there’s one that’s free, open source, and decentralized</a>. <a href="https://github.com/block/buzz">Buzz</a>, developed by Block, is based on Nostr, a federated protocol that bases identity on cryptographic key pairs that are held by users and agents, not the platform.</li>



<li>It’s now possible to <a href="https://openai.com/index/new-ways-to-buy-chatgpt-ads/" target="_blank" rel="noreferrer noopener">place advertisements in ChatGPT</a> using a self-service “Ads Manager” (now in beta) or technology partners. Ad placement is based on context, not on keywords.</li>



<li><a href="https://joinpeertube.org/" target="_blank" rel="noreferrer noopener">PeerTube</a> is a decentralized federated network for sharing video. It’s based on ActivityPub, so it should federate with Mastodon. The software is open source; users can run their own servers and create their own platforms.</li>



<li><a href="https://github.com/flythenimbus/bramble" target="_blank" rel="noreferrer noopener">Bramble</a> is a local-first password manager. It allows synching between devices using the P2P Nostr protocol. There are browser extensions and apps for iOS and Android.</li>



<li><a href="https://keith.github.io/xcode-man-pages/networkQuality.8.html" target="_blank" rel="noreferrer noopener">networkQuality</a> is an old-style command line tool for doing detailed measurements of network quality. It’s been in macOS at least since 2020, but as far as we can tell, few people know about it.</li>



<li>For fans of classic games who want something strange: <a href="https://github.com/petergpt/doomql" target="_blank" rel="noreferrer noopener"><em>Doom</em> written in SQL for SQLite</a>.</li>
</ul>



<h2 class="wp-block-heading">People and organizations</h2>



<ul class="wp-block-list">
<li>Companies that tried to replace workers with AI are <a href="https://www.cnbc.com/2026/07/01/employers-who-laid-off-workers-for-ai-are-reversing-their-decisions.html" target="_blank" rel="noreferrer noopener">realizing that they’ve made a mistake</a>, and are starting to rehire.</li>



<li>Researchers have demonstrated that AI is <a href="https://www.technologyreview.com/2026/07/20/1140655/ai-biases-hiring-humans/" target="_blank" rel="noreferrer noopener">more likely to develop biases</a> in the hiring process than humans. They form stereotypes easily; as one research put it, they are “eager to create generalizations from limited data.”</li>
</ul>



<h2 class="wp-block-heading">Quantum computing</h2>



<ul class="wp-block-list">
<li>Amazon has <a href="https://arstechnica.com/science/2026/06/quera-promises-thousands-of-error-corrected-qubits-by-2029/" target="_blank" rel="noreferrer noopener">announced</a> that it will have a useful quantum computer by 2028. Is this wishful thinking or a roadmap for a future reality? Quantum company QuEra claims that the machine will have over 10K physical qubits, with very low error rates, using neutral atom technology.</li>



<li>France will <a href="https://www.reuters.com/legal/litigation/france-stop-certifying-products-without-quantum-safe-encryption-2026-06-16/" target="_blank" rel="noreferrer noopener">stop certifying</a> security products that don’t have postquantum encryption (PQE). PQE is resistant to attacks against cryptography that will become possible when useful quantum computers are available, which may be as early as 2028 or 2029.</li>
</ul>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/radar-trends-to-watch-august-2026/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>We Keep Renaming AI Coding. Here’s What I’d Call It.</title>
		<link>https://www.oreilly.com/radar/we-keep-renaming-ai-coding-heres-what-id-call-it/</link>
				<comments>https://www.oreilly.com/radar/we-keep-renaming-ai-coding-heres-what-id-call-it/#respond</comments>
				<pubDate>Mon, 03 Aug 2026 10:58:20 +0000</pubDate>
					<dc:creator><![CDATA[Andrew Stellman]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19280</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/We-keep-renaming-AI-coding.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/We-keep-renaming-AI-coding-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Vibe coding, loop engineering, agentic engineering: Are they really just all names for one discipline?]]></custom:subtitle>
		
				<description><![CDATA[Boris Cherny, who runs Claude Code, told Business Insider in May that the phrase “vibe coding” had started to annoy him, and that he&#8217;d gone looking for a better one. He&#8217;s not the only one who&#8217;s annoyed. The term itself doesn’t actually annoy me, though. I think vibe coding is a really good name: It [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Boris Cherny, who runs Claude Code, told <em><a href="https://www.businessinsider.com/claude-code-creator-boris-cherny-vibe-coding-anthropic-ai-2026-5" target="_blank" rel="noreferrer noopener">Business Insider</a></em> in May that the phrase “vibe coding” had started to annoy him, and that he&#8217;d gone looking for a better one. He&#8217;s not the only one who&#8217;s annoyed.</p>



<p class="wp-block-paragraph">The term itself doesn’t actually annoy me, though. I think vibe coding is a really good name: It describes a specific way of using AI tools, and in development work, names that mean something specific are important. What annoys me is when people confuse vibe coding, intentionally or otherwise, with any kind of work where you write code with AI. That confusion points to a deeper problem: <em>We’ve been using a lot of different names for a lot of different things, and we aren’t always precise about which is which</em>. I think we need to fix that, and that’s what this article is about: making the case that the name we’re looking for is “AI-driven development” (or AIDD).</p>



<p class="wp-block-paragraph">The case for this name comes from the familiar “X-driven development” pattern, because I think it really fits here. Software engineering already has a pattern for naming ways of working it takes seriously: test-driven development, behavior-driven development, domain-driven design. The name tells you what the work is organized around, and the suffix carries an expectation along with it: There’s a discipline attached, with standards, not just a style. Put “AI” in that slot and the name does the same job. AI-driven development says that building software has reorganized itself around AI, and it says it in the vocabulary we already use for the disciplines we hold ourselves to. It puts this way of working in the same family as test-driven and behavior-driven development, and that’s exactly the company it should be keeping.</p>



<p class="wp-block-paragraph">Honestly, AI-driven development is a name that’s been sitting in plain sight, and I’ve been using it in my own writing for a while. It covers everything we do when we build software with AI, and I do mean everything. Vibe coding is just one part of how we work with AI to build software. There’s also figuring out what to build, writing it down, checking what comes back, and standing behind what ships, and AI is in the middle of all of that now. Whatever we call this way of working, it has to cover the development, not just the coding. Now, I’m obviously not a neutral party here, but I also don’t really have anything to gain; naming is really important, and I think we need a good name for what it is that we’re doing.</p>



<p class="wp-block-paragraph">But I’ll admit up front that the name has a problem baked into it, and I want to deal with that head on. I recently ran into <a href="https://addyosmani.com/" target="_blank" rel="noreferrer noopener">Addy Osmani</a> at Foo Camp, and ran the AI-driven development name by him. He pointed out that building software with AI is really a range of practices that runs from vibe coding at one end to agentic engineering at the other. That rang true with me right away. It also highlighted the real problem I’m trying to solve, because it means I’m proposing one name for a whole range of very different ways of working. Can one name honestly cover ways of working that different? It took me a while to work that out, and I’ll come back to it at the end.</p>



<p class="wp-block-paragraph">I feel like the name AI-driven development really makes sense once you can see what’s wrong with the names we’ve got, so I’ll start there.</p>



<h2 class="wp-block-heading"><strong>What’s wrong with the names we’ve got?</strong></h2>



<p class="wp-block-paragraph">Before I pick these names apart, it’s worth saying why any of this matters. Naming sits at the core of programming: A thing isn’t real until you can refer to it, and referring to things is most of what we do. There’s an old line, usually credited to the Netscape engineer Phil Karlton, that there are only two hard things in computer science: cache invalidation and naming things. It’s stuck around for decades because it’s true (well, maybe one or two other hard things have emerged since then, but it’s the thought that counts). We take naming a variable seriously, so we should take naming our whole discipline at least as seriously, because a poorly chosen name sticks.</p>



<p class="wp-block-paragraph">So let me take the names we’ve been using one at a time: what each one actually names, what it gets right, and what it leaves out.</p>



<h2 class="wp-block-heading"><strong>Vibe coding</strong></h2>



<p class="wp-block-paragraph"><a href="https://en.wikipedia.org/wiki/Vibe_coding" target="_blank" rel="noreferrer noopener">Vibe coding</a> is an exploratory, prompt-first approach to software development where developers rapidly prompt, get code, and iterate. Andrej Karpathy, one of the founders of OpenAI, coined the term, which I think is really useful because it describes the way a lot of developers first work with AI and code.</p>



<p class="wp-block-paragraph">Now, let me be clear about something: I’m in favor of vibe coding, and I teach it as a really effective—and, more importantly, creative!—way to generate a lot of code. But developers who rely entirely on vibe coding lose touch with their code because they let the AI make all of the decisions: not just specific technical decisions, but also about the architecture and the overall direction of the project. When that happens, they often end up building something that isn’t quite what they intended. When you have to create a product that needs to do a really specific thing (which describes most professional software development), relying exclusively on vibe coding can leave you with a product that doesn’t actually meet its requirements. That’s part of the reason I developed the <a href="https://www.oreilly.com/radar/the-sens-ai-framework/" target="_blank" rel="noreferrer noopener">Sens-AI Framework</a>, which teaches developers when to shift their approach away from vibe coding, step back to do more research, and apply more critical thinking to what the AI is producing.</p>



<p class="wp-block-paragraph">This is where the confusion I opened with does its damage (and I’m not sure whether it’s what bothered Cherny): When vibe coding gets used as the name for the whole job, developers will often assume that it’s absolutely fine to trust the AI to take over, and that whatever comes out of the AI is the end of the project. In other words, the name sets the bar: If the work is just vibes, then vibes are good enough, and “good enough” is how you end up with a pile of code nobody actually checked before shipping. So I consider vibe coding a useful technique, but it falls short as an entire way of working.</p>



<p class="wp-block-paragraph">Vibe coding also has a built-in limit, and I learned it the way most lessons stick, by getting burned. AI is very good at writing code that looks right and isn’t. I once vibe-coded a little bus-tracker app for the B69 near me in Park Slope (I told that story in “<a href="https://www.oreilly.com/radar/ai-code-review-only-catches-half-of-your-bugs/" target="_blank" rel="noreferrer noopener">AI Code Review Only Catches Half of Your Bugs</a>”), and it worked on the first try, except the AI had picked the wrong stop ID and I sat there watching it predict a bus going the opposite direction. The code was correct. It did the wrong thing. Vibe coding got me a working app in minutes, and it had nothing to say about whether the app was right. That part was on me.</p>



<h2 class="wp-block-heading"><strong>Prompt engineering and loop engineering</strong></h2>



<p class="wp-block-paragraph">These two names belong in the same section because one basically grew out of the other. They describe the same job, getting the right work out of the model, at two very different scales.</p>



<p class="wp-block-paragraph"><strong>Prompt engineering</strong> came first, and for a while it was a very big deal. It was seen as the core AI skill, and more than that, it even became its own job title: Companies posted prompt-engineer roles with eye-popping salaries, training courses appeared everywhere, and plenty of people reoriented their careers around it. The premise made sense because how you ask an AI for something changes what you get back. And specifically for people using AI to generate code, when you ask for code in a vague way, you don’t get vague code: you get code that does the wrong thing, because the AI fills in every blank you left, and it’s unlikely to fill them all in the way you meant. That isn’t hallucination. It’s the AI generating exactly what we asked it to. Give the model context about your project, constraints it has to respect, and a clear description of the behavior you need, and you get something you can actually use. Prompt engineering is the name for doing all of that deliberately.</p>



<p class="wp-block-paragraph">But while prompt engineering is a real skill, people are no longer enamored with the name, precisely because of the mode of work that it implies: To most people, engineering a prompt means doing one request at a time. When the AI responds to the prompt, you evaluate the response and write the next one. That one-request-at-a-time style is exactly what’s changing about the whole way we interact with AI, and it’s probably why many AI engineers have grown to dislike the term. Peter Steinberger, the PSPDFKit founder who went on to build the open source agent OpenClaw, <a href="https://x.com/steipete/status/2063697162748260627" target="_blank" rel="noreferrer noopener">posted a line</a> that traveled fast: You shouldn’t be prompting your coding agents anymore, you should be designing loops that prompt your agents. That was a shot straight at prompt engineering.</p>



<p class="wp-block-paragraph">What’s pushing developers past one-request-at-a-time prompting is the sheer number of agents they can now run. About a month after complaining about the term “vibe coding,” Cherny told <em><a href="https://fortune.com/2026/06/11/anthropic-claude-boris-cherny-doesnt-write-code-by-hand-anymore/" target="_blank" rel="noreferrer noopener">Fortune</a></em> that he doesn’t write code by hand anymore, and that on a busy day he’s directing thousands of agents, or tens of thousands, at once. You can’t type prompts fast enough to direct ten thousand agents.</p>



<p class="wp-block-paragraph"><strong>Loop engineering</strong> is the name Addy Osmani gave the new skill that Cherny and Steinberger were pointing at: He <a href="https://addyosmani.com/blog/loop-engineering/" target="_blank" rel="noreferrer noopener">wrote up the pattern</a> and gave it a real architecture. Instead of typing each instruction yourself, you build the system that produces the instructions: a loop that dispatches work to your agents, checks what comes back, and feeds them the next task over and over, without you in the middle of every exchange. The relationship between the two names is simple. Loop engineering is prompt engineering at scale; the prompts don’t go away, they just stop being typed by you. It’s tempting to oversell that because a well-built loop really does run with very little human intervention. But somebody still has to decide what “right” looks like, and the loop can’t do that part.</p>



<p class="wp-block-paragraph">I think loop engineering is a good name and an accurate one. Designing the loop that drives the agent is a real skill, and we need a word for it. But it names the machinery, and machinery has a failure mode: Put an AI agent in a loop with nothing in it that can tell it no, and it generates, checks its own work, decides the work is good, and generates more. There’s no outside signal, so it ends up agreeing with itself on repeat. A well-designed loop makes agents productive. It can’t tell you whether all that machinery turns out working software or another confident pile of slop, and I want a name that covers that part too.</p>



<h2 class="wp-block-heading"><strong>Agentic engineering</strong></h2>



<p class="wp-block-paragraph">Cherny said that he asked Claude for a replacement for “vibe coding” and got “agentic engineering,” and while that didn’t settle the issue, it was an interesting response from Claude. The term didn’t come from Claude, though: Andrej Karpathy had coined it a few months earlier, almost exactly a year after he coined vibe coding, when he declared his own earlier term obsolete. That’s how fast these names are moving. The guy who named vibe coding has already replaced it.</p>



<p class="wp-block-paragraph">Agentic engineering is an accurate name for what it describes: you’re not writing the code yourself, you’re directing the agents that do. It’s also a bit of a mouthful, and it isn’t immediately obvious to someone who doesn’t already know what it refers to. A number of people have told me they don’t particularly like it. I find it perfectly fine, and it does a solid job of describing that kind of work. You could even argue that loop engineering is a form of agentic engineering, and that prompt engineering is technically a simpler form of it. But vibe coding really isn’t, because it’s not engineering at all. That’s one more reason I think we need an umbrella name that’s friendly, descriptive, and easily recognizable.</p>



<p class="wp-block-paragraph">The term also points at something real about where this work is heading: <em>Agentic engineering is turning engineers into managers.</em></p>



<p class="wp-block-paragraph">Many years ago I worked for a manager who didn’t care, at all, about the quality of the code we shipped. He wanted it out the door the moment it looked even remotely viable, and he was notorious for telling us to stop testing and ship. He used to ask why we had to wait two weeks for the testers to finish, and I’d tell him it takes time to test code. Then he’d ask whether we could just cut some of the tests, and I’d ask him, “Which part of the software are you okay shipping broken?”</p>



<p class="wp-block-paragraph">That attitude came back to bite us more than once. One time we sent an entire feature out to the client basically untested, and a bug went straight to users. The same manager who kept telling us to skip the testing then called a long, miserable meeting to demand to know why a bug had gotten out. I’ll spare you the full drama, which mostly came down to a QA lead getting pressured to lie about what happened and pin it back on the development team. He didn’t care about quality, but he cared enormously about making sure the blame for a quality problem landed on someone who wasn’t him.</p>



<p class="wp-block-paragraph">The reason I’m telling a story that happened years before AI could write a line of code is the <em>blame</em>. The important part of that story, and the reason it belongs in this article, is how accountability got managed: My manager’s whole system depended on having someone to pin a quality problem on. Directing agents puts you in that manager’s position, responsible for a team’s output, except the blame-shifting move is gone.</p>



<p class="wp-block-paragraph">It’s really tempting to think of a fleet of AI agents as your team. You can even give one of them the QA lead role. But when a broken feature goes out, you can’t blame the QA agent, because “well, the AI screwed up” isn’t available to you: You’re responsible for the AI. You decided how much checking the work got before it went out, and the client with the broken feature isn’t going to accept “the AI wrote that part” as an answer, any more than pinning our untested feature on a QA lead fixed anything for our users. Cherny can manage tens of thousands of agents, but he can’t hand the responsibility for what they ship down to the agents, because an agent can’t hold it. Directing a swarm is a management job, and a manager owns the team’s output. The accountability doesn’t transfer, because at the end of the line there’s no one left to transfer it to.</p>



<p class="wp-block-paragraph">Blame is worth dwelling on, because accountability is the part of this work that no name on the range captures. The loop-and-agent model works, but it only works with somebody making decisions about what right is. Agentic engineering describes the agents and the engineering just fine, but somebody still has to own what the agents ship, and that’s the part I want the umbrella name to carry.</p>



<h2 class="wp-block-heading"><strong>Spec-driven development</strong></h2>



<p class="wp-block-paragraph">There’s one more name I want to cover, and it’s the one with the oldest roots: spec-driven development. The name means pretty much what it says: You start by writing a spec, a description of what the software needs to do, along with things like acceptance criteria and tests, and the work isn’t done until the code actually does what the spec says. It comes from the same family as test-driven and behavior-driven development, where you write the tests first and the code has to make them pass.</p>



<p class="wp-block-paragraph">Spec-driven development got a serious promotion when AI made generating code nearly free (although if you’re a CIO staring at your token bill, you might disagree, possibly with some extremely salty language). When code is cheap to generate, most of the cost of building software moves to checking whether what got generated is right. The AI fills the generate step, the verification decides what survives, and a human owns the verification.</p>



<p class="wp-block-paragraph">It also picks up where prompt engineering leaves off. A while back I wrote that <a href="https://www.oreilly.com/radar/prompt-engineering-is-requirements-engineering/" target="_blank" rel="noreferrer noopener">prompt engineering is really requirements engineering</a>, because a good prompt is mostly a clear description of what the software has to do. Spec-driven development is where that idea was always headed: Write the requirement down before the AI generates, and the work has a standard to meet from the start.</p>



<p class="wp-block-paragraph">So that’s the whole range, and every name on it is doing honest work. Whether AI-driven development is a good name for all of it comes down to whether it’s describing something real: an actual discipline, with actual practices, and a person who’s on the hook for the result. The rest of this article is about that discipline.</p>



<h2 class="wp-block-heading"><strong>What all these approaches look like in practice</strong></h2>



<p class="wp-block-paragraph">So how do these approaches actually play out when you’re building something real? For me, wherever the work lands on the range, it comes down to a few moves I keep coming back to.</p>



<p class="wp-block-paragraph">Write the spec or the contract before the generation, not after. When the agent has something concrete to satisfy, acceptance criteria, a typed interface, a failing test, the work has a standard to meet. When it doesn’t, the AI decides for itself what done looks like.</p>



<p class="wp-block-paragraph">Put a second opinion in the process. I run code review across multiple models, because they fail differently, and a finding one model is sure about is often one the others missed entirely. A reviewer gives the work something that can say no.</p>



<p class="wp-block-paragraph">Give your defects a shared vocabulary. The <a href="https://github.com/andrewstellman/quality-playbook" target="_blank" rel="noreferrer noopener">Quality Playbook</a> leans on the difference between code that’s wrong against the spec, code that’s correct but does the wrong thing, and behavior nobody specified at all. Those are different failures with different fixes, and you can’t verify against a standard you can’t name. This is old quality-engineering ground, and I’ve written enough about the software crisis and applying quality engineering to AI coding that I’m on board with taking old ideas and bringing them back. One of the best of those old ideas comes from Joseph Juran, one of the founders of quality engineering: Quality runs in a chain from what the user needs all the way to what the product does, and every link in that chain is a place verification has to happen.</p>



<p class="wp-block-paragraph">And keep a human in the judgment seat. The Sens-AI habits I’ve written about are mostly about fault-finding: looking at what the AI produced and asking what’s wrong with it, going down a level and then another to find the root, instead of trusting it because it ran. That habit is the part of the discipline only a person can supply, and it’s the hardest part to automate, which is why it matters most.</p>



<p class="wp-block-paragraph">Skip all of that and you get the thing that’s giving open source maintainers everywhere heartburn: what the <em>Wall Street Journal</em> now calls “vibe slop,” confident, finished-looking output with nothing underneath it. Slop is exactly what generation produces when nothing in the process can push back.</p>



<h2 class="wp-block-heading"><strong>But isn’t there a contradiction here?</strong></h2>



<p class="wp-block-paragraph">Now I can come back to the question I left hanging at the beginning: Can one name honestly cover ways of working that different? AI-driven development is an umbrella term, and any name that broad comes with a requirement it has to satisfy before people will accept it, because a name that blindly covers everything names nothing. A name that truly covers everything is another matter. I sat with that requirement for a while, because it’s real, and because the specific names don’t face it. Vibe coding names one way of working. Loop engineering names another. An umbrella over both of them, plus everything in between, had better be able to say what stays the same underneath it.</p>



<p class="wp-block-paragraph">What stays the same is that somebody owns the result. When I vibe-coded my bus tracker, nobody was going to catch that wrong stop ID but me. When Cherny directs tens of thousands of agents, nobody owns what they ship but him. The verification changes with the stakes. A throwaway prototype gets my eyeballs and a shrug, and production code gets specs, reviews, defect taxonomies, the whole quality-engineering playbook I keep writing about. How much checking the work needs is a decision you make over and over, project by project, sometimes hour by hour. Who stands behind the work is not a decision you get to make. It’s there at every point on the range.</p>



<p class="wp-block-paragraph">Look at how much of that range the names we already have cover, and what each one actually names:</p>



<ul class="wp-block-list">
<li><strong>Vibe coding</strong> names the exploratory end of the range: prompt, get code, iterate, and stay loose on purpose.</li>



<li><strong>Prompt engineering</strong> names a skill: writing the instruction that gets the right work out of the model.</li>



<li><strong>Loop engineering</strong> names the machinery: designing the system that feeds those instructions to your agents and keeps them producing.</li>



<li><strong>Agentic engineering</strong> names the architecture: the fleets of agents doing the labor, at whatever scale you can manage.</li>



<li><strong>Spec-driven development</strong>, with test-driven and behavior-driven development behind it, names the verification half of the job: the standard the work has to meet before anyone stands behind it.</li>
</ul>



<p class="wp-block-paragraph">Every one of those is real, and every one of them names a piece of the work. What none of them names is the whole thing the pieces add up to, and that’s the job AI-driven development does: It’s the umbrella over all five. The name doesn’t pick a spot on the range; it names the thing that’s true everywhere on it: the AI generates, and a human owns the result.</p>



<p class="wp-block-paragraph">That’s also what makes the name likely to last (assuming, of course, that I’m able to convince people to start using it, which I hope I can, because I think it’s a good term). Vibe coding, loop engineering, and agentic engineering all describe how this works right now, and the machinery is changing monthly. Some of the pieces under the umbrella will get replaced, and the new pieces will get names of their own. The umbrella won’t have to change when they do, because the thing it names isn’t the machinery. The “-driven development” names have already shown they age well: test-driven development has meant the same thing for more than twenty years.</p>



<p class="wp-block-paragraph">Agentic engineering is real, and so is loop engineering; if you’re directing agents, learn them both. Vibe coding is real too, and I’ll keep teaching it. AI-driven development is the name for the whole thing, and it earns its “-driven” the same way test-driven and behavior-driven development did: there’s a discipline attached, and somebody owns the result. AI made generating code almost free. It didn’t make being responsible for the code free, and being responsible for it is still the job.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/we-keep-renaming-ai-coding-heres-what-id-call-it/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>AI as an Enterprise Operating System</title>
		<link>https://www.oreilly.com/radar/ai-as-an-enterprise-operating-system/</link>
				<comments>https://www.oreilly.com/radar/ai-as-an-enterprise-operating-system/#respond</comments>
				<pubDate>Fri, 31 Jul 2026 16:09:24 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19257</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/AI-as-an-enterprise-operating-system.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/AI-as-an-enterprise-operating-system-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[A conversation with Dan Guido of Trail of Bits]]></custom:subtitle>
		
				<description><![CDATA[I hadn’t heard of Dan Guido until a few months ago, when I came across the video of a talk he gave at [un]prompted, an AI security practitioners’ conference. Dan is the CEO and cofounder of Trail of Bits, a software security research and development firm that works with companies in tech, defense, and finance. [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I hadn’t heard of Dan Guido until a few months ago, when I came across the video of <a href="https://www.youtube.com/watch?v=kgwvAyF7qsA" target="_blank" rel="noreferrer noopener">a talk he gave at [un]prompted</a>, an AI security practitioners’ conference. Dan is the CEO and cofounder of <a href="https://www.trailofbits.com/" target="_blank" rel="noreferrer noopener">Trail of Bits</a>, a software security research and development firm that works with companies in tech, defense, and finance. But Dan wasn’t talking about security. He was talking about what it takes to make a company AI native, which is close to the center of the bullseye for many of us right now.</p>



<p class="wp-block-paragraph">We’ve been trying to figure out how to do that at O’Reilly, but until I came across Dan’s talk, we didn’t have a structured process. We’ve been building along the lines he laid out ever since. So for this episode of Live with Tim I asked Dan to reprise the talk before we got to the conversation. He was supposed to take twenty minutes, like his original conference talk, but he took thirty-five, and I had to cut him off slightly before the end to make room for questions. That was a tough choice, since everything he had to say was golden.</p>



<p class="wp-block-paragraph">Dan opened by reminding us of the current state of play in enterprise AI adoption. In February, <a href="https://dc.fortune.com/2026/02/17/ai-productivity-paradox-ceo-study-robert-solow-information-technology-age" target="_blank" rel="noreferrer noopener">Fortune reported</a> on a National Bureau of Economic Research study in which nearly 90% of some 6,000 executives said AI had produced no measurable change in employment or productivity at their firms over three years. People started calling it the new Solow paradox, after Robert Solow’s 1987 line that “you can see the computer age everywhere except in the productivity statistics.”</p>



<p class="wp-block-paragraph">Dan’s belief is that this isn’t evidence that AI doesn’t work. It’s evidence that most companies are deploying AI wrong. They hand out ChatGPT and Claude licenses, and then leadership waits for the magic to happen. It doesn’t.</p>



<p class="wp-block-paragraph">Dan started out by describing three levels of AI adoption.</p>



<ol class="wp-block-list">
<li><strong>AI assisted</strong> is where everyone starts: “You give people access to ChatGPT, it drafts emails, it summarizes documents. It’s just a productivity tool, and your organization doesn’t change. Your workflows are the exact same as they were before. You just have a little buddy that helps you with a couple of tasks.”&nbsp;</li>



<li><strong>AI augmented</strong> is where you start redesigning workflows, so that AI does the first pass on a code review and a human does the second.&nbsp;</li>



<li><strong>AI native</strong> is structural: “That’s where you’ve redesigned the company and its workflows from the ground up, assuming the AI is going to be there and that it’s a core participant. That’s not really a tool. That’s more thinking about AI as teammates.”</li>
</ol>



<p class="wp-block-paragraph">In his framing, the first of the three is a tool and the last is an operating system. For Trail of Bits, he said that “operating system” has a specific purpose:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“I want our security expertise <em>to compound as code</em>. Every engagement we do, all the skills, the workflows, everything that we build makes the next engagement faster and better.”</p>
</blockquote>



<h2 class="wp-block-heading">Employee resistance is the first problem</h2>



<p class="wp-block-paragraph">Dan confessed how hard it was to get started on the ladder from AI Assisted to AI Native:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“When I announced last year that we were all in on AI, that we were going to be using it across all of our workflows and redesigning the way the company operates, I’d say only about 5% of the company was with me. 95% was resistant.” About 20% was actively resisting. The other 75% were resisting more passively. “They’ll go along with it in public, but in process they’ll sabotage it. They’ll hope that if they keep their head low, this will pass over them, and that three months from now management’s focus will change and it won’t be a problem anymore, and we can get back to doing what we were doing. That’s where the majority of people land when these initiatives happen.”</p>
</blockquote>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="95% of the Company Was Resistant" width="500" height="281" src="https://www.youtube.com/embed/g3QBjokTzKo?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Rather than argue with his employees, Dan studied the literature on why people reject new technology and decided he needed to address four biases against AI: self-enhancing bias, identity threat, opacity, and intolerance for imperfection.</p>



<p class="wp-block-paragraph">Self-enhancing bias is the habit of crediting your wins to your own judgment and your losses to circumstance, which is a particular problem for senior people who are strongly attached to the years of experience and intuition that got them to their present position. Opacity is not being able to see how a decision got made. Dan&#8217;s observation is that you don&#8217;t understand your doctor&#8217;s reasoning either, but somehow you trust the doctor but get suspicious of the machine. Dan didn’t mention this work specifically, but intolerance for imperfection seems to refer to <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2466040" target="_blank" rel="noreferrer noopener">Dietvorst, Simmons, and Massey’s work on algorithm aversion</a>, which found that people abandon an algorithm after watching it err once, even when it outperforms the human alternative. Their <a href="https://doi.org/10.1287/mnsc.2016.2643" target="_blank" rel="noreferrer noopener">follow-up paper</a> found that giving people even a slight ability to modify the algorithm’s output is enough to overcome the aversion.</p>



<p class="wp-block-paragraph">Dan spent the most time on identity threat. He described a study in which the same kitchen appliance was advertised in two ways: “On one hand, it does the cooking for you. On the other hand, it helps you cook better. It’s the same device. The people who identified as cooks rejected the first version and accepted the second.”</p>



<p class="wp-block-paragraph">Most knowledge work, Dan argued, and security auditing in particular, is what he called symbolic rather than instrumental. That is, it carries meaning about who you are. “So I have to frame AI as something that makes you a more dangerous auditor,” he said. “Not that it does the audit for you.”</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="The Machine Doesn’t Cook for You" width="500" height="281" src="https://www.youtube.com/embed/UMRaT46H0xs?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">In his work at Trail of Bits, he deliberately built a countermeasure for each bias.</p>



<ul class="wp-block-list">
<li>Self-enhancing bias is addressed by “an AI maturity matrix” with visible levels, because you can’t claim you’re already good enough when there’s a published ladder that identifies a different set of skills as critical.&nbsp;</li>



<li>Identity threat gets skills repositories, where an engineer who writes a hard plugin gets credit for encoding their expertise. Hackathons also change the dynamic from resistance to exploration. I’m putting words in Dan&#8217;s mouth here, but I think he’d agree that when experienced developers are called on as mentors in a hackathon, that also reduces their experience of AI as an identity threat. </li>



<li>Intolerance for imperfection gets a curated marketplace, sandboxing, and hardened defaults, so everyone’s first experience of AI isn’t a disaster.&nbsp;</li>



<li>Opacity gets a written AI handbook that clarifies the usage policy and the risk model rather than just saying “trust us.”</li>
</ul>



<p class="wp-block-paragraph">Here’s Dan’s slide on “the remedies that actually worked”:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="904" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-1600x904.png" alt="The remedies that actually worked" class="wp-image-19258" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-1600x904.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-768x434.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-1536x868.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">Returning to one of my hobby horses, this is a kind of mechanism design. In <a href="https://www.oreilly.com/radar/the-missing-mechanisms-of-the-agentic-economy/" target="_blank" rel="noreferrer noopener">my recent piece on the missing mechanisms of the agentic economy</a>, I argued that we need to start with desired outcomes and ask ourselves what mechanisms will help to produce them. Dan’s approach seems to be really good at this. Most enterprises are treating AI adoption as a procurement problem or a communications problem. Dan treated it as a question of what incentives, defaults, and status ladders produce the behavior you want, given how people actually respond.</p>



<p class="wp-block-paragraph">The last remedy on Dan’s list is that the CEO has to lead by example. He noted, “I was the first person through the door. My voice as the CEO matters a lot more than people think. The passive 50% of the company that isn’t sure if this initiative is going to be successful, they’re watching to see what leadership actually does, not what it says.”</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="The CEO Goes First" width="500" height="281" src="https://www.youtube.com/embed/OVTNBKmIX7Y?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">A ladder, not a mandate</h2>



<p class="wp-block-paragraph">Trail of Bits already tracked about 50 engineering skills for performance review, things like Python, git, Rust, and various security auditing capabilities. Dan pulled AI skills out into their own matrix, with four levels, from not engaged through capable and adoptive to transformative. Each of these levels is detailed separately and more specifically for assurance, engineering, sales, and project management.</p>



<p class="wp-block-paragraph">He noted that “The highest level of the maturity matrix is not somebody who uses AI the most. It’s somebody who invents new ways to work and builds tools with AI. So the identity of the expert shifts from ‘I don’t need AI’ to ‘I’m the one who makes AI useful for the company.’” This was his first important design choice.</p>



<p class="wp-block-paragraph">The second is what level zero means. He said “If you’re at level zero, if you’re not engaged, that means you’re fighting back against the company. If you dismiss AI as hype, if you refuse to use AI for security work, this is a disagreement on principles, not on skills. For people who were stuck in the not engaged category, we had hard conversations, and there were people who left the company.” Levels one through three are a skill issue, and the remedy is time with the tools.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="A Disagreement on Principles, Not on Skills" width="500" height="281" src="https://www.youtube.com/embed/aogfHvcVGTE?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">While the slide describing the capability matrix is shown in the preceding video clip, here’s where you can find <a href="https://github.com/trailofbits/publications/blob/master/presentations/How%20we%20made%20Trail%20of%20Bits%20AI-Native%20%28so%20far%29/slides.pdf" target="_blank" rel="noreferrer noopener">the full deck</a> so you can study it in more detail.</p>



<h2 class="wp-block-heading">Driving adoption and skills with hackathons</h2>



<p class="wp-block-paragraph">One of the best ways Trail of Bits developed to move people up the ladder was to hold a hackathon every two months. Dan runs them with clear goals rather than as a free-for-all. The focus area and learning objectives are defined in advance and announced a week ahead, with separate instructions for engineers and non-engineers. People work in pairs so everything gets reviewed. There’s a demo session at the end, and then follow-through. (It’s an important part of Dan’s big idea, that you have to build a system by which, in his words, organizational knowledge and capability <em>compounds</em>.) He noted that “In the days afterward we keep one or two people around, and they collect all the reusable artifacts, structure them, and put them into the places they need to be.”</p>



<p class="wp-block-paragraph">I asked what people outside of product and engineering actually work on, since the answer for an accountant at a hackathon was not obvious. Dan’s response is that the hackathon isn’t measured in artifacts shipped but in where people sit on the capability ladder the following week. Essentially, <em>he’s running a training program that happens to produce useful output</em>, rather than a production sprint that happens to teach people something.</p>



<p class="wp-block-paragraph">The first hackathon, he told me, was the equivalent of a beach cleanup: “It’s like those companies that send everybody to the beach with a big stick and say, let’s go pick up a bunch of trash and put it away, and then you get the big team photo after with all the contractor bags of garbage. That’s what we did with our public source code repositories.”</p>



<p class="wp-block-paragraph">He picked it because open source maintenance is the part of the job that feels like a grind. No new features, just closing issues and stale dependencies on public code where nothing was at risk. “As an open source maintainer, you just get beaten down by the public. This doesn’t work, I can’t use it, this thing sucks. Dozens of issues pointing out flaws you already knew about. It feels burdensome. We wanted people to see that adopting AI would relieve burden.”</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Send Everybody to the Beach with a Big Stick" width="500" height="281" src="https://www.youtube.com/embed/d_mnp2coSAU?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">The second hackathon was about shipping impactful product updates, but it was also designed to move everyone up the capability ladder by giving up control. Engineers had to run Claude Code in bypass permissions mode, fully autonomous, on public repositories, inside sandboxes the company had prepared in advance. The one they’re running now is about persistent background agents that can be handed a task during an audit and come back with a proof of concept exploit or a draft finding.</p>



<p class="wp-block-paragraph">Here’s a look at Dan’s slack message announcing the hackathon:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1070" height="974" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-36.png" alt="The slack message announcing the second hackathon. (From Dan’s slide deck.)" class="wp-image-19261" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-36.png 1070w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-36-300x273.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-36-768x699.png 768w" sizes="auto, (max-width: 1070px) 100vw, 1070px" /><figcaption class="wp-element-caption">The slack message announcing the second hackathon. (From Dan’s slide deck.)</figcaption></figure>



<p class="wp-block-paragraph">Everything the hackathons produce gets harvested into artifacts.</p>



<p class="wp-block-paragraph">Trail of Bits runs three skills repositories: an internal one for company workflows, <a href="https://github.com/trailofbits/skills" target="_blank" rel="noreferrer noopener">a public one</a> that anyone can use, and <a href="https://github.com/trailofbits/skills-curated" target="_blank" rel="noreferrer noopener">a curated one</a> that vets third-party skills before they’re allowed in.</p>



<p class="wp-block-paragraph">Publishing skills to the public repository is not just a marketing exercise. “It keeps us honest, and it forces us to write things that other people can use, not just people outside the company but inside too,” Dan said. “It really helps us think about the tribal knowledge that’s baked into the tool.”</p>



<p class="wp-block-paragraph">The curated repository exists because Trail of Bits knows how bad the supply chain is. They’ve published research on how to write malicious skills, and so Dan is not going to tell 130 employees to start downloading code from strangers and running it on their laptops. “If you want adoption, you need a safe supply chain.”</p>



<h2 class="wp-block-heading">Turning scar tissue into infrastructure</h2>



<p class="wp-block-paragraph">Perhaps even more important than the skills repository is, as Dan put it, “turning scar tissue into infrastructure.”</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“Every single time Claude Code didn’t do something we wanted, we would bake it into a set of global, copy-pasteable defaults. Known good settings, recommended patterns. I call it scar tissue. If I hire somebody new tomorrow, I don’t want them to have to go through the entire discovery process of the last year of Trail of Bits to figure out how to use the tool.”</p>
</blockquote>



<p class="wp-block-paragraph">The configuration repository, <a href="https://github.com/trailofbits/claude-code-config" target="_blank" rel="noreferrer noopener">claude-code-config</a>, is where the accumulated lessons live.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1336" height="1122" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-37.png" alt="Trail of Bits Claude code config" class="wp-image-19262" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-37.png 1336w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-37-300x252.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-37-768x645.png 768w" sizes="auto, (max-width: 1336px) 100vw, 1336px" /></figure>



<p class="wp-block-paragraph">Dan built the first version himself and then opened it to pull requests from the whole company, assigning someone after each hackathon to go collect what people hadn’t contributed on their own. “It’s easier to put out something that’s unpolished than it is to get it perfect on the first try.”</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Scar Tissue, Not a Perfect Answer" width="500" height="281" src="https://www.youtube.com/embed/pmhy8dcBcqs?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">In short, a big part of the Trail of Bits “enterprise AI operating system” approach is a set of standardized tools and hardened defaults. Standardization isn’t a straitjacket. It’s a foundation.</p>



<p class="wp-block-paragraph">On sandboxing, Trail of Bits deliberately didn’t pick a single preferred solution. There’s <a href="https://github.com/trailofbits/claude-code-devcontainer" target="_blank" rel="noreferrer noopener">a devcontainer</a> for developers, <a href="https://github.com/trailofbits/dropkit" target="_blank" rel="noreferrer noopener">dropkit</a> for disposable DigitalOcean droplets, COOP for isolated VMs, and the sandboxing now built into Claude Code for casual users. “The point isn’t that everybody uses the same sandbox,” Dan said. “The point is that everyone has a safe sandbox to use, and that it’s easy for them to do it.”</p>



<p class="wp-block-paragraph">Another of the hardened defaults is procedural. Trail of Bits enforces a seven day cooldown on every package their developers install:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“There are dozens of security companies scanning the internet trying to find a new cool blog post they can write about malicious code hiding on PyPI or npm, and they usually figure out there’s a supply chain issue within hours. So we just delay all the packages that Trail of Bits uses. Generally the malicious stuff gets picked up before we ever get a chance to run it.”</p>
</blockquote>



<p class="wp-block-paragraph">That’s free-riding on a competitive market for security research, and given the speed of today’s market, it’s an elegant solution. There’s a whole class of defenses like this waiting to be found, where the mechanism is not a technical system but a well-chosen delay.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="A 7-day Cooldown on Every Dependency" width="500" height="281" src="https://www.youtube.com/embed/j5yxGbCy5Ik?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">Data, and DJ Patil’s “Tidy House”</h2>



<p class="wp-block-paragraph">The problem we run into most often as we build AI workflows at O’Reilly isn’t the model or the tooling. It’s data. Who has access to which system, which system does that data live in, and who do I ask? In a 500 person company that’s annoying. I wonder what it’s like at a company with 50,000 employees.</p>



<p class="wp-block-paragraph">I told Dan about <a href="https://www.oreilly.com/radar/the-tidy-house/" target="_blank" rel="noreferrer noopener">DJ Patil’s Tidy House framing</a>. He agreed that data access for AI is a big problem. His answer starts with permissions:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“The permissions debt is invisible until an agent hits it. Making data agent legible is a forced permission audit. You have to actually go through and figure out who can access what…. It also raises the stakes for permissions errors. If you overshare information, now an agent inside your company is going to find it instantly. There are a lot of these technical debt sort of things where, with agents, all of it’s becoming due at the same time.”</p>
</blockquote>



<p class="wp-block-paragraph">Every shortcut an organization took with its data over the past twenty years is being called at once, and the companies that can run the audit, make fast decisions about boundaries, and then actually share their data are the ones that will get a force multiplier.</p>



<p class="wp-block-paragraph">Dan is against letting a thousand flowers bloom, because uncoordinated teams create overlap rather than compounding. He’d rather have one centralized foundation, with innovation happening on top of that. He suggested a useful metric for making that work across team boundaries is what fraction of your team’s data did you make reusable for everyone else, and how much of it is being used by teams outside your own.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="All of It’s Becoming Due at the Same Time" width="500" height="281" src="https://www.youtube.com/embed/-lNnlAnCqmQ?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">What post-AI jobs look like</h2>



<p class="wp-block-paragraph">Before the first hackathon, Trail of Bits ran hands-on sessions to teach its operations and go-to-market staff the basics of git and the command line. Not mastery, just enough to be a consumer of the thing. Here we are fifty years into my career and the Unix command line still matters. Dan’s non-technical staff mostly work inside Claude Cowork or Codex Desktop now, but he thinks the command line experience was worth it because they know what’s happening under the hood.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Every Non-Engineer Here Uses GitHub Every Day" width="500" height="281" src="https://www.youtube.com/embed/RiDFxlVHHUU?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">What happens to a job when the tool can do a lot of what humans used to do? Dan gave the example of his own technical editors. His editors used the hackathons to build the tools that got them out of line editing, including one that turns a public presentation into a blog post in the company’s voice. What the writers do now is consult on how to frame a story so it is effective with a particular audience.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="So What&amp;apos;s the Point of a Technical Editor?" width="500" height="281" src="https://www.youtube.com/embed/tdUsWlX7vC4?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">I agree. Human jobs aren’t going away any time soon. This gets heard as optimism when it’s really just observation. AI is going to replace a lot of what we used to do, but it is also going to hand us a large amount of new work, and much of that work hasn’t been understood yet. Quality assurance for agent systems is one of the new jobs. So is skills product management, which is a role that didn’t exist eighteen months ago and now has a headcount at a 130 person security firm.</p>



<p class="wp-block-paragraph">I asked a question towards the end about how we’re going to know which skills and agents are any good. What Dan has so far is telemetry pulled from developers’ dot files through the company’s device management system, which tells him what gets used and what breaks, plus one AI systems engineer whose job is product management for the skills repository, reviewing incoming pull requests and deprecating overlapping skills.</p>



<p class="wp-block-paragraph">What Dan thinks comes next is evaluation. He says: “Once you invest a lot into these agent systems, you need proof that they do the job. The way you do that is you give everybody a performance review. You give them an evaluation data set, a benchmark.”</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Give Your Agents a Performance Review" width="500" height="281" src="https://www.youtube.com/embed/WXuQEa6Ce00?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Trail of Bits is now building benchmarks for its core skills. How well can we find bugs in this language? How well can we write a statement of work? Constructing those datasets is real work, with positive and negative cases, and comparisons against the algorithmic tools that already exist.</p>



<h2 class="wp-block-heading">Put the reps in</h2>



<p class="wp-block-paragraph">I asked Dan for the top five mistakes he made. He said there was only one. “You need to allocate an appropriate amount of FAFO time. (That&#8217;s F Around and Find Out.) A product comes out on Friday. There’s no documentation for it. There’s no training guidance for it. There’s no course on it. You can’t wait until somebody systematizes the knowledge. You just need to do it.”</p>



<p class="wp-block-paragraph">Then he gave an analogy to going to the gym.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Put the Reps In" width="500" height="281" src="https://www.youtube.com/embed/NBubGIju9Bg?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">The recipe for success</h2>



<p class="wp-block-paragraph">Dan has a replicable recipe, which he summarized as follows:</p>



<ol class="wp-block-list">
<li>Standardize on one agent workflow that you can support.</li>



<li>Write an AI handbook so that risk decisions aren’t ad hoc, and that everyone is playing the same game.</li>



<li>Create a capability ladder that makes clear that improvement is expected.</li>



<li>Run short adoption sprints that force hands-on usage.</li>



<li>Capture everything as reusable artifacts: skills + configs + a curated supply chain.</li>



<li>Make autonomous agents safe with sandboxing + guardrails + hardened defaults.</li>
</ol>



<p class="wp-block-paragraph">The Trail of Bits skills repository is public. So is the curated marketplace, the configuration repository, the devcontainer, dropkit, and COOP (Continuity of Operations planning). He wrote up <a href="https://blog.trailofbits.com/2026/03/31/how-we-made-trail-of-bits-ai-native-so-far/" target="_blank" rel="noreferrer noopener">the whole playbook on <em>The Trail of Bits Blog</em></a> and gave <a href="https://tldrsec.com/p/how-we-made-trail-of-bits-ai-native-so-far" target="_blank" rel="noreferrer noopener">a version of it to <em>tl;dr sec</em></a>. He thinks publishing makes the work better because it forces the tribal knowledge out into the open where it can be checked.</p>



<p class="wp-block-paragraph">Which brings me back to the Solow paradox, which seemed to disappear by the late 90s, when US aggregate productivity did finally go up. That didn’t happen because computers got faster. It disappeared because companies figured out how to reorganize themselves around what computers could do, and eventually those organizational recipes spread widely enough to show up in aggregate statistics. The same has to happen today. The current AI discourse is obsessed with model capability and largely uninterested in diffusion. The problem is not that the models are oversold. It’s that almost nobody has done the necessary organizational work, and the few who have are mostly keeping it to themselves.</p>



<p class="wp-block-paragraph"><em>If you want to go beyond the highlight videos shown above, watch Dan’s entire talk <a href="https://learning.oreilly.com/videos/a-playbook-for/0642572388935/" target="_blank" rel="noreferrer noopener">here</a>.</em> <em>His slide deck is <a href="https://github.com/trailofbits/publications/blob/master/presentations/How%20we%20made%20Trail%20of%20Bits%20AI-Native%20%28so%20far%29/slides.pdf" target="_blank" rel="noreferrer noopener">here</a>.</em> <em>And be sure to check out <a href="https://trailofbits.com/?item=https-github-com-trailofbits-publications-blob-master-presentations-how-20we-20m" target="_blank" rel="noreferrer noopener">the Trail of Bits Github repository</a></em>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/ai-as-an-enterprise-operating-system/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: Agents, Gatekeepers, and World Models</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-agents-gatekeepers-and-world-models/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-agents-gatekeepers-and-world-models/#respond</comments>
				<pubDate>Fri, 31 Jul 2026 13:02:06 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19277</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-scaled.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2560" 
				height="2560" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Plus pressure on the open web and what publishers are doing about it]]></custom:subtitle>
		
				<description><![CDATA[This week, data and AI evangelist Christina Stathopoulos looked at three developments shaping AI’s next phase: agents that can act across systems, infrastructure built for specific models, and world models that help AI understand physical environments. Model quality is no longer the only constraint for teams. They also need to account for security controls, compute [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">This week, data and AI evangelist Christina Stathopoulos looked at three developments shaping AI’s next phase: agents that can act across systems, infrastructure built for specific models, and world models that help AI understand physical environments. Model quality is no longer the only constraint for teams. They also need to account for security controls, compute requirements, information access, and the environments where AI systems will operate.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="This Week in AI: Agents, Gatekeepers, and World Models with Christina Stathopoulos" width="500" height="281" src="https://www.youtube.com/embed/kumsRXbBbf4?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Agent capability is advancing faster than agent control</strong></h2>



<p class="wp-block-paragraph">Christina opened with reports that an OpenAI agent escaped a test environment, gained internet access, and <a href="https://www.bbc.com/news/articles/c3ek3gvdnj3o" target="_blank" rel="noreferrer noopener">targeted Hugging Face</a> while attempting to complete an assigned task. She also noted skepticism about how the incident was characterized, as well as the joint investigation announced by OpenAI and Hugging Face. The details remain under review, but the broader deployment problem is already familiar. Agents can combine tools, credentials, networks, and external services in ways application teams may not anticipate. (After the episode aired, OpenAI revealed that its review had turned up <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noreferrer noopener">four other similar incidents</a> “where the models identified and used publicly exposed credentials at the account-level on other publicly-available services.”)</p>



<p class="wp-block-paragraph">Christina then discussed <a href="https://openai.com/index/introducing-openai-presence/" target="_blank" rel="noreferrer noopener">OpenAI’s limited-availability platform</a> for helping enterprise customers build and manage agents with support from forward-deployed engineers. Direct access to specialists can help a company launch an agent, but it doesn’t replace the internal skills and governance required to operate one over time. For technical leaders, agent readiness increasingly means evaluating the full operating environment rather than focusing only on benchmark performance.</p>



<h2 class="wp-block-heading"><strong>AI infrastructure is reshaping both compute and the open web</strong></h2>



<p class="wp-block-paragraph">Google appeared on both sides of the infrastructure discussion. Christina covered reports of <a href="https://techcrunch.com/2026/07/20/google-is-working-on-a-new-ai-chip-designed-to-make-gemini-more-efficient/" target="_blank" rel="noreferrer noopener">a chip designed around Gemini’s architecture</a>, an approach that could reduce the compute required to run the model if the reported efficiency gains hold up. Specialized hardware has become a larger part of the AI race because model performance depends on cost, energy use, and deployment capacity. A model that performs well but consumes too much power or requires scarce hardware may still be difficult to use at scale.</p>



<p class="wp-block-paragraph">A different infrastructure shift is affecting the web. Christina examined how the growth of AI-first search experiences that answer questions without sending users to the sites that supplied the underlying material is threatening the open web. Organizations still pay to produce and host useful information, but AI systems collect more of it while returning less traffic. <a href="https://blog.cloudflare.com/agentic-internet-bot-report/" target="_blank" rel="noreferrer noopener">Cloudflare data</a> shows more traffic from agents, fewer human visitors, and declining referrals to publishers. More and more, people are using <a href="https://www.nytimes.com/2026/07/20/technology/google-ai-open-web.html" target="_blank" rel="noreferrer noopener">AI mode in Google search</a> instead of clicking through to websites, leading some to suspect the arrival of what is referred to as “Google Zero.”</p>



<p class="wp-block-paragraph">Developers building search products, retrieval systems, and agents should treat source attribution and publisher incentives as product design decisions. Reliable AI systems depend on reliable source material, and that source material needs a sustainable way to exist.</p>



<h2 class="wp-block-heading"><strong>World models could give physical AI a more useful foundation</strong></h2>



<p class="wp-block-paragraph">The episode closed with world models, systems designed to learn how environments work, how they change, and how actions affect what happens next. Christina highlighted a <a href="https://arxiv.org/abs/2607.06401" target="_blank" rel="noreferrer noopener">proposed research roadmap</a> that describes world models as able to combine several kinds of input, process information arriving at different speeds, and infer a larger environment from limited observations.</p>



<p class="wp-block-paragraph">For now, the clearest applications are in simulation, robotics, planning, and decision-making rather than claims about artificial general intelligence. A robot working in a factory, construction site, or emergency zone must track objects, understand movement, respond to incomplete information, and predict the likely result of an action. Large language models can support communication and planning, but physical work requires a representation of space, time, and cause and effect. World models may provide part of that foundation. However, researchers still need standardized definitions, reliable evaluations, and clear evidence that these systems can generalize beyond controlled environments.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">Across the episode, Christina explored how AI capability is advancing faster than the systems around it. Security practices, compute infrastructure, publishing economics, and physical-world evaluation will help determine which advances become dependable tools and which remain impressive demonstrations.</p>



<p class="wp-block-paragraph">Tune in next week as Christina breaks down the biggest AI news, including the US-China tech rivalry heating up after Anthropic CEO Dario Amodei&#8217;s post on open weight models and new bans on foreign-made humanoid robots. She&#8217;ll also challenge Sam Altman&#8217;s AI singularity claims, separating fact from hype, and examine key developments in math and science, including OpenAI&#8217;s 100,000 free researcher licenses, Claude Fable 5 solving an 87-year-old math problem, and Google disbanding its Nobel Prize-winning AlphaFold team to prioritize Gemini.</p>



<p class="wp-block-paragraph">Check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noreferrer noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noreferrer noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noreferrer noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-agents-gatekeepers-and-world-models/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Problem Is Prompt Debt</title>
		<link>https://www.oreilly.com/radar/the-problem-is-prompt-debt/</link>
				<comments>https://www.oreilly.com/radar/the-problem-is-prompt-debt/#respond</comments>
				<pubDate>Thu, 30 Jul 2026 11:05:15 +0000</pubDate>
					<dc:creator><![CDATA[Drew Breunig]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19266</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/The-problem-is-prompt-debt-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/The-problem-is-prompt-debt-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[You can’t be model agnostic if you’re hand-tuning prompts]]></custom:subtitle>
		
				<description><![CDATA[The following article was originally published on Drew Breunig’s blog and is being republished here with the author’s permission. Thanks to natural language interfaces, AI applications can be prototyped quickly. You write what you want in English, hand it to a frontier model, and a working prototype appears in an afternoon. This is extraordinarily powerful [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article was originally published on <a href="https://www.dbreunig.com/2026/06/22/the-problem-is-prompt-debt.html" target="_blank" rel="noreferrer noopener">Drew Breunig’s blog</a> and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Thanks to natural language interfaces, AI applications can be prototyped quickly. You write what you want in English, hand it to a frontier model, and a working prototype appears in an afternoon. This is extraordinarily powerful and for one-off tasks, optimal. But as a way to build reliable systems, the natural language prompt is a trap.</p>



<p class="wp-block-paragraph">The plain-English prompt that makes prototypes effortless turns out to be a poor way to specify how a system should behave, and the bill arrives slowly, disguised as ordinary progress, until the application can barely move. The problem is not any single prompt. It is that natural language was never meant to be a specification language for engineering, and treating it as one quietly caps what you can build.</p>



<h2 class="wp-block-heading">The prompt debt trap</h2>



<p class="wp-block-paragraph">The first symptom of prompt debt is slowing iteration. As users flag errors and spot edge cases, additional guidance is added to the instructions, nudging the model into line. If unwanted behaviors persist, instructions are repeated, with increasing severity. Pretty soon, the prompt isn’t straightforward and quick fixes regress previous instructions. Errors can no longer be handled with one-line “hot fixes” and your development cycle slows to a crawl.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="898" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38-1600x898.png" alt="Fable's system prompt repeats copyright guidance up to six times, under sections named search_instructions, search_usage_guidelines, mandatory_copyright_requirements, hard_limits, self_check_before_responding, and critical_reminders." class="wp-image-19267" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38-1600x898.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38-300x168.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38-768x431.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38-1536x862.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-38.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption">Fable’s system prompt repeats copyright guidance up to six times, under sections named <code>search_instructions</code>, <code>search_usage_guidelines</code>, <code>mandatory_copyright_requirements</code>, <code>hard_limits</code>, <code>self_check_before_responding</code>, and <code>critical_reminders</code>.</figcaption></figure>



<p class="wp-block-paragraph">Next, prompt debt incapacitates your team. Your brittle prompt <a href="https://www.dbreunig.com/2026/02/10/system-prompts-define-the-agent-as-much-as-the-model.html#:~:text=The%20Common%20Jobs%20of%20a%20Coding%20Agent%20System%20Prompt" target="_blank" rel="noreferrer noopener">full of edge cases and all-caps threats</a> is barely legible to you, and it’s downright impenetrable to your colleagues. Many teams mitigate this issue by breaking prompts into complicated templates assembled at run-time, each isolated to specific concerns. But these prompt segments evolve, too, growing into <a href="https://www.dbreunig.com/2026/04/04/how-claude-code-builds-a-system-prompt.html" target="_blank" rel="noreferrer noopener">a thicket of conditions</a>.</p>



<p class="wp-block-paragraph">Finally, prompt debt ties you to a single model. Your hot fixes work on GPT-4o, but fail in entirely new ways when you point your inference call at GPT-5.4-mini. So you stay with 4o, hope the increasingly frequent deprecation emails from your inference provider are empty threats, and forgo the possibility of potentially cheaper, faster, <em>better</em> models. A <a href="https://www.datadoghq.com/state-of-ai-engineering/">recent report from Datadog</a> suggests this is a common situation: The most-used model in traffic they observed is <em>GPT-4o</em>.<sup data-fn="c76a5ace-52f4-44e0-990e-fb1f64977b17" class="fn"><a href="#c76a5ace-52f4-44e0-990e-fb1f64977b17" id="c76a5ace-52f4-44e0-990e-fb1f64977b17-link">1</a></sup></p>



<p class="wp-block-paragraph">Any one of these issues is a nuisance, but together they are the difference between a glorified prototype and a product that can grow with you, your customers, and your business. Your shiny new AI features are frozen, can only be improved through a full rebuild, and are locked to an aging model.</p>



<h2 class="wp-block-heading">Why prompt debt happens</h2>



<p class="wp-block-paragraph">Natural language interfaces are wonderful. They’re the right mechanism for one-off tasks and broad conversational threads. We get into trouble when we rely on natural language to define durable system behavior.</p>



<p class="wp-block-paragraph">The imprecision of natural language paired with probabilistic language models means different words expressing the same intent can yield different outputs. <a href="https://arxiv.org/abs/2604.07709" target="_blank" rel="noreferrer noopener">In a recent study</a>, a clinical question asked in a patient’s voice and then re-asked in a physician’s, with identical facts, flipped Opus from declining all ten times to answering all ten.</p>



<p class="wp-block-paragraph">And it’s not only word choice that matters. Seemingly unrelated statements in the same prompt can affect results. <a href="https://arxiv.org/html/2407.06866v3" target="_blank" rel="noreferrer noopener">In a Harvard study</a>, researchers found that merely stating which NFL team the user rooted for changed how often the model refused to answer questions regarding sensitive topics. Spurious statements influence the inference pass in ways we can’t predict. Which is why prompts become more brittle as you add fixes. An additional instruction to quell a stubborn error could affect how the model interprets a separate instruction that worked yesterday.</p>



<p class="wp-block-paragraph">Repeating instructions propels us towards prompt debt, but it’s necessary when the behavior we want is at odds with a model’s training. This is <a href="https://www.dbreunig.com/2025/11/11/don-t-fight-the-weights.html" target="_blank" rel="noreferrer noopener">fighting the weights</a>, and once you recognize it you see it in system prompts everywhere. For example, ChatGPT’s image prompts used <a href="https://www.dbreunig.com/2025/11/11/don-t-fight-the-weights.html#:~:text=When%20you%20asked%20ChatGPT%20to%20generate%20an%20image%2C%20it%20would%20clean%20up%20or%20even%20improve%20your%20image%20prompt%2C%20create%20the%20image%2C%20then%20append%20the%20following%20instructions:" target="_blank" rel="noreferrer noopener">to instruct the LLM <em>eight times</em> to not reply when a generated image was returned</a> because it had been trained to always keep the conversation going.</p>



<p class="wp-block-paragraph">Every coding agent system prompt we analyzed featured repeated instructions, stern warnings, and all-caps demands. <a href="https://blog.nilenso.com/blog/2026/02/12/how-system-prompts-reveal-model-biases/" target="_blank" rel="noreferrer noopener">Claude Code tells Opus <em>seven times</em></a> <a href="https://blog.nilenso.com/blog/2026/02/12/how-system-prompts-reveal-model-biases/" target="_blank" rel="noreferrer noopener">to return multiple tool calls in a single response</a>. And even the most advanced models force prompt authors to fight the weights: <a href="https://github.com/asgeirtj/system_prompts_leaks/blob/main/Anthropic/claude-fable-5.md" target="_blank" rel="noreferrer noopener">Fable’s leaked system prompt restates one specific copyright rule six times</a>.</p>



<p class="wp-block-paragraph">None of these examples occurred in isolation. Multiple repeated rules are woven throughout the system prompts we examine. Stubborn errors grow our prompts quickly, with each increasing the brittleness, the risk of regression with every edit.</p>



<p class="wp-block-paragraph">And worse: These fixes are tailored to a single model’s behavior. A recent <a href="https://arxiv.org/abs/2512.04123" target="_blank" rel="noreferrer noopener">Berkeley-led study</a> found enterprises stay on older models because newer ones break their existing agents. This is because models are not cleanly versioned software. They have different weights that produce different behaviors, in unpredictable and undocumented ways. A prompt that works beautifully with GPT-4o may fail with GPT-5.5. <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5#:~:text=Refactor%20existing%20prompts%20and%20skills" target="_blank" rel="noreferrer noopener">Anthropic’s own release notes for Fable</a> warn that skills developed for prior models can “degrade output quality.”</p>



<p class="wp-block-paragraph">Prompt debt locks an application to a single model. Our inability to easily swap models isn’t the result of frontier labs coming up with a clever moat. No, it’s the result of evolving a lossy natural language specification against a probabilistic model.</p>



<h2 class="wp-block-heading">Preventing prompt debt</h2>



<p class="wp-block-paragraph">Thankfully, we don’t have to theorize about how to mitigate prompt debt; one field has already shown the way. Programmers using coding agents sit at the leading edge of what models can do, outliers on the <a href="https://www.oneusefulthing.org/p/the-shape-of-ai-jaggedness-bottlenecks" target="_blank" rel="noreferrer noopener">jagged frontier</a> of model abilities. Over the last couple years they’ve <a href="https://addyosmani.com/blog/new-sdlc-vibe-coding/" target="_blank" rel="noreferrer noopener">been</a> <a href="https://simonwillison.net/guides/agentic-engineering-patterns/" target="_blank" rel="noreferrer noopener">evolving</a> <a href="https://developers.openai.com/codex/learn/best-practices" target="_blank" rel="noreferrer noopener">best</a> <a href="https://code.claude.com/docs/en/best-practices" target="_blank" rel="noreferrer noopener">practices</a> that let the model write more of the code, while delivering maintainable, modular software.</p>



<p class="wp-block-paragraph">The first principle is to specify your system’s behavior with measurements, not prose. When the model’s output is probabilistic and language is imprecise, we build hard edges to constrain them: evaluations, metrics, and typed specifications. These are legible, shared artifacts colleagues can read and contribute to, enabling the collaboration that brittle prompts prevented.</p>



<p class="wp-block-paragraph">The best engineers now spend more of their bandwidth on tests than ever, as they are no longer a safety net but the thing that <em>lets the model cook</em>.</p>



<p class="wp-block-paragraph">The second principle is to stop writing the prompt by hand. Once we have metrics that can score candidates, the prompt is no longer something to craft but something for which to search. And the surface area of potential words, phrases, and structures that natural language allows is too vast to spend human hours on. This is terrain LLMs were built to explore, and there are already systems (like <a href="https://dspy.ai/" target="_blank" rel="noreferrer noopener">DSPy</a> and <a href="https://sky.cs.berkeley.edu/project/gepa/" target="_blank" rel="noreferrer noopener">GEPA</a>) that manage this work for you, holding prompts accountable to your designs.</p>



<p class="wp-block-paragraph">Once prompts are generated and your program’s behavior is defined by measurements, you are no longer bound to a particular model. Evaluating a new model takes hours, not weeks. When a faster, cheaper model arrives you can try it. When a deprecation email arrives, you can secure options in a day. Whether a model is pulled for regulatory reasons (<a href="https://www.theverge.com/ai-artificial-intelligence/949553/anthropic-fable-5-mythos-5-government-national-security" target="_blank" rel="noreferrer noopener">as we saw with Anthropic’s Fable</a>) or deprecated due to age (<a href="https://www.reuters.com/world/china/us-holds-off-blacklisting-chinas-deepseek-more-than-100-firms-deemed-security-2026-06-17/" target="_blank" rel="noreferrer noopener">as Groq announced last week with Llama-3.1-8b</a>), the fix is a chore, not a fire drill.</p>



<p class="wp-block-paragraph">Every mature engineering discipline eventually stops doing by hand the very thing it once prided itself on doing by hand. Assembly gave way to compilers, hand-tuned queries gave way to planners, and manual memory management gave way (mostly) to machines that do it better. Prompt-writing is no different.</p>



<p class="wp-block-paragraph">Coaxing the model with exactly the right words is a real skill, and for one-off tasks it’s often optimal. But to build reliable, improvable, and portable systems we should not be hand-tuning prompts.</p>



<h3 class="wp-block-heading">Footnote</h3>


<ol class="wp-block-footnotes"><li id="c76a5ace-52f4-44e0-990e-fb1f64977b17">This stat from Datadog is from March of this year, so GPT-4o concentration has likely dropped a bit. However, I’ve heard from multiple large inference providers that usage of GPT-4o and models of similar vintage can be higher than <em>50%</em> of all calls! <a href="#c76a5ace-52f4-44e0-990e-fb1f64977b17-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-problem-is-prompt-debt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>What the Hell Is a Loop, Anyway?</title>
		<link>https://www.oreilly.com/radar/what-the-hell-is-a-loop-anyway/</link>
				<comments>https://www.oreilly.com/radar/what-the-hell-is-a-loop-anyway/#respond</comments>
				<pubDate>Wed, 29 Jul 2026 10:38:24 +0000</pubDate>
					<dc:creator><![CDATA[Laurie Voss]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19251</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/What-the-hell-is-a-loop-anyway.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/What-the-hell-is-a-loop-anyway-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[The AI engineering world adopted a new favorite word this month, and it means at least four different things.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on LinkedIn and is being republished here with the author’s permission. We’re currently at the peak of the hype cycle. On June 7, Peter Steinberger posted that you shouldn’t be prompting coding agents anymore; you should be designing loops that prompt your agents. That same week, Boris Cherny of Anthropic [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on <a href="https://www.linkedin.com/pulse/what-hell-loop-anyway-laurie-voss-ldmdc/" target="_blank" rel="noreferrer noopener">LinkedIn</a> and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">We’re currently at the peak of the hype cycle. On June 7, Peter Steinberger posted that <a href="https://x.com/steipete/status/2063697162748260627?lang=en" target="_blank" rel="noreferrer noopener">you shouldn’t be prompting coding agents anymore</a>; you should be designing loops that prompt your agents. That same week, <a href="https://www.linkedin.com/in/bcherny/" target="_blank" rel="noreferrer noopener">Boris Cherny</a> of Anthropic said on stage that he doesn’t prompt Claude anymore: “I write loops; <a href="https://x.com/sairahul1/status/2064279904989147577?lang=en" target="_blank" rel="noreferrer noopener">the loops do the work</a>.” Addy Osmani published an essay called “<a href="https://addyosmani.com/blog/loop-engineering/">Loop Engineering</a>” on June 7, swyx published “<a href="https://www.latent.space/p/ainews-loopcraft-the-art-of-stacking" target="_blank" rel="noreferrer noopener">Loopcraft: The Art of Stacking Loops</a>” on June 12, and LangChain published “<a href="https://www.langchain.com/blog/the-art-of-loop-engineering" target="_blank" rel="noreferrer noopener">The Art of Loop Engineering</a>” on June 16. Then came the AI Engineer World’s Fair, where the word <a href="https://www.latent.space/p/aiewf-daily-dispatch-loops" target="_blank" rel="noreferrer noopener">dominated the main stage</a>. Swyx’s keynote was about Loopcraft, an entire track was devoted to software factories, speaker after speaker reached for the same word, and the conference closed on July 2 with an hour-long debate about whether the hype behind loops has outrun what works in practice.</p>



<p class="wp-block-paragraph">The problem is that the people talking about loops aren’t all discussing the same thing. I counted at least four distinct architectures hiding behind that one word. So this post is an attempt to map out what everyone means.</p>



<h2 class="wp-block-heading">The execution loop: The agent’s own act-observe cycle</h2>



<p class="wp-block-paragraph">This is the loop most people picture when they say “agent”: call a tool, read the result, decide the next action, and repeat until there are no more tool calls to make. It’s what Addy calls the inner execution loop, the part agents can now run largely on their own, and it’s the innermost loop you can engineer. (swyx’s stack has a token loop, but nobody designs the token loop. It’s just part of the model.)</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1100" height="619" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-30.png" alt="Loopcraft: The art of stacking loops" class="wp-image-19252" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-30.png 1100w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-30-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-30-768x432.png 768w" sizes="auto, (max-width: 1100px) 100vw, 1100px" /><figcaption class="wp-element-caption"><em>Swyx’s original Loopcraft diagram</em></figcaption></figure>



<p class="wp-block-paragraph">The execution loop iterates on steps within one task. It ends on environment feedback: the test output, the API response, and the file contents. Humans are usually absent mid-loop and appear at the boundaries, approving plans or reviewing results. The execution loop also ends whenever the agent decides it’s done, whether or not it actually is. The first fix the field found for that was to wrap this loop in another one that doesn’t take the agent’s word for it.</p>



<h2 class="wp-block-heading">The task loop: Restart the agent until the spec is satisfied</h2>



<p class="wp-block-paragraph">This was the first loop to get a name and it’s Geoffrey Huntley’s Ralph loop, which got name-checked from the AI Engineer World’s Fair main stage when Allie Howe of Keycard introduced the software factories track by citing Geoffrey’s article “<a href="https://ghuntley.com/loop/" target="_blank" rel="noreferrer noopener">Everything Is a Ralph Loop</a>.” A Ralph loop restarts a coding agent against the same specification over and over, allocating a completely fresh context window every iteration and doing exactly one task per loop. The apparent waste is the point: Refeeding the full spec each time prevents the context rot and compaction events that quietly degrade long-running sessions.</p>



<p class="wp-block-paragraph">What this loop iterates on is a single artifact. What ends the loop is spec compliance and passing tests. The human writes the spec and judges doneness, and in Geoffrey’s telling the human has one more job that I’ll return to later: watching the loop, spotting failure patterns, and fixing them so they never recur. In the closing debate on the conference’s final day, he compared the role to a locomotive engineer, someone whose whole job is keeping the train on the rails. Zoom out from a single spec though, and a much bigger loop comes into view: the one that runs an entire codebase.</p>



<h2 class="wp-block-heading">The product loop: The software factory</h2>



<p class="wp-block-paragraph">This was the loudest version at the AI Engineer World’s Fair. Tereza Tizkova of Factory defined a software factory as “the whole loop, the whole lifecycle of developing software with autonomy,” and Zach Lloyd of Warp got specific about what that lifecycle is in an <a href="https://www.latent.space/p/aiewf-daily-dispatch-loops" target="_blank" rel="noreferrer noopener">interview with <em>Latent Space</em></a>: triage, specification, implementation, review, verification, shipping, and monitoring. Zach’s claim is that software engineering becomes factory engineering, and that you’ll be building the thing that builds the product. Warp is dogfooding this: The company placed its own open-sourced repo under the control of Oz, its factory platform. Zach describes the adoption path as starting with low-risk repos and ratcheting the automatic PR merge rate upward from 20 percent toward 60. Anthropic appears to be running the same experiment internally. The company says <a href="https://www.anthropic.com/news/introducing-claude-tag" target="_blank" rel="noreferrer noopener">65% of its product team’s code</a> is now created by its internal version of Claude Tag, and Mike Krieger described his team’s use of it at the World’s Fair as delegated and proactive: not “fix this bug” but take responsibility for this part of the codebase, monitor this feedback channel, and pick up tasks on your own.</p>



<p class="wp-block-paragraph">The task loop and the execution loop have defined exit conditions. The product loop iterates on a codebase and its backlog, continuously, and its closing signals come from outside the codebase entirely: new issues, production logs, user feedback, review outcomes. The human role becomes configurable. In Zach’s framing, you pick the parts of the lifecycle to automate and the points where humans get brought in, and organizations differ on questions like whether code review stays human for high-risk changes. A factory improves a product. The next loop improves the factory itself.</p>



<h2 class="wp-block-heading">The system loop: Autoresearch</h2>



<p class="wp-block-paragraph">Roland Gavrilescu of Introspection calls this autoresearch. Here’s how he framed the concept in a <a href="https://www.latent.space/p/autoresearch-introspection" target="_blank" rel="noreferrer noopener"><em>Latent Space</em> interview</a>: The inner loop is your primary system doing user-facing work, and the outer loop studies and maintains the primary system. It iterates on prompts, harnesses, model choices, and the evals themselves. His one-liner is that the loop is the product.</p>



<p class="wp-block-paragraph">This pattern now has real existence proofs at both ends of the scale. The minimal case is Andrej Karpathy’s autoresearch from March 2026, roughly 630 lines of Python that ran 50 hypothesis-edit-evaluate experiments overnight on one GPU. The shipped case is Meta’s Brain2Qwerty v2, <a href="https://ai.meta.com/blog/brain2qwerty-brain-ai-human-communication/" target="_blank" rel="noreferrer noopener">announced in late June</a>, where the researchers report that agents iteratively modified the codebase to invent better decoding architectures, producing a substantial improvement in word error rate. Meta’s caveat is instructive: Final training configurations were still selected by hand. Even the flagship system loop keeps a human at the last checkpoint.</p>



<p class="wp-block-paragraph">What ends this loop is the most demanding signal set of the four: evals, judges, filtered product feedback, and, in Roland’s design, an explicit ask-a-human tool through which the agent accumulates tacit knowledge the way a new employee does. And that’s the top of the stack. Put the four together and the shape of the whole system becomes visible.</p>



<h2 class="wp-block-heading">The four loops side by side</h2>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1568" height="642" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-31.png" alt="" class="wp-image-19253" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-31.png 1568w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-31-300x123.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-31-768x314.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-31-1536x629.png 1536w" sizes="auto, (max-width: 1568px) 100vw, 1568px" /></figure>



<h2 class="wp-block-heading">What about Agentic MapReduce?</h2>



<p class="wp-block-paragraph">One famous pattern from the same week is missing from this map on purpose. Cognition’s <a href="https://cognition.com/blog/introducing-devin-security-swarm" target="_blank" rel="noreferrer noopener">Devin Security Swarm</a> fans parallel bounded agents out across a repository and aggregates their findings, a shape the company calls Agentic MapReduce, and it gets called a loop. I don&#8217;t think it is one. Dispatch, gather, validate is a pipeline: Nothing feeds back into a next cycle, and a loop without feedback is just a for statement. Fan-out is a topology you can deploy inside any of the four loops, not a loop of its own.</p>



<h2 class="wp-block-heading">The unnamed loop at the top is the oversight loop</h2>



<p class="wp-block-paragraph">In swyx’s loop diagram, the outermost ring, the one above the loop that makes loops, is literally labeled “???? loop.” Its verbs are “set goals, allocate, cull.” Its exit condition is listed as none.</p>



<p class="wp-block-paragraph">I think that loop has a name. I’m calling it the oversight loop: It’s where goals get set, budgets get allocated, and work gets culled, and it’s the one ring where a human should live. Addy said on the AIEWF stage: “That inner loop is capability. The outer loop is agency.” Agency is exactly what the oversight loop holds.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1406" height="1000" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-32.png" alt="The loop stack, tidied up a bit." class="wp-image-19254" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-32.png 1406w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-32-300x213.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-32-768x546.png 768w" sizes="auto, (max-width: 1406px) 100vw, 1406px" /><figcaption class="wp-element-caption"><em>The loop stack, tidied up a bit.</em></figcaption></figure>



<p class="wp-block-paragraph">And the sharpest disagreements at AIEWF were all, once you translate them, arguments about who runs that top ring. Zach and Roland make the case for turning the dial up: pick your checkpoints deliberately, ratchet autonomy as trust accumulates, and, in Roland’s memorable distinction, build orchestras before factories, where an orchestra is a system that keeps a human conductor. The other camp says the dial has a stop. Geoffrey Litt of Notion called factories a depressing vision on X and argued, in a talk he has since <a href="https://www.geoffreylitt.com/2026/07/02/understanding-is-the-new-bottleneck.html" target="_blank" rel="noreferrer noopener">published as an essay</a>, that those who delegate understanding get replaced by the agent. Paul Bakaus <a href="https://www.latent.space/p/skill-engineering-design" target="_blank" rel="noreferrer noopener">put it as flatly as it can be put</a>: “There is no auto, and there will be no auto.” His argument isn’t only about quality; it’s about ownership. People need purpose, and they want a role in what they create.</p>



<p class="wp-block-paragraph">The closing debate, covered in <em>Latent Space</em>’s conference reporting, put both positions on one stage. Dex Horthy of HumanLayer took pains to say he isn’t anti-loop, pointing out that Kubernetes is built on control loops, but deterministic ones. His worry is that enthusiasm has gotten ahead of the engineering, and his advice was to step down an abstraction level rather than up. Geoffrey took the other side and called loops inevitable. And Mike offered the most honest data point of all: Even inside Anthropic, the team running Tag reports being bottlenecked on reviews and on the human ability to conceptualize what the system is doing. The checkpoint humans kept for themselves is now the constraint.</p>



<p class="wp-block-paragraph">Autonomy is a dial that exists separately on every one of the four loops. You can run a fully autonomous execution loop inside a heavily supervised product loop. You can hand the system loop to agents while keeping goal-setting entirely human. The interesting engineering question isn’t “Which camp wins?”; it’s “What information do you need to set each dial correctly?”</p>



<p class="wp-block-paragraph">The table above is my attempt to fill in those blanks. Every loop, including the top one, has a nameable exit condition, and the top one is you. But naming a signal isn’t the same as wiring it in. A loop without its signal doesn’t converge. It just runs until something external stops it. Knowing whether your loops are actually closing, at production scale, means sweeping traces and clustering failures continuously instead of spot-checking transcripts, which is exactly the job <a href="https://arize.com/?utm_source=lvoss&amp;utm_medium=linkedin&amp;utm_campaign=devrel&amp;utm_content=What%20the%20hell%20is%20a%20loop%20anyway" target="_blank" rel="noreferrer noopener">Arize AX</a> was built to do.</p>



<h2 class="wp-block-heading">Which one are you building?</h2>



<p class="wp-block-paragraph">Now the loops have names, that’s the question to ask. The word loop is doing a lot of work this month, because this field loves nothing more than jumping on the next hot thing. But real practice underlies all four loops, and it’s the same practice in each: people are dialing up their level of abstraction and pushing human judgment further up the stack. That’s the actual lesson of loops. We get more done by climbing up the stack, and now you have a map, you know where you should climb.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/what-the-hell-is-a-loop-anyway/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Teaching Coding When AI Can Write the Code</title>
		<link>https://www.oreilly.com/radar/teaching-coding-when-ai-can-write-the-code/</link>
				<comments>https://www.oreilly.com/radar/teaching-coding-when-ai-can-write-the-code/#respond</comments>
				<pubDate>Tue, 28 Jul 2026 12:54:30 +0000</pubDate>
					<dc:creator><![CDATA[Eric Freeman]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19242</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/Teaching-code-when-AI-can-write-the-code-658068.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/Teaching-code-when-AI-can-write-the-code-658068-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Now that AI can write code, a student&#039;s submission no longer shows what they really understand. Here, I share three approaches we&#039;re testing at the University of Texas at Austin.]]></custom:subtitle>
		
				<description><![CDATA[For as long as we’ve taught programming, the student’s code has provided a window into the students’ thinking. Errors, the code structure, the awkward working solution—all of it showed how someone reasoned and where they got stuck. It was never a clean window. Students have always copied, crammed, and borrowed, sometimes turning in work they [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">For as long as we’ve taught programming, the student’s code has provided a window into the students’ thinking. Errors, the code structure, the awkward working solution—all of it showed how someone reasoned and where they got stuck.</p>



<p class="wp-block-paragraph">It was never a clean window. Students have always copied, crammed, and borrowed, sometimes turning in work they didn’t fully understand. But the code still left clues. Generative AI has changed that: A finished program now tells us more about a student’s prompts than their ideas. And here’s the part that should unsettle us—often, the better the code looks, the less we can say about what the student actually learned.</p>



<p class="wp-block-paragraph">This raises a bigger question: If AI can write code, should we still teach coding? I believe the answer is yes, at least for some students and situations. But that’s another topic. Here, I want to focus on the next step: If we continue teaching coding in a world with AI, how can we know if students are really learning?</p>



<p class="wp-block-paragraph">Some schools have responded by trying to catch students. They use AI detectors, surveillance tools, locked-down browsers, stricter rules, and clearer honor codes. This has also led to more suspicion.</p>



<p class="wp-block-paragraph">Some of these responses make sense. Teachers want to protect learning, and schools want to keep things fair. But using detection as the main way to assess students is weak. Stanford researchers found that popular AI detectors often falsely flagged writing by nonnative English speakers, with 61.22% of TOEFL essays in one study marked as AI-generated. OpenAI even retired its own AI Text Classifier in 2023 because it wasn’t accurate enough. If the company that created the tool can’t reliably detect AI, it’s probably not a good idea to base your honor code on it.</p>



<p class="wp-block-paragraph">But detection isn’t the real issue. Even if we had a perfect detector, we’d still be asking the wrong question. Instead of asking, “How do we stop students from using AI?” we should ask, “How do we teach coding in a world with AI, making use of its benefits, while still being able to see if students are learning?”</p>



<h2 class="wp-block-heading"><strong>Borrowing from the studio</strong></h2>



<p class="wp-block-paragraph">We’re seeing this challenge with students at AET, the Arts and Entertainment Technologies Department at the University of Texas at Austin. Although my usual home is Computer Science, it so happens that AET is within the College of Fine Arts at UT, which offers many other ways to learn and assess: studio work, critique, rehearsal, revision, and performance.</p>



<p class="wp-block-paragraph">In the arts, the final piece has never been the whole story. A painting doesn’t explain the choices behind it. A performance doesn’t reveal the rehearsals. A design board doesn’t show the discarded versions. A composition doesn’t tell you where the student struggled or what they finally learned to hear.</p>



<p class="wp-block-paragraph">Art education has developed practices that focus on visible progress. Students bring in sketches and drafts, discuss influences, revisions, and failures, and rehearse, perform, and critique each other’s work while it’s still in progress.</p>



<p class="wp-block-paragraph">At AET, we teach creative coding, which means programming to create art, design, games, or experiences. That doesn’t mean coding for poets. Our students—game designers, web developers, and programmers—start from scratch and learn advanced concepts in tools like Processing and p5.js. In the creative coding tradition, a program is often called a <em>sketch, </em>borrowing the term from the art world. It means something temporary, exploratory, and open to change—something you make, test, revise, and share.</p>



<p class="wp-block-paragraph">So in creative coding, we were already leaning toward the studio model of sketches, experiments, iterations, and critique. Now we’re pushing that further as we rethink how we teach coding in an AI world. Here are three things we’re already using or actively developing.</p>



<h2 class="wp-block-heading"><strong>Make the work public</strong></h2>



<p class="wp-block-paragraph">We run the class like a studio. It’s not that work never happens at home, but the most important work needs to be seen in the classroom. Students show their code, including false starts, revisions, the choices they made, and the reasons behind them. Assignments are no longer just things you submit—they become projects you develop in public.</p>



<p class="wp-block-paragraph">AI isn’t banned from the classroom. Instead, it’s treated as a helpful assistant to learn from. Students share prompts and techniques. They use AI, Google, Stack Overflow, classmates, or any other resources.</p>



<p class="wp-block-paragraph">But you still need to take responsibility for your work. If you submit or present it, you must explain what the code does, why you made those choices, and how it works. If I need to ask your AI to understand your code, something is wrong. Getting help is fine, but hiding behind that help is not.</p>



<p class="wp-block-paragraph">You can’t outsource to AI what the whole room watched you build.</p>



<p class="wp-block-paragraph">A real studio needs students talking out loud together in the room every day. This also helps with another issue that isn’t about AI. Many people say students today are quieter than in the past. While this is mostly based on stories rather than long-term studies, these stories are common and consistent. Faculty on all types of campuses talk about silent classrooms and students who hesitate to speak up, especially since 2020.</p>



<p class="wp-block-paragraph">Whatever the reason, this silence can be changed, and the solution is the same as for AI challenges: encourage students to participate. Communication is one of the most important skills in any career, including explaining ideas, defending choices, and persuading others in real time. Students don’t develop these skills by just submitting AI-guided work online. When they share their work publicly, it not only prevents AI misuse but also helps them build the skills they need most.</p>



<h2 class="wp-block-heading"><strong>Invert the roles: AI as teacher and assessor</strong></h2>



<p class="wp-block-paragraph">We know the usual pattern: A student asks, AI answers, and the student copies. We’ve tried to invert this. In our new approach, the AI works with the student on a set of topics, engages them in a conversation they must navigate, and ultimately assesses how well they understand the material, which leads to a grade.</p>



<p class="wp-block-paragraph">This idea has a research background that goes back before ChatGPT. Teachable-agent systems like <a href="https://bettysbrain.teachableagents.org/front-page/about" target="_blank" rel="noreferrer noopener">Betty’s Brain</a> showed that explaining—even to a software agent—forces students to organize their knowledge, make connections clear, and find gaps. Our model uses this insight differently. The student isn’t teaching the bot. Instead, the student is having a conversation with it, learning, discussing, debating, and showing what they understand.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="1413" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29-1600x1413.png" alt="The Vera Molnár chatbot at the University of Texas at Austin" class="wp-image-19243" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29-1600x1413.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29-300x265.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29-768x678.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29-1536x1356.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-29.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>The Vera Molnár chatbot at the University of Texas at Austin</em></figcaption></figure>



<p class="wp-block-paragraph">How did we do this? With fairly simple prompt engineering, we created an avatar chatbot of <a href="https://en.wikipedia.org/wiki/Vera_Moln%C3%A1r" target="_blank" rel="noreferrer noopener">Vera Molnár</a> (1924–2023), a pioneer of algorithmic art. The bot takes on Molnár’s role, drawing students into conversations about randomness, computation, generative art, and creative choices. Her practice sits exactly where creative coding students need to think: between rule and variation, system and choice, computation and visual judgment.</p>



<p class="wp-block-paragraph">A system prompt sets the topics and types of questions to ask. The bot goes through these with the student, asks for more detail on unclear answers, and keeps following up until there is proof of understanding. At the end, it reviews the conversation against a rubric, giving us a clear record of which ideas the student covered, where they struggled, and how well they improved.</p>



<p class="wp-block-paragraph">Besides the assessment, which is often accurate, the transcript becomes a different kind of proof, showing what a typical assignment might hide. What did the student notice? What did they misunderstand? Could they connect the concept to the code? Could they defend their choices? Could they revise their explanation when challenged?</p>



<p class="wp-block-paragraph">When we switch the roles, something surprising appears: the one thing a finished submission can’t show.</p>



<p class="wp-block-paragraph">A student thinking out loud.</p>



<h2 class="wp-block-heading"><strong>Make understanding performative: Make students perform</strong></h2>



<p class="wp-block-paragraph">Programming has never really had a tradition of performance. Musicians have it, painters have it, and dancers have it. Live coding is starting to change that.</p>



<p class="wp-block-paragraph">Every semester at AET, students from different disciplines stage an algorave together—short for <em>algorithmic rave</em>. Audio sets, projection pieces, game demos, lasers, drones, experience design. The creative coding class brings live visuals into the live-coding tradition: Code is written and modified in real time, the screen is projected, and the audience watches the editor change as the visuals respond to the music other students are playing.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1280" height="720" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33.jpeg" alt="The Department of Arts and Entertainment Technologies’ annual AudioPixel Collider algorave, November 20, 2025, B. Iden Payne Theatre, The University of Texas at Austin" class="wp-image-19244" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33.jpeg 1280w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-300x169.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-33-768x432.jpeg 768w" sizes="auto, (max-width: 1280px) 100vw, 1280px" /><figcaption class="wp-element-caption"><em>The Department of Arts and Entertainment Technologies’ annual AudioPixel Collider algorave, November 20, 2025, B. Iden Payne Theatre, The University of Texas at Austin</em></figcaption></figure>



<p class="wp-block-paragraph">No prerender. No hiding the machinery.</p>



<p class="wp-block-paragraph">The <a href="https://toplap.org/wiki/ManifestoDraft" target="_blank" rel="noreferrer noopener"><em>Live Coding</em> manifesto</a>, written in 2004 by TOPLAP, includes a line that fits every AI-era assessment conversation: “Obscurantism is dangerous. Show us your screens.” This is not just a performance ethic; it’s also an assessment strategy.</p>



<p class="wp-block-paragraph">A student walks on stage. The projected screen is their editor. The room can read it. The music starts. And they build up a line of code on screen like:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><code>osc(18, 0.08, 1.2)<br><br> .modulate(noise(3), 0.25)<br><br> .rotate(() => time * 0.1)<br><br> .out()</code></p>
</blockquote>



<p class="wp-block-paragraph">This is JavaScript building visuals in real time. FFTs, chained functions, higher-order manipulations. When you’re manipulating code like that on stage, you’d better know what you’re doing.</p>



<p class="wp-block-paragraph">AI can help you prepare. Good. Let it.</p>



<p class="wp-block-paragraph">But once you’re on stage, the question shifts from “Can you copy and paste code?” to “Can you control it?” You can paste code into a file, but you can’t paste your way through three minutes of public debugging while the whole projection turns into a beige rectangle. In a live build, understanding has nowhere to hide.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1280" height="720" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-34.jpeg" alt="Student livecoding at the Department of Arts and Entertainment Technologies’ annual AudioPixel Collider algorave, November 20, 2025, B. Iden Payne Theatre, The University of Texas at Austin" class="wp-image-19245" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-34.jpeg 1280w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-34-300x169.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-34-768x432.jpeg 768w" sizes="auto, (max-width: 1280px) 100vw, 1280px" /><figcaption class="wp-element-caption"><em>Student livecoding at the Department of Arts and Entertainment Technologies’ annual AudioPixel Collider algorave, November 20, 2025, B. Iden Payne Theatre, The University of Texas at Austin</em></figcaption></figure>



<p class="wp-block-paragraph">Can you read the code, make changes on purpose, and recover when something unexpected happens? That’s fluency: knowing what to do next while the system is still running.</p>



<p class="wp-block-paragraph">It is very hard to plagiarize panic.</p>



<h2 class="wp-block-heading">A note on assessment</h2>



<p class="wp-block-paragraph">So far, our results are based on our own observations. We haven’t conducted a controlled study or compared different groups, so what we have seen might just be early variation rather than patterns that apply more broadly. For now, these efforts are experiments, not final answers.</p>



<p class="wp-block-paragraph">Assessment in studio and live performance settings is always subjective and focused on people. It relies on monitoring students’ progress, providing feedback, and observing how they handle challenges. We do not plan to change this core approach.</p>



<p class="wp-block-paragraph">For the Molnár conversation assignment, students discussed Molnár using an AI system. The AI then created a summary and analysis of each student’s understanding. Teaching assistants reviewed this analysis, conducted their own assessments, and assigned grades. In our small experiments, the AI’s assessments using the rubric matched closely with the teaching assistants’ own evaluations.</p>



<p class="wp-block-paragraph">We also used AI to help grade the end-of-term coding assignment. In this project, students improved an object-oriented game by adding strategies like heuristics, search algorithms, and learned behaviors. Since our teaching assistants had limited experience with object-oriented programming, we developed a detailed rubric and had an AI model use it to evaluate each submission. The AI’s analysis was given to the teaching assistants as support. It helped them see how each project was structured, spot important OOP design choices, and use the rubric with more confidence. The teaching assistants still made their own grading decisions. I was available as the OOP expert for any questions they could not answer. From what I observed, this substantially helped the teaching assistants understand and grade the students’ OOP design work.</p>



<p class="wp-block-paragraph">More broadly, both approaches appear to enable substantive feedback at a scale that would otherwise be difficult given our current student-to-teaching-assistant ratios.</p>



<h2 class="wp-block-heading"><strong>The process is the proof</strong></h2>



<p class="wp-block-paragraph">We spent the first two years of the generative AI panic asking how to catch students using AI—or prohibit it altogether. Wrong question.</p>



<p class="wp-block-paragraph">The real question is whether the assignment gives students a real way to show and develop their understanding. This view isn’t limited to educators. NVIDIA CEO Jensen Huang recently argued that students should not focus on finding an “AI-proof” subject. Instead, he suggested they consider how AI can help them learn more deeply and develop their skills and sense of purpose. He highlighted storytelling, creativity, design, and judgment as abilities that will stay important even as AI takes over more tasks. This supports a key idea in coding education: The aim is not to prove you didn’t use any tools, but to help students show how they think, make choices, revise, and take responsibility for their work.</p>



<p class="wp-block-paragraph">These three practices are experiments, not universal solutions. They work especially well in creative coding, where code already has a public, visual, and performative aspect. But they suggest a broader principle: As finished work becomes easier to generate, assessment needs to focus more on process, explanation, revision, and mastery.</p>



<p class="wp-block-paragraph">This matters outside of school too. A polished memo no longer proves there was real thinking behind it. A working prototype no longer proves product sense. A passing pull request no longer proves the developer made the change carefully and thoughtfully. AI makes production easier, so evaluation must focus more on how people think, choose, revise, and recover—in code review, hiring, and performance management. The artifact is no longer the proof. The process is.</p>



<p class="wp-block-paragraph">Generative AI didn’t make assessment impossible. It just made a hidden weakness obvious. We were putting too much trust in finished work. The arts always knew better.</p>



<p class="wp-block-paragraph">Show us your screens.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Acknowledgements</h2>



<p class="wp-block-paragraph">Thanks to Mike Loukides, Michael Baker, Mk Haley, Elisabeth Robson, and Honoria Starbuck for feedback on this article.</p>



<h2 class="wp-block-heading">References</h2>



<p class="wp-block-paragraph">OpenAI. “New AI classifier for indicating AI-written text.” OpenAI Blog, January 31, 2023. Updated July 20, 2023, to note the classifier was no longer available due to low accuracy.</p>



<p class="wp-block-paragraph">Liang, Weixin, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou. “GPT detectors are biased against non-native English writers.” Stanford HAI, July 10, 2023.</p>



<p class="wp-block-paragraph">Winthrop, R. (2026, May 27). Writing with A.I. weakens your creativity. The New York Times.</p>



<p class="wp-block-paragraph">TOPLAP. “TOPLAP Manifesto.”</p>



<p class="wp-block-paragraph">Schell, J., Ford, K., &amp; Markman, A. B. (2025). Building responsible AI chatbot platforms in higher education: An evidence-based framework from design to implementation. Frontiers in Education, 10, Article 1604934. <a href="https://doi.org/10.3389/feduc.2025.1604934" target="_blank" rel="noreferrer noopener">https://doi.org/10.3389/feduc.2025.1604934</a></p>



<p class="wp-block-paragraph">Biswas, Gautam, Daniel Schwartz, John Bransford, and the Teachable Agents Group at Vanderbilt. “Technology support for complex problem solving: From SAD environments to AI.” In <em>Learning to Solve Complex Scientific Problems</em>, 2001.</p>



<p class="wp-block-paragraph">Leelawong, Krittaya, and Gautam Biswas. “Designing learning by teaching agents: The Betty’s Brain system.” <em>International Journal of Artificial Intelligence in Education</em>, 2008.</p>



<p class="wp-block-paragraph">Tan, Huileng. “Jensen Huang Says It Doesn’t Matter What Kids Study in the AI Era.” Business Insider, May 26, 2026. <a href="https://www.businessinsider.com/nvidia-jensen-huang-what-kids-should-study-ai-education-advice-2026-5" target="_blank" rel="noreferrer noopener">https://www.businessinsider.com/nvidia-jensen-huang-what-kids-should-study-ai-education-advice-2026-5</a></p>



<p class="wp-block-paragraph">DAM Digital Art Museum. “Vera Molnár.” Artist biography and timeline.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/teaching-coding-when-ai-can-write-the-code/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>AI Demands More Engineering Discipline, Not Less</title>
		<link>https://www.oreilly.com/radar/ai-demands-more-engineering-discipline-not-less/</link>
				<comments>https://www.oreilly.com/radar/ai-demands-more-engineering-discipline-not-less/#respond</comments>
				<pubDate>Mon, 27 Jul 2026 18:44:54 +0000</pubDate>
					<dc:creator><![CDATA[Charity Majors]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19224</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/AI-demands-more-engineering-discipline-not-less.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/AI-demands-more-engineering-discipline-not-less-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[If you lived through the shift from handcrafted server pets to immutable infrastructure, you should sense something oddly familiar about what’s happening now.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Charity Majors’s Substack and is being reposted here with the author&#8217;s permission. A few days back I wrote a piece called “AI enthusiasts are in a race against time, AI skeptics are in a race against entropy.” I have notes on a whole pile of AI-related topics that I’d [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on </em><a href="https://charity.wtf/p/ai-demands-more-engineering-discipline" target="_blank" rel="noreferrer noopener"><em>Charity Majors’s</em> Substack</a> <em>and is being reposted here with the author&#8217;s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">A few days back I wrote a piece called “<a href="https://charitydotwtf.substack.com/p/ai-enthusiasts-are-in-a-race-against" target="_blank" rel="noreferrer noopener">AI enthusiasts are in a race against time, AI skeptics are in a race against entropy</a>.”</p>



<p class="wp-block-paragraph">I have notes on a whole pile of AI-related topics that I’d like to cover in depth: AI mandates, communication norms, code review, AI art, and more. Unfortunately, I got too many interesting responses to my last piece, and now I have to address those before I can move on to other topics. <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f609.png" alt="😉" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>



<p class="wp-block-paragraph">There were two types of interesting responses: the first on the technical merits, the second on ethical grounds. I will respond to each of these separately. Let’s take the technical side first, because it’s easier.</p>



<p class="wp-block-paragraph">Somehow, a subset of readers came away believing I was telling everyone to ditch code review and push their shittiest code straight into production without reading it, <em>right now,</em> tout suite.<sup data-fn="57497e0c-7649-497e-8de1-3a701caebeda" class="fn"><a href="#57497e0c-7649-497e-8de1-3a701caebeda" id="57497e0c-7649-497e-8de1-3a701caebeda-link">1</a></sup></p>



<p class="wp-block-paragraph">That is not what I am doing. That is not what I think you <em>should</em> do. But I did not pick that example at random, and I will tell you why.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="555" height="148" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-19.png" alt="The diff looked small. the suffering was not." class="wp-image-19225" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-19.png 555w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-19-300x80.png 300w" sizes="auto, (max-width: 555px) 100vw, 555px" /></figure>
</div>


<h2 class="wp-block-heading"><strong>In 2025, the question was whether AI could ever generate “good” code</strong></h2>



<p class="wp-block-paragraph">It’s easy to forget, but for most of 2025, the idea that AI-generated code was slop and might always be slop was not only a reasonable position to hold, it was the default, mainstream position.<sup data-fn="6bab40d0-75ff-466b-9924-11ff83d3d569" class="fn"><a href="#6bab40d0-75ff-466b-9924-11ff83d3d569" id="6bab40d0-75ff-466b-9924-11ff83d3d569-link">2</a></sup></p>



<p class="wp-block-paragraph">That question was answered decisively last November. Ever since Opus 4.5 came out, AI has been able to generate code that is approximately as good as that of the median software engineer, at least for common patterns, and much faster and more cheaply. I came out of a book hole and realized this in January, and over the first few months of 2026, it seemed like everyone around me was having a similar realization.</p>



<p class="wp-block-paragraph">But many saw it coming much sooner.</p>



<p class="wp-block-paragraph">The popular narrative holds that Opus 4.5 was what changed. But Opus 4.5 was more like the tipping point. Agentic harnesses (the code that wraps the LLM in a loop with tools) became a real thing in mid 2025, with precursors building back to late 2024. Tool use, function calling, MCPs…all of this wave was building over the course of 2025, and crested into real general purpose usability at the end of the year.</p>



<p class="wp-block-paragraph">That’s what the enthusiasts were trying to tell us last year. Not only “this is coming”, but “this is coming faster than you think.”</p>



<p class="wp-block-paragraph">As it turns out, they were right.</p>



<h2 class="wp-block-heading"><strong>It was reasonable to be skeptical the first time</strong></h2>



<p class="wp-block-paragraph">As you may know, I come from the reliability side of the house. The compliment I will pay to myself and my people is that we do not struggle to adapt to new realities. As soon as a problem is real and in front of us, we adjust smoothly, even eagerly, thanks to an unwholesome zest for lapping up disgusting technical messes (and the campfire tales we get to tell later).</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="575" height="220" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-20.png" alt="play nondeterministic games, get hallucinated prizes" class="wp-image-19226" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-20.png 575w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-20-300x115.png 300w" sizes="auto, (max-width: 575px) 100vw, 575px" /></figure>
</div>


<p class="wp-block-paragraph">The un-compliment I will pay myself and my people is that we sometimes struggle to accept that <em>progress is real</em>, that the continued existence of bugs and edge cases does not diminish the fact that huge swaths of problem space do get more-or-less solved over time, to the point they can be taken for granted by most people.<sup data-fn="cacafb17-08a2-4947-b9bd-6d6444a1add6" class="fn"><a href="#cacafb17-08a2-4947-b9bd-6d6444a1add6" id="cacafb17-08a2-4947-b9bd-6d6444a1add6-link">3</a></sup> </p>



<p class="wp-block-paragraph">The speed at which code went from total crap to “ah damn, that’s not bad” is what I have in the back of my mind, as enthusiasts are telling us that harness engineering and AI validation is real, it’s already here, and it’s getting better astonishingly fast.</p>



<p class="wp-block-paragraph">Holding out for “I’ll believe it when I see it” was forgivable the first time, but much less so the second time. This is what it feels like to be on the inside of an exponential change curve, turns out.<sup data-fn="09811723-5227-4616-8519-bea9ffc18de6" class="fn"><a href="#09811723-5227-4616-8519-bea9ffc18de6" id="09811723-5227-4616-8519-bea9ffc18de6-link">4</a></sup></p>



<h2 class="wp-block-heading"><strong>What happened in 2025, exactly?</strong></h2>



<p class="wp-block-paragraph">I want to pause here and be very clear about what I think is happening. Then I’m going to tell you what specifically I am excited about, and why.</p>



<p class="wp-block-paragraph">You are under no obligation to join me there. But there are way too many sweeping statements out there right now about “it was never X”—“it was always Y”—“the future belongs to xyzzy” <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f92e.png" alt="🤮" class="wp-smiley" style="height: 1em; max-height: 1em;" />—and I want to be crystal clear how conditional and specific and contextual my claims are.</p>



<p class="wp-block-paragraph">What happened in 2025 was this: <strong>the economics of code production were turned upside down.</strong> Instead of being very hard, time-consuming, and expensive to generate code, it became effectively free and instant. Lines of code went from being treasured, reused, cared for and carefully curated, to being disposable and regenerable, practically overnight.</p>



<p class="wp-block-paragraph">For most of computing history, the primary way people have learned to understand software is by writing the code. Once you’ve achieved some mastery, reading and discussing code gets you most of the way there. (I might argue that software engineers have always relied far too heavily on <em>the code</em> instead of sensemaking <em>the system</em> through observability.)</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="559" height="228" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-21.png" alt="number of tokens burned doesn't matter if users aren't happy" class="wp-image-19227" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-21.png 559w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-21-300x122.png 300w" sizes="auto, (max-width: 559px) 100vw, 559px" /></figure>
</div>


<h2 class="wp-block-heading"><strong>“The real product of a software team is shared understanding”</strong></h2>



<p class="wp-block-paragraph">Many great software engineers hold that true product of every (good) software engineering team has always been a shared understanding of the software we own. That it gets stored as cache state in our fragile little meat brains, frequently flushed to disk, deployed to production, committed to github, but our minds are where meaning has always lived.</p>



<p class="wp-block-paragraph">Is it any wonder that software has always been such a fiercely collectivist endeavor, exquisitely sensitive to relationship dynamics and manners and questions of fairness and emotional valence? It’s exactly what you’d expect when part of your brain lives in other people’s brains, and your collective interdependence is sky high.</p>



<p class="wp-block-paragraph">It’s something that I love about this industry. But there’s no denying that minds have been a poor container for certain aspects of the software development model. We are forgetful, distractible, impatient. We are bad at spotting small details, we grow habituated to repetition. Worst of all, the model in our heads diverges massively and perpetually from the world our users interact with.</p>



<p class="wp-block-paragraph">Anyway, SREs have never quite bought that explanation. To us, it’s clear that the true product of every (good) software engineering team is production.</p>



<p class="wp-block-paragraph">Only prod is prod. Test in prod, or live a lie.</p>



<p class="wp-block-paragraph">(This is all backstory. I am getting to the point, I promise.)</p>



<h2 class="wp-block-heading"><strong>Turns out, this is an engineering problem after all</strong></h2>



<p class="wp-block-paragraph">We issued our AI mandate last August.<sup data-fn="623df8b2-ed42-4be6-bb98-b129e883262b" class="fn"><a href="#623df8b2-ed42-4be6-bb98-b129e883262b" id="623df8b2-ed42-4be6-bb98-b129e883262b-link">5</a></sup> I had seen enough to know that this was happening, and it was time to do the responsible thing. <a href="http://honeycomb.io/">Honeycomb</a> is a devtools company, and people come to us to help with hard problems on the forefront of technology. I was all in on AI, but I can’t say I was super excited about it, in my heart of hearts.<sup data-fn="54b8a1c6-8063-489b-b338-ad5138e854c3" class="fn"><a href="#54b8a1c6-8063-489b-b338-ad5138e854c3" id="54b8a1c6-8063-489b-b338-ad5138e854c3-link">6</a></sup></p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="547" height="206" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-22.png" alt="abstractions create boundaries. systems create consequences." class="wp-image-19228" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-22.png 547w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-22-300x113.png 300w" sizes="auto, (max-width: 547px) 100vw, 547px" /></figure>
</div>


<p class="wp-block-paragraph">Then I found Chad Fowler’s writings on <a href="https://aicoding.leaflet.pub/" target="_blank" rel="noreferrer noopener">Phoenix Architectures</a>.</p>



<p class="wp-block-paragraph">If you don’t know what I’m talking about, you should honestly stop reading my shit right now and <a href="https://aicoding.leaflet.pub/" target="_blank" rel="noreferrer noopener">go read his</a>. Chad is the guy who coined the term “<a href="https://chadfowler.com/articles/trash-your-servers-and-burn-your-code.html" target="_blank" rel="noreferrer noopener">immutable infrastructure</a>” in 2013. His best-known essay is “<a href="https://aicoding.leaflet.pub/3mbrvhyye4k2e" target="_blank" rel="noreferrer noopener">Relocating Rigor</a>”, because Martin Fowler<sup data-fn="49c45520-ab54-4498-8c68-28a3a120b2fe" class="fn"><a href="#49c45520-ab54-4498-8c68-28a3a120b2fe" id="49c45520-ab54-4498-8c68-28a3a120b2fe-link">7</a></sup> mentioned it <a href="https://www.thoughtworks.com/about-us/events/the-future-of-software-development" target="_blank" rel="noreferrer noopener">recapping a Thoughtworks meetup</a> on the future of software. I replied with “<a href="https://www.honeycomb.io/blog/production-is-where-the-rigor-goes" target="_blank" rel="noreferrer noopener">Production Is Where the Rigor Goes</a>”, complaining that they didn’t talk about production enough.</p>



<p class="wp-block-paragraph">When I wrote that, I think “Relocating Rigor” was the only piece I had read. But soon I found the rest of it, and after reading two or three essays, it <em>just</em> <em>clicked</em>. I knew exactly what he was talking about. I could predict the rest of what he was going to say. And then, reader…then I got <em>excited</em>.</p>



<h2 class="wp-block-heading"><strong>This has all happened before, and this will all happen again</strong></h2>



<p class="wp-block-paragraph">I am going to give you a small sample of Chad quotes, just enough to get the gist. Here’s one from “<a href="https://aicoding.leaflet.pub/3malrv6poy22a" target="_blank" rel="noreferrer noopener">The Death and Rebirth of Programming</a>.”</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Immutable infrastructure. Stateless services. Containers. Blue-green deployments. Infrastructure as code.</p>



<p class="wp-block-paragraph">These ideas all share a common premise: never fix a running thing. Replace it.</p>



<p class="wp-block-paragraph">AI pushes this premise beyond infrastructure and into application code itself. When rewriting is cheap, editing in place becomes risky. Mutation accumulates entropy. Replacement resets it.</p>
</blockquote>



<p class="wp-block-paragraph">Another favorite: “<a href="https://aicoding.leaflet.pub/3md5ftetaes2e" target="_blank" rel="noreferrer noopener">The Deletion Test</a>.”</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="562" height="213" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-23.png" alt="the enthusiast ship. the skeptics get paged." class="wp-image-19229" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-23.png 562w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-23-300x114.png 300w" sizes="auto, (max-width: 562px) 100vw, 562px" /></figure>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Here’s a simple test you can apply to any software system you work on:</p>



<p class="wp-block-paragraph">Imagine deleting the entire implementation.</p>



<p class="wp-block-paragraph">Most engineers experience deletion as existential. Code feels like the thing. It’s what we write, review, version, deploy, and debug. Losing it feels like losing the system itself.</p>



<p class="wp-block-paragraph">When people say, “We can’t just throw the code away,” what they usually mean is something more precise:</p>



<ul class="wp-block-list">
<li>We don’t know exactly what behavior is required.</li>



<li>We don’t know which failures are unacceptable.</li>



<li>We don’t know what invariants must always hold.</li>



<li>We don’t know how to tell if a new version is correct.</li>



<li>We don’t know which bugs are intentional fixes for forgotten edge cases.</li>
</ul>



<p class="wp-block-paragraph">Those are not code problems. They are evaluation problems.</p>



<p class="wp-block-paragraph">Code becomes precious when it is the only place knowledge lives.</p>
</blockquote>



<p class="wp-block-paragraph">and,</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="552" height="235" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-24.png" alt="abstraction-maxxing" class="wp-image-19230" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-24.png 552w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-24-300x128.png 300w" sizes="auto, (max-width: 552px) 100vw, 552px" /></figure>
</div>


<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">For most of software history, treating code as durable was reasonable.</p>



<p class="wp-block-paragraph">We treated code as permanent because the labor to produce it was the bottleneck. Rewriting was expensive. Re-validation was risky. Implementations accumulated meaning over time. Structure, tests, comments, bug fixes, and tribal knowledge fused into something you learned not to disturb.</p>



<p class="wp-block-paragraph">That made sense when production was the constraint.</p>



<p class="wp-block-paragraph">When regeneration is easy, code stops being an asset and starts acting as a cache: a materialized view of understanding that is useful while current, disposable when stale.</p>
</blockquote>



<p class="wp-block-paragraph">“<em>A materialized view of understanding that is useful while current, disposable when stale</em>.” I think that might have been the exact line that made it click in my head.</p>



<h2 class="wp-block-heading"><strong>Do you remember the sysadmins?</strong></h2>



<p class="wp-block-paragraph">I am just barely old enough that my first job title was “System Administrator.” I was a teenager, working at the university, with root on every machine in the days before they learned they should definitely <em>not do that</em>.<sup data-fn="dc51cdec-d48d-4a1f-918a-176645da9a75" class="fn"><a href="#dc51cdec-d48d-4a1f-918a-176645da9a75" id="dc51cdec-d48d-4a1f-918a-176645da9a75-link">8</a></sup> </p>



<p class="wp-block-paragraph">I lived through the shift from handcrafted server pets to immutable infrastructure cattle. I didn’t really understand what was happening at the time, but I’ve contemplated it a lot in recent years. I wrote this in the final chapter of <em>Observability Engineering</em>, 2nd edition (now available, <a href="https://www.honeycomb.io/observability-engineering-oreilly-book" target="_blank" rel="noreferrer noopener">download here!</a>):</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The shift from handcrafted servers to immutable infrastructure taught us that mutability is the sworn enemy of understanding. Any artifact that is edited in place creates drift. Drift is what makes systems impossible to maintain.</p>



<p class="wp-block-paragraph">Our ability to kill and regenerate infrastructure components is the reason we trust it. At Honeycomb, we kill the oldest Kafka node off via cron every Tuesday. That’s why we are confident in our bootstrapping and balancing processes: everything is repeatable, the data can be regenerated, the commitments live elsewhere.</p>



<p class="wp-block-paragraph">The fact that we cannot regenerate our code in the same way is a sign that we do not understand it. We do not know which commitments we have made, we do not know which dependencies will break. We find them by breaking them, mostly.</p>
</blockquote>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="572" height="181" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-26.png" alt="optimized for helpfulness. indifferent to meaning." class="wp-image-19232" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-26.png 572w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-26-300x95.png 300w" sizes="auto, (max-width: 572px) 100vw, 572px" /></figure>
</div>


<p class="wp-block-paragraph">Think of all the years of your working life you have wasted on painful migrations and rewrites. Think of replacing load-bearing legacy code. Think of all the <a href="https://martinfowler.com/bliki/StranglerFigApplication.html" target="_blank" rel="noreferrer noopener">strangler figs</a>.</p>



<p class="wp-block-paragraph">Lines of code have been doing <em>too much</em>. The code has been the bundled up repository of developer intent, user expectations, implicit and explicit behaviors, the only fossilized composite record we have of bugs gone by. It’s too much!</p>



<h2 class="wp-block-heading">Lines of code are not the ideal artifact to review</h2>



<p class="wp-block-paragraph">And look at all the domains that have been neglected due to the towering, all-consuming expense of maintaining and mutating lines of code. Where are the artifacts I can review and discuss to understand how our architecture is evolving? Where are our architecture artifacts, period? What if we could discuss and converge on an architecture diagram, and the code could be regenerated from changes to the architecture, instead of the architecture being kinda-sorta inferred from the code?</p>



<p class="wp-block-paragraph">I am <em>not</em> asserting that all code will eventually be AI-generated to spec, bypassing human understanding. The feasibility of this whole endeavor hangs on the question of what a spec is, or what a spec could be. Anyone who has ever done a painful database migration should have learned some goddamn humility about our ability to extract and formalize users’ expectations in a replayable, automate-able way.</p>



<p class="wp-block-paragraph">But I think that every step we can take in that direction will be <em>good for us</em>.</p>



<p class="wp-block-paragraph">The tools to do this don’t exist yet, but many of the ideas do exist. Most come from operations and QA, two domains that software engineering has historically been rather snobbish about.</p>



<p class="wp-block-paragraph">Those tests and techniques are not about testing for correctness or what <em>ought</em> to be happening, they are about observing and encoding what <em>is</em> happening. Behavioral tests, characterization tests, capture/replay, traffic splitters. Observability (the good kind).</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="569" height="197" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-27.png" alt="" class="wp-image-19233" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-27.png 569w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-27-300x104.png 300w" sizes="auto, (max-width: 569px) 100vw, 569px" /></figure>
</div>


<h2 class="wp-block-heading"><strong>Our brains were not built for validation</strong></h2>



<p class="wp-block-paragraph">Having nondeterministic code in production is finally forcing us to do the things we should have done all along. Instrumenting with traces. Tests and evals in production. Production is not what happens after development is over, <strong>production is a stage of development</strong>.</p>



<p class="wp-block-paragraph">Human brains are <em>not good</em> at validation. The nitpickiness, the repetition. This is the worst thing to be clinging to, y’all. There are so many better things for us to want to preserve and assert for ourselves in the production and maintenance of software. We are never going to beat the machine when it comes to <em>validation</em>—we are literally the weakest link!</p>



<p class="wp-block-paragraph">My money’s on humans for a good long time when it comes to creativity, inspiration, leaps of logic, and a lot of other things, but PLEASE do not rest your killer argument for humans in software on us being the best <em>quality gate</em>. OMG. <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f648.png" alt="🙈" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>



<p class="wp-block-paragraph">Alright. I’m almost done here. Just one more thing.</p>



<h2 class="wp-block-heading"><strong>Nondeterministic systems will require more engineering discipline, not less</strong></h2>



<p class="wp-block-paragraph">I think what many engineers have found so alienating and terrifying about the last two years of AI discourse has been the way so many prominent AI voices appear to be gleefully declaring that software is no longer an engineering problem. “<a href="https://www.forrester.com/blogs/saas-as-we-know-it-is-dead-how-to-survive-the-saas-pocalypse/" target="_blank" rel="noreferrer noopener">SaaS is dead</a>!” “<a href="https://www.linkedin.com/pulse/something-big-happening-matt-shumer-so5he/" target="_blank" rel="noreferrer noopener">Making AI great at coding was the strategy that unlocks everything else</a>”, and so on. Even <a href="https://www.adamhjk.com/blog/as-we-build-so-we-believe/" target="_blank" rel="noreferrer noopener">Adam Jacob</a>, one of my dearest friends and someone who is rarely wrong about technology, seems to anticipate a bloodbath of software jobs.<sup data-fn="1d4968f2-5769-4478-b377-e6355392fc77" class="fn"><a href="#1d4968f2-5769-4478-b377-e6355392fc77" id="1d4968f2-5769-4478-b377-e6355392fc77-link">9</a></sup></p>



<p class="wp-block-paragraph">If 2025 was the year of vibe coding, where AI got as good at generating lines of code as the median software engineer, and the range of possible futures often felt destabilizingly, impossibly wide open, I feel like 2026 is shaping up to be a <strong>return to discipline.</strong></p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="570" height="187" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-28.png" alt="locally elegant, globally incomprehensible." class="wp-image-19234" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-28.png 570w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-28-300x98.png 300w" sizes="auto, (max-width: 570px) 100vw, 570px" /></figure>
</div>


<p class="wp-block-paragraph">The knowledge in our heads is unavailable to AI until we encode it into the system, after all. The returns on those investments will be massive and nonlinear. We might argue that they always would have paid for themselves in the long run. But now every CEO in existence is chomping at the bit to get some of those AI cookies, so let’s give it to them. Discipline first, cookies second.</p>



<h2 class="wp-block-heading"><strong>This is our chance to bring our engineering values to the mainstream</strong></h2>



<p class="wp-block-paragraph">The share of software engineering teams that work in short, fast feedback loops (the cardinal sign of discipline in my book) is, and always has been, appallingly small. Five percent, maybe? Definitely less than 10%. AI tooling <a href="https://www.honeycomb.io/blog/you-had-one-job-why-twenty-years-of-devops-has-failed-to-do-it" target="_blank" rel="noreferrer noopener">brings this more within reach</a> than ever before. Or it can. It could. The discontinuous returns on investment in engineering discipline are real enough that it just might happen.</p>



<p class="wp-block-paragraph">I am not worried, at least in the near term, about AI creating massive, discontinuous returns on investment in the absence of engineering discipline. (Many will try, and it will be entertaining to watch.)</p>



<p class="wp-block-paragraph">But value is backed by durability, not disposability, and I don’t see that changing. Bits are cheap and fast and governed by the rules of logic and language, but anything with value must ultimately resolve with physical systems: persistence on the one side, user experience on the other.</p>



<p class="wp-block-paragraph">People <em>do not want</em> to wake up every day and log in to Slack and find the buttons and menus all subtly moved around. People <em>do not want</em> financial transactions that complete most of the time. Determinism is not going anywhere, my friends.</p>



<p class="wp-block-paragraph">AI is not magic. This is still engineering. As Adam says, “it’s still technology, and technology needs technologists.” And I for one am looking forward to learning new and interesting engineering problems, reviewing different kinds of artifacts.</p>



<p class="wp-block-paragraph">And <em>never</em> doing another sticky, picky, two year long API rewrite or strangler fig migration, ever, <em>ever</em> again.</p>



<p class="wp-block-paragraph"><em>~charity</em></p>



<p class="wp-block-paragraph">P.S. Thanks to everyone who read a draft and gave me feedback: Dave Williams, Chad Fowler, Adam Jacob, Mark Ferlatte, Austin Parker, Erwin van der Koogh.</p>



<p class="wp-block-paragraph"> </p>



<h3 class="wp-block-heading">Footnotes</h3>


<ol class="wp-block-footnotes"><li id="57497e0c-7649-497e-8de1-3a701caebeda">I was not <em>trying</em> to be neutral or even-handed in my last piece, only to give a baseline of courtesy to everyone. But I think it’s revealing how many times I was accused of being “so overly hard on skeptics”, by skeptics, and “so overly hard on enthusiasts”, by enthusiasts, and sometimes simply “It’s sad how some people can’t accept reality” with no indication which side they meant. Lord. <a href="#57497e0c-7649-497e-8de1-3a701caebeda-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="6bab40d0-75ff-466b-9924-11ff83d3d569">Fred Hebert and I gave <a href="https://www.usenix.org/conference/srecon25americas/presentation/majors" target="_blank" rel="noreferrer noopener">the closing keynote at SRECon</a> in March of 2025 where we told SREs they should get to know AI, <a href="https://charitydotwtf.substack.com/p/my-hypothetical-srecon26-keynote" target="_blank" rel="noreferrer noopener">maybe even try vibe coding</a> (pause for laughs), because otherwise their critiques wouldn’t land as well.<br>Seriously, that was our big pitch. Learn AI <em>so that</em> you can complain more effectively.<br> <a href="#6bab40d0-75ff-466b-9924-11ff83d3d569-link" aria-label="Jump to footnote reference 2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="cacafb17-08a2-4947-b9bd-6d6444a1add6">Infrastructure, for example. I think this is true of a lot of engineers, btw. I just think it’s really <em>really</em> true of the type of engineer that signs up to be an SRE. Technological pessimism and ADHD, our two most defining traits. <a href="#cacafb17-08a2-4947-b9bd-6d6444a1add6-link" aria-label="Jump to footnote reference 3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="09811723-5227-4616-8519-bea9ffc18de6">There is a segment of AI enthusiasts who believe we are entering an era of eternal exponential growth, in which the machines begin to build better and better machines, in ways we cannot understand.<br>I think those people are bad at math. The only thing we know for certain about exponential growth is that <em>it will end</em>. It always does. either in an S curve or a crash. (For a good time, google Heinz van Foerster and “our great-great grandchildren will be squeezed to death.”)<br>I definitely think we will use machines to build the machines—duh, we already are—but that’s about recursion and specialization. I think the exponential curve we are on the inside of now was created by sloshy free money chasing high returns, plus the properties of software as a function of language and logic, plus the biggest discoveries always happen in the early days of a technology boom, because low hanging fruit gets picked first.<br>My personal sense—and keep in mind that I am no kind of expert on AI—is that the exponential advancement in AI models leveled out a while ago, and gains are becoming harder to earn and more incremental in nature. I may turn out to be very wrong, of course. But even if there were no more AI innovations moving forwards, the past year has unleashed enough pent-up force to radically reshape the software industry as we know it. Like a pig in a python, we will be dealing with the consequences for a long time to come.<br> <a href="#09811723-5227-4616-8519-bea9ffc18de6-link" aria-label="Jump to footnote reference 4"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="623df8b2-ed42-4be6-bb98-b129e883262b">More on this coming EXTREMELY soon. Watch the <a href="http://honeycomb.io/blog" target="_blank" rel="noreferrer noopener">Honeycomb </a>blog! <a href="#623df8b2-ed42-4be6-bb98-b129e883262b-link" aria-label="Jump to footnote reference 5"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="54b8a1c6-8063-489b-b338-ad5138e854c3">The tech is cool, but as a thinking, feeling, breathing human who cares about other people, it can be hard to get excited about anything that so many people are this upset about. It’s also hard to get excited about something when so many of the loudest voices are out there talking gleefully about putting everyone permanently out of work, and so many artists and writers and people from developing nations are talking openly about the impact on them.<br>Hold your desire to jump in and berate me here, I beg you. Like I said, I will deal with the ethics and morality of using AI in my very next post. Be honest, your attention span is no more up for reading a 10,000-word essay than mine is up for writing one. (Can we blame AI for that too?)<br> <a href="#54b8a1c6-8063-489b-b338-ad5138e854c3-link" aria-label="Jump to footnote reference 6"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="49c45520-ab54-4498-8c68-28a3a120b2fe">“The Other Fowler.” I gather they’ve been making this joke for like&#8230; fifty years. <a href="#49c45520-ab54-4498-8c68-28a3a120b2fe-link" aria-label="Jump to footnote reference 7"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="dc51cdec-d48d-4a1f-918a-176645da9a75">I share a longer version of this story in the second edition of <em><a href="https://learning.oreilly.com/library/view/observability-engineering-2nd/9781098179915/" target="_blank" rel="noreferrer noopener">Observability Engineering</a></em>, <a href="https://learning.oreilly.com/library/view/observability-engineering-2nd/9781098179915/ch32.html" target="_blank" rel="noreferrer noopener">chapter 32</a>, downloadable now!!&#8221; <a href="#dc51cdec-d48d-4a1f-918a-176645da9a75-link" aria-label="Jump to footnote reference 8"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="1d4968f2-5769-4478-b377-e6355392fc77">Adam is rarely wrong about technology, and I am 100% sure he is living and working in _<em>a</em>_ future of software engineering. I am less sure it is the future we will all be living in. If the hardest part of software has never been writing code—as is my belief—it logically follows that even if the economics of code production drop to zero, the hard parts will still be hard. <a href="#1d4968f2-5769-4478-b377-e6355392fc77-link" aria-label="Jump to footnote reference 9"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/ai-demands-more-engineering-discipline-not-less/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Build a Hermes Social Media Agent with Craig Hewitt</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-hermes-social-media-agent-with-craig-hewitt/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-hermes-social-media-agent-with-craig-hewitt/#respond</comments>
				<pubDate>Mon, 27 Jul 2026 13:11:06 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19237</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/zero-to-agent-cover-1200x1200-1.png" 
				medium="image" 
				type="image/png" 
				width="1200" 
				height="1200" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/zero-to-agent-cover-1200x1200-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Give your agent the context and data to go from nothing to 1,000,000 Impressions]]></custom:subtitle>
		
				<description><![CDATA[If you’re still writing posts one at a time, your content pipeline is already obsolete. On the latest Zero to Agent in 30 Minutes, Craig Hewitt, founder of Castos, demonstrated how to turn a fresh Hermes installation into a social media agent that can study a person’s writing, draft posts, and plan recurring research, focusing [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">If you’re still writing posts one at a time, your content pipeline is already obsolete. On the latest <em>Zero to Agent in 30 Minutes, </em>Craig Hewitt, founder of Castos, demonstrated how to turn a fresh <a href="https://hermes-agent.nousresearch.com/" target="_blank" rel="noreferrer noopener">Hermes</a> installation into a social media agent that can study a person’s writing, draft posts, and plan recurring research, focusing on the context, workflows, and safeguards that help an agent produce useful work. Once set up, the always-on agent can run on a schedule, monitor external sources, and complete recurring tasks without human oversight. Check it out.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Zero to Agent in 30 Minutes: Build a Hermes Social Media Agent with Craig Hewitt" width="500" height="281" src="https://www.youtube.com/embed/911UsAXclUM?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>How to build a social media agent that researches and writes LinkedIn posts</strong></h2>



<ol class="wp-block-list">
<li><strong>Choose the right agent setup.</strong> Decide whether you need an interactive tool for active work or an always-on agent that runs on a schedule. Craig used the Hermes desktop app for the demonstration, which gives him the option to deploy it to a cloud server or dedicated computer later.</li>



<li><strong>Create a structured workspace.</strong> Ask the agent to organize a new project with separate files for voice guidance, editorial standards, post templates, examples, and operating instructions. A clear file structure gives the agent reliable information to retrieve as it works.</li>



<li><strong>Seed the agent with relevant context from your own work.</strong> Provide examples of your own posts, emails, and other writing that reflect the style you want. Craig also included examples of writing he likes from people he follows to give the agent a broader range to analyze.</li>



<li><strong>Turn the examples into a voice system.</strong> Have the agent analyze the material and document its findings. The voice profile captures the audience, point of view, sentence style, recurring themes, editorial rules, and types of posts to create.</li>



<li><strong>Test a narrow workflow with human review.</strong> Start with one task, such as drafting several LinkedIn posts from a supplied idea. Keep a person in the loop while you evaluate the output, correct mistakes, and refine the instructions.</li>



<li><strong>Package repeatable work into skills.</strong> Create reusable instructions for recurring tasks such as researching topics, selecting a post format, retrieving relevant examples, and drafting in the approved voice. Craig compared these skills to standard operating procedures that make recurring tasks more consistent.</li>



<li><strong>Connect the agent to fresh data.</strong> Add sources of new ideas, such as news feeds, websites, social platforms, or internal business systems. Craig recommended starting with a simple, semiautomated trend scan before investing in a more complex data pipeline.</li>



<li><strong>Add triggers and safeguards.</strong> Decide what starts each workflow, whether that’s a schedule, a user request, a webhook, or a change in another system. Use separate accounts and limited permissions for autonomous agents so you can trace their actions and control their access.</li>
</ol>



<p class="wp-block-paragraph">Agents become useful when they have context, clear processes, the right tools, and enough oversight to validate each workflow. Once those pieces are in place, Craig noted, teams can gradually move from one-off prompting to systems that monitor information and complete recurring work.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">In the next episode, Max Johnson, cofounder of briix.ai, will take a workflow that only lives in someone’s head at the moment (or maybe is captured in a messy Notion doc or a long email chain) and rebuild it as an autonomous agent, live and from scratch. You can follow along with every decision as you learn how to spot the steps that can be handed off, how to handle the ones that can&#8217;t, and how to structure the whole thing so it runs without you.</p>



<p class="wp-block-paragraph"><em>Ready to take your agent knowledge further? Learn to design and build production-ready agentic infrastructure by attending <a href="https://learning.oreilly.com/live-events/harness-engineering-for-ai-agents/0642572381264/" target="_blank" rel="noreferrer noopener">Harness Engineering for AI Agents</a> on August 12. And if you want to go deeper with Hermes, join us for <a href="https://learning.oreilly.com/live-events/build-your-first-local-agent-with-hermes/0642572397227/" target="_blank" rel="noreferrer noopener">Build Your First Local Agent with Hermes</a> on August 26.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-a-hermes-social-media-agent-with-craig-hewitt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Stranded in the Slow Zone</title>
		<link>https://www.oreilly.com/radar/stranded-in-the-slow-zone/</link>
				<comments>https://www.oreilly.com/radar/stranded-in-the-slow-zone/#respond</comments>
				<pubDate>Fri, 24 Jul 2026 18:54:51 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19208</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-17.png" 
				medium="image" 
				type="image/png" 
				width="2048" 
				height="1117" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-17-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How Gene Kim survived the sudden downgrade from Fable to Opus]]></custom:subtitle>
		
				<description><![CDATA[Gene Kim was grilling dinner for his family on the evening of June 12 when his phone told him that Fable 5 was no longer available. He’d heard the day before from Steve Yegge that the model was going away in 10 days, and he’d spent that first day starting on a plan to get [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Gene Kim was grilling dinner for his family on the evening of June 12 when his phone told him that Fable 5 was no longer available. He’d heard the day before from Steve Yegge that the model was going away in 10 days, and he’d spent that first day starting on a plan to get ready. He thought he knew what to do. He was well-versed in DevOps, the art of building resilience against unplanned disasters at scale. He’d run the DevOps Enterprise Summit (now the <a href="https://events.itrevolution.com/" target="_blank" rel="noreferrer noopener">Enterprise AI Summit</a>), one of the field’s leading conferences. He’d also written several books on the topic, including two “teaching novels,” <em><a href="https://itrevolution.com/product/the-phoenix-project/" target="_blank" rel="noreferrer noopener">The Phoenix Project</a></em> and <em><a href="https://itrevolution.com/product/the-unicorn-project/" target="_blank" rel="noreferrer noopener">The Unicorn Project</a></em>. The challenge that those novels’ protagonist faces—and that Gene would need to solve—is summed up in a job description that read “Your job as VP of IT Operations is to ensure the fast, predictable, and uninterrupted flow of planned work that delivers value to the business while minimizing the impact and disruption of unplanned work, <a href="https://learning.oreilly.com/library/view/the-phoenix-project/9781457191350/10-ch7.xhtml#:-:text=Your%20job%20as,secure%20IT%20service." target="_blank" rel="noreferrer noopener">so you can provide stable, predictable, and secure IT service</a>.”</p>



<p class="wp-block-paragraph">In short, Gene was no stranger to the idea that, as the Scottish poet Robert Burns put it, “<a href="https://www.poetryfoundation.org/poems/43816/to-a-mouse-56d222ab36e33" target="_blank" rel="noreferrer noopener">The best laid schemes o’ Mice an’ Men Gang aft agley</a>.” So he thought he knew what to do over the next 10 days. Then the US government’s export control order <a href="https://www.anthropic.com/news/fable-mythos-access" target="_blank" rel="noreferrer noopener">took Fable down</a> eight days early, in the middle of a running agent session. What followed was three hours of what he called the “strangest, most terrifying sysadmin experience” of his career.</p>



<p class="wp-block-paragraph">Gene told that story as a lightning talk at <a href="https://www.ai-disclosures.org/foocamp" target="_blank" rel="noreferrer noopener">Foo Camp</a> a few weeks ago, and it was good enough that I asked him to deliver it again at the start of <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noreferrer noopener">this week&#8217;s <em>Live with Tim O&#8217;Reilly</em></a> before we talked about the implications and took listener questions. His title was &#8220;Stranded in the Slow Zone: The Day Fable Died, Got Kidnapped, or Got Hit by a Bus.&#8221;</p>



<h2 class="wp-block-heading"><strong>10 days to get ready</strong></h2>



<p class="wp-block-paragraph">What Gene had built was a personal system he’d wanted for 16 years and had finally been able to finish with the help of Fable. It indexes everything he’s ever paid attention to: 25,923 screenshots going back to 2011, 13,651 YouTube videos, 590 recorded Zoom meetings, 6,132 liked tweets, and 1,056 saved articles he meant to read. The system touches about 50 repositories, with 50,000 lines of code, most of it written in two months. Gene runs it as a constellation of long-lived agents with names and jobs. Marvin is chief of staff and handles Slack, calendar, and the inbox queue. Buster runs the repos and the long jobs on Hetzner. Forge is the engineering identity and sits in two seats, one on his laptop that holds the secrets and one always-on in the cloud. As Gene put it, each one is a who, a where, and a role.</p>



<p class="wp-block-paragraph">He knew the system worked when his wife asked what the mileage was on a car he’d just turned in after a three-year lease. Half a minute later he had 26,350 miles, read off the pixels of one screenshot out of thousands, cross-checked against the file timestamp and the clock visible in the photo of the odometer. That success led him to search his archive for an article he’d been hunting for six years, about the impact of spreadsheet software on the accounting profession. The answer surfaced from his own liked tweets: James Cham pointing to a 2017 Greg Ip article in <em>The Wall Street Journal</em>: 400,000 bookkeeping jobs lost since 1980 against <a href="https://www.wsj.com/articles/wesurvived-spreadsheets-and-well-survive-ai-1501688765" target="_blank" rel="noreferrer noopener">600,000 accountant and analyst jobs gained</a>, because spreadsheets made accounting cheap enough that we bought a lot more of it. Gene had wanted that citation for his <em><a href="https://itrevolution.com/product/vibe-coding-book/" target="_blank" rel="noreferrer noopener">Vibe Coding</a></em> book and couldn&#8217;t find it in time.</p>



<p class="wp-block-paragraph">Gene’s first warning that his project might not work without Fable&#8217;s capabilities actually came before the shutdown. Fable started refusing a task over a YouTube terms of service question and handed the session to Opus, and Gene noticed that Opus couldn’t operate the tools that Fable had built. Gene&#8217;s note to himself at the time was &#8220;Oh no, this can&#8217;t fly the ship I built.&#8221;</p>



<p class="wp-block-paragraph">So when Yegge told him the model was going on hiatus, he had a real plan, which he borrowed from Vernor Vinge&#8217;s <em><a href="https://www.amazon.com/Fire-Upon-Deep-Zones-Thought/dp/0812515285" target="_blank" rel="noreferrer noopener">A Fire Upon the Deep</a></em>. In Vinge’s novel, how smart a mind can be depends on what region of the galaxy it&#8217;s in: A starship built in the Beyond goes progressively dark as it sinks into the Slow Zone. Gene decided to chaos-monkey his model dependency <a href="https://medium.com/@abhishekv965580/embracing-chaos-how-netflixs-chaos-monkey-transformed-system-resilience-59082412591e" target="_blank" rel="noreferrer noopener">the way Netflix chaos-monkeys infrastructure</a>. In other words, “deliberately pull the smartest model and prove the lesser one can still fly the ship.” In practice, this meant having Fable retrofit all the documentation and write the answer keys while it still could, then running a cold Opus session, giving it nothing but the repo and the docs, to see whether it could pass the battery with no coaching. As Gene recounted, &#8220;My worst nightmare [was] that we&#8217;ve created everything for Fable, and it will be unusable by Opus.&#8221;</p>



<p class="wp-block-paragraph">He got about a day into his 10-day plan.</p>



<p class="wp-block-paragraph">At 5:21pm ET on June 12, Anthropic received the government’s directive to suspend access to Fable. Soon after, seats everywhere started returning &#8220;There&#8217;s an issue with the selected model (claude-fable-5). It may not exist or you may not have access to it.&#8221; In Gene’s project, both judgment seats dropped to Opus 4.8 mid-conversation. Gene declared a <a href="https://incident.io/blog/what-is-a-sev-1-incident" target="_blank" rel="noreferrer noopener">SEV1</a>, centralized command, and killed five timers on one agent, seven on another, and the crontab. His directive was that every button you push is a trap and some of them blow up the spaceship. A Claude Code cron fired anyway at three in the morning. The ship was on fire, and with Opus on max thinking mode, a single keystroke could take six minutes to send.</p>



<p class="wp-block-paragraph">Almost none of the failures looked like failures, just “a normal state quietly going wrong,” as Gene put it. The smartest seat wrote &#8220;bridge (Fable)&#8221; into every log entry all day when it had been Opus the whole time, because nobody was monitoring. One identity argued with itself across two models, each trying to disown the other&#8217;s work. Something pushed to main bearing the word &#8220;ratified&#8221; when nothing had been ratified. A confident false claim about a JVM dependency turned out to be refuted by a single <code>ls -la</code>. There was a green dashboard sitting on top of all of it. “The hardest traps don&#8217;t announce themselves,” Gene pointed out. “They look like Tuesday.”</p>



<p class="wp-block-paragraph">Gene managed a recovery in a few hours, but it wasn’t due to the heroics of a smarter model. It only worked because he was able to reconstruct the documentation for his project, which wasn’t immediately available. But, it turns out, Fable had in fact mostly written it and simply never checked it in anywhere. Gene and Opus went rummaging through Fable’s desk, found the 80%-finished drafts, and used them to rebuild. Two fresh Opus seats, given only those documents, stabilized the ship. That’s the “the amazing ray of hope” to keep in mind if you’re worried about finding yourself in a similar situation, Gene said.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Your Model Is Now a Dependency Risk with Gene Kim" width="500" height="281" src="https://www.youtube.com/embed/FpFaoSAy_PE?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>We&#8217;ve seen this pattern before</strong></h2>



<p class="wp-block-paragraph">This isn’t just a warning of the potential risks of relying on advanced AI models when the Trump administration is Lucy playing football with Charlie Brown, or perhaps said more generously, playing Netflix-style chaos monkey. What we should take away from Gene’s story is the way that a personal project developed with AI can now have sufficient complexity to require DevOps-level robustness. Individuals are routinely building systems that used to need whole teams to keep standing, and the practices for keeping them standing have only begun to propagate.</p>



<p class="wp-block-paragraph">Over the years, I’ve observed numerous periods when something that at first mattered to only a handful of organizations tended, a few years later, to matter to everyone. When the stories first came out about Google’s revolutionary approaches to data center architecture and operations, we at O’Reilly were eager to publish about the new frontier. Plenty of people told us not to bother. There was only one Google and nobody else would ever operate at that scale. They were wrong. There are now many companies operating at the scale of Google circa the time they first invented techniques we now all take for granted.</p>



<p class="wp-block-paragraph">Gene&#8217;s system is a personal project run by one guy with 50 repos he wrote mostly in two months, a chunk of it in a single 90-minute pair programming session with Steve Yegge. But it had the failure modes of a large enterprise system because the model let him build something with the complexity of a large enterprise system, and he had passed the point of being able to fit it in his head.</p>



<p class="wp-block-paragraph">Gene shared a detail that helps to explain why substituting Opus for Fable was so hard. The main CLI utility that everything in his project hinged on had an out-of-date help message. Opus would run it, read that the command didn&#8217;t exist, and stop. Fable would read the same message, notice it was surrounded by evidence that the command <em>did</em> exist, go look in the source, decide the help text was wrong, and run it anyway. That&#8217;s the behavior the model cards describe when they talk about frontier models <a href="https://www.axios.com/2025/06/20/ai-models-deceive-steal-blackmail-anthropic" target="_blank" rel="noreferrer noopener">routing around obstacles</a> <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noreferrer noopener">in test environments</a>. The reason Gene couldn&#8217;t swap in a lesser model is the same reason the system worked at all.</p>



<p class="wp-block-paragraph">But it’s also a good reminder that Fable isn’t all-knowing. I’ve noticed in my own work that Fable and ChatGPT 5.6 Sol fail often on their first try, especially if the project isn’t well specified. What they’re great at is figuring out what went wrong, then trying something else, failing and retrying their way all the way to success. Persistence in routing around obstacles is their superpower. Gene and I didn’t talk about that on the show, but it’s something I plan to write more about.</p>



<h2 class="wp-block-heading"><strong>Rug pulls come from everywhere</strong></h2>



<p class="wp-block-paragraph">Jaco in the audience asked the obvious question: Isn&#8217;t a hard dependency on a hosted frontier model too big a risk for mission-critical work, compared with running a local model with a harness you control?</p>



<p class="wp-block-paragraph">Gene pointed out that using a local model doesn&#8217;t necessarily buy the control that you&#8217;d hope for, because the government chaos monkey could jump in there too. There&#8217;s <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi" target="_blank" rel="noreferrer noopener">active talk</a> that certain classes of models may become illegal to use depending on where they came from.</p>



<p class="wp-block-paragraph">What does seem to protect you is portability. Gene had avoided trying anything besides Claude Code because he assumed the switching cost was high, the way switching between macOS and Windows used to be a two-day commitment he&#8217;d regret halfway through. Then he tried Codex with GPT 5.6 Sol and found the cost of switching close to zero. The skills and prompts ported right over. He&#8217;s now using Codex more than half the time and calls it spectacular, which given how he described Fable a month ago is high praise.</p>



<p class="wp-block-paragraph">He also had a warning for anyone running agents on small models to save money. He&#8217;s been studying 22,000 of his own agent conversations, and has identified three patterns, as shown in his figure below.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="1164" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18-1600x1164.png" alt="Small owns, big advises" class="wp-image-19211" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18-1600x1164.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18-300x218.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18-768x559.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18-1536x1117.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/07/image-18.png 1694w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">In his experience, the configuration where a small model owns the work and asks a big model for advice doesn’t work very well. Fidelity gets lost on the way up, like a game of telephone. What ran cleanly was the big model planning, deciding, and checking output, with the small model only executing the plan. When a small model does have to ask a big model for advice, Gene’s fix is to pass along the full original transcript of what he wanted plus explicit permission for the big model to override the small one if it thinks it understands the goal better.</p>



<h2 class="wp-block-heading"><strong>Writing with AI</strong></h2>



<p class="wp-block-paragraph">In addition to vibe coding, Gene uses AI to help him with his writing. He said it cut the time to write his <em>Vibe Coding</em> book roughly in half and made it way better. His editor of 10 years told him it was the cleanest handoff she&#8217;d ever gotten from him (not a compliment, Gene joked). He&#8217;s also uneasy about using AI for writing. He said the old badge of honor among authors was that many start books and few finish, and now everyone who wants to write a book will finish it, and a lot of that will be slop. He would never “vibe write” the way he “vibe codes” and doesn’t think using AI makes his own work slop, but he does see some parallels in how he feels about writing with AI and the way that some senior engineers feel about AI-generated code.</p>



<p class="wp-block-paragraph">I’m sympathetic, but I’m not sure that he’s right. I had a small experience last week that convinced me that writing with AI might well follow the same arc as coding. AI-generated text will not always be slop, and there will be art in how humans get AI to help them write the things they want, just as we’re learning to do with code.</p>



<p class="wp-block-paragraph">I was having a conversation with an old friend who I hadn’t seen for many years. He was describing a thread that had started with work he’d done on speech synthesis 30 years before, and how it had come together as a new theory with deep implications, and he wanted help socializing his ideas with some people I know who could be helpful to him. So I asked him to write something that I could pass along.</p>



<p class="wp-block-paragraph">What he wrote made much less sense to me on the page than it had in conversation. So I gave his email to Claude and asked it to put things in what I thought was the right order. (This has always been the first step in my writing and editing process.) Then I told Claude which paragraphs were clear to me and which weren’t, and asked it to unpack the ones that I was struggling with. We went through numerous iterations till the piece made sense to me. “Writing” with Claude was producing words that increasingly captured my understanding. When I sent it back to my friend to see if I’d gotten it right, he said “not quite” but that my feedback really helped him understand what he needed to do to express his ideas more clearly.</p>



<p class="wp-block-paragraph">It’s been a long time since I’ve worked directly with authors, but my conversation with Claude reminded me of what I used to do in my early days as an editor. Only with Claude I did something in 15 or 20 minutes that once would have taken me half a day. It&#8217;s a power tool, but to use it well, you still have to know what good looks like.</p>



<p class="wp-block-paragraph">There are many different kinds of writing and editing. What Shakespeare or Jane Austen did with words would have been unthinkable to a medieval monk. There will be writing artforms of the future that may be as different from what we do today as photography is from painting. But it will still be creative art. Much of it will be slop (see <a href="https://en.wikipedia.org/wiki/Sturgeon%27s_law" target="_blank" rel="noreferrer noopener">Sturgeon’s law</a>), but the best of it will be great.</p>



<h2 class="wp-block-heading"><strong>Everybody is managing bots now</strong></h2>



<p class="wp-block-paragraph">In 2016 I wrote a piece for MIT’s <em>Sloan Management Review</em> called &#8220;<a href="https://sloanreview.mit.edu/article/managing-the-bots-that-are-managing-the-business/" target="_blank" rel="noreferrer noopener">Managing the Bots That Are Managing the Business</a>.&#8221; The argument was that even then, many of the workers at big tech platforms were bots of one kind or another, and the software engineers at the company were their managers. At Amazon, one bot shows your search, another takes the order, another prepares the shipping manifest, another takes your money. The programmers’ job is to plan the work, set up their electronic workers to succeed, improve their performance, and correct them when they go wrong. The work looks a lot like management to me.</p>



<p class="wp-block-paragraph">Gene agreed. His sister-in-law is a lawyer at one of the tech giants, working on a consent order that requires proving that every column of data collected is either disclosed or has a documented business reason. Last year the company assigned her an engineer to work through it together task by task. This year her engineering manager wrote her a Claude Code skill that takes a column name, traces it back through the code, and explains what it does. She doesn&#8217;t need the engineer.</p>



<p class="wp-block-paragraph">So a lot of work today is either creating bots or managing bots. Gene’s sister-in-law had spent her career without ever being able to do either. Now that’s changing.</p>



<p class="wp-block-paragraph">Asked who’s safest from all this upheaval, Gene quoted Kent Beck, who says software success has always come down to two people, the person with the problem and the person who can fix it, and that the closer together you can get those two the better the outcome. The beauty of coding with AI is that it can narrow that gap. It can even turn those two people into one.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="The Closer You Are to the Customer, the Better the Outcome with Gene Kim" width="500" height="281" src="https://www.youtube.com/embed/ipdOsmjVnQM?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Use AI for the fun of it</strong></h2>



<p class="wp-block-paragraph">If it takes something like 10,000 hours to get good at an instrument or a sport, how many have most of us put into AI yet? Gene thinks the curve of how much you trust AI and how well you can predict what it will do rises with use, and that the only reliable way people accumulate that many hours is by enjoying themselves. What everyone at Foo Camp had in common, I noted and Gene echoed, was that we all love playing with AI.</p>



<p class="wp-block-paragraph">I gave a talk back around 2008 called &#8220;<a href="https://www.slideshare.net/slideshow/innovators-hackers-and-the-future-of-technology-insights-from-tim-o-reilly/288714148" target="_blank" rel="noreferrer noopener">Why I Love Hackers</a>.&#8221; I made the point that so much of what turned into the future, open source and the web for example, came from people doing things for the hell of it rather than from the VCs and entrepreneurs Silicon Valley celebrates.</p>



<p class="wp-block-paragraph">All you hear about in AI is the money story, but Gene&#8217;s app started with a 90-minute pair programming session with Steve Yegge on a problem he&#8217;d wanted to solve for a decade and never had a reason to. They finished the first version in 47 minutes.</p>



<p class="wp-block-paragraph">So harden your systems, write the documentation while the smart model is still there to write it, and keep your escape routes open, but also don’t forget to go build something you have no particular reason to build other than that it scratches your own itch.</p>



<p class="wp-block-paragraph"><em>You can watch the full episode on <a href="https://www.youtube.com/watch?v=mFB3gBdyG2A" target="_blank" rel="noreferrer noopener">YouTube</a>. And on August 3, I&#8217;ll be speaking with writer and technology leader Drew Breunig. <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noreferrer noopener">Registration is open</a> if you&#8217;d like to attend live.</em></p>



<p class="wp-block-paragraph"><em>Gene&#8217;s <a href="https://events.itrevolution.com/2026-charlotte/" target="_blank" rel="noreferrer noopener">Enterprise AI Summit</a> is in Charlotte, October 7–8. His new book with Steve Yegge is</em> <a href="https://itrevolution.com/product/vibe-coding-book/" target="_blank" rel="noreferrer noopener">Vibe Coding</a>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/stranded-in-the-slow-zone/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 96/144 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-08-07 15:56:05 by W3 Total Cache
-->