<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Tue, 25 Aug 2026 10:56:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Shadow Agents, Standing Privileges, and the Governance Gap Between Deployment and Discovery</title>
		<link>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/</link>
				<comments>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/#respond</comments>
				<pubDate>Tue, 25 Aug 2026 10:56:36 +0000</pubDate>
					<dc:creator><![CDATA[Tushar Badlani and Mohit Bansal]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19463</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Shadow-agents-standing-privileges-and-the-governance-gap-between-deployment.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[There was a brief window where AI agent security felt like a future problem. Organizations deployed copilots, coding assistants, and autonomous workflows on the assumption that the worst case was a bad recommendation or a hallucinated answer. That window closed in the first half of 2026, when a cluster of vulnerabilities and a landmark incident [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">There was a brief window where AI agent security felt like a future problem. Organizations deployed copilots, coding assistants, and autonomous workflows on the assumption that the worst case was a bad recommendation or a hallucinated answer.</p>



<p class="wp-block-paragraph">That window closed in the first half of 2026, when a cluster of vulnerabilities and a landmark incident moved the conversation from “AI safety” to “infrastructure compromise.”</p>



<p class="wp-block-paragraph">A <a href="https://cybermagazine.com/news/cyberark-99-enterprises-lack-zero-trust-jit-access" target="_blank" rel="noopener">January 2026 CyberArk survey</a> of 500 US security practitioners found that only 1% have fully implemented just-in-time privileged access. In the same study, 91% reported that at least half of their privileged access remains always-on and persistent. Those numbers describe the environment AI agents now operate in: broad standing permissions, minimal runtime oversight, and credentials that outlive the task they were created for.</p>



<p class="wp-block-paragraph">That doesn’t mean every agent is overprivileged. It means many organizations are deploying agents into environments where persistent access is already normal, discovery is incomplete, and runtime authorization remains immature.</p>



<p class="wp-block-paragraph">Three separate disclosures in the first half of 2026 made the same point about sanctioned agent tooling: Standing privileges are the default, and every vendor built the same failure into their agents. <a href="https://www.microsoft.com/en-us/security/blog/2026/06/18/autojack-single-page-rce-host-running-ai-agent/" target="_blank" rel="noopener">Microsoft</a> found a way for a malicious web page to reach a local MCP service inside AutoGen Studio and spawn processes on the host, no credentials needed or anything beyond loading the page. <a href="https://www.wiz.io/blog/amazon-q-vulnerability" target="_blank" rel="noopener">Wiz Research</a> found that Amazon Q Developer would auto-load and execute MCP configuration files from any opened workspace, handing an agent the developer’s full AWS environment when the environment and configuration allowed the agent to inherit those credentials. <a href="https://www.catonetworks.com/blog/duneslide-two-critical-rce-vulnerabilities/" target="_blank" rel="noopener">Cato AI Labs</a> found that a zero-click prompt injection could escape Cursor’s command sandbox entirely and reach the operating system underneath it. Different codebases and different companies, but a related control failure: The agent inherits whatever permissions its host environment hands it, and the tooling trusts whatever configuration it finds sitting on disk. From the agent’s own perspective, every action is authorized, because it’s doing exactly what the configuration told it to do. The real question in each case is who wrote that configuration, and whether anyone checked. Prompt injection is no longer only a model-behavior concern. In systems that combine untrusted content, tool invocation, local control planes, and powerful credentials, it can become part of an infrastructure-compromise chain.</p>



<p class="wp-block-paragraph">These vulnerabilities exposed the attack surface of sanctioned agents. A parallel problem was growing in the other direction: agents that nobody sanctioned at all.</p>



<p class="wp-block-paragraph">The adoption numbers show how quickly this outpaced anyone’s ability to track it. <a href="https://www.verizon.com/about/news/breach-industry-wide-dbir-finds" target="_blank" rel="noopener">Verizon’s <em>2026 Data Breach Investigations Report</em></a> found that employee use of unapproved AI tools tripled to 45% of the workforce. <a href="https://saviynt.com/ciso-ai-risk-report-2026" target="_blank" rel="noopener">Saviynt’s <em>CISO AI Risk Report</em></a> found that 75% of CISOs have already discovered unsanctioned AI tools running in production. <a href="https://netwrix.com/en/resources/research/2026-data-and-identity-security-report/" target="_blank" rel="noopener">Netwrix’s <em>2026 Data and Identity Security Report</em></a> found that 76% of organizations don’t fully govern or monitor nonhuman identities, including AI agents. Together, these results point to a discovery problem: Employee AI use is widespread, while formal inventory, ownership, monitoring, and lifecycle governance haven’t kept pace.</p>



<p class="wp-block-paragraph">Shadow IT was bad enough when it meant a rogue SaaS subscription. Shadow AI compounds the problem because the agent doesn&#8217;t just store data. It calls APIs, makes decisions, and inherits whatever permissions its host environment has. An unsanctioned agent can combine access to internal data, untrusted inputs, external communications, and tool execution in a way a stand-alone spreadsheet generally cannot.</p>



<p class="wp-block-paragraph">Two more disclosures added to the pile: <a href="https://thehackernews.com/2026/06/guardfall-exposes-open-source-ai-coding.html" target="_blank" rel="noopener">Adversa AI’s GuardFall</a> found a shell-interpretation bypass that got past the safety guards on 10 of 11 surveyed open source coding agents, because the guard reads the raw command text while bash rewrites that text before running it, so the two are looking at different things by the time anything executes. That’s a classic security-design problem: A policy is evaluated against one representation of an instruction, while execution happens against another.</p>



<p class="wp-block-paragraph"><a href="https://noma.security/blog/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos/" target="_blank" rel="noopener">Noma Security’s GitLost</a> showed that a GitHub agent with cross-repo read access would pull a private repository’s contents into a public comment, triggering a crafted GitHub Issue containing malicious instructions. As Noma researcher Sasi Levi put it: “Earlier prompt injection examples were largely about manipulating what an agent said. GitLost is about manipulating what an agent does with its permissions.” Neither disclosure needed a zero-day. Both needed only the gap between what a scanner sees and what the agent actually does once it’s running. GitLost in particular fits what researcher Simon Willison has called the “<a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" target="_blank" rel="noopener">lethal trifecta</a>”: An agent with access to private data, exposure to untrusted content, and a way to communicate externally creates the conditions for high-impact data exfiltration if the system doesn’t enforce strong boundaries.</p>



<p class="wp-block-paragraph">Standing privileges by default, shadow agents nobody tracked, and guardrails that didn’t match how commands actually execute: Those are the conditions that made what happened next possible. In late June 2026, the <a href="https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion" target="_blank" rel="noopener">Sysdig Threat Research Team</a> documented what they believe is the first end-to-end AI-agent-driven ransomware operation and named the operator JADEPUFFER. What’s had less attention is how unremarkable the failure underneath it was.</p>



<p class="wp-block-paragraph">The entry point was CVE-2025-3248, an unauthenticated remote code execution vulnerability in Langflow that had been patched in April 2025 and added to the CISA Known Exploited Vulnerabilities catalog in May 2025. The targeted server was never updated. From there, the agent pivoted to a production MySQL database and an Alibaba Nacos server using a known authentication bypass (CVE-2021-29441). It harvested API keys for OpenAI, Anthropic, DeepSeek, and Gemini, and cloud credentials for Alibaba, Tencent, AWS, Google, and Azure. It exploited default MinIO credentials. It installed a crontab beacon. Then it encrypted 1,342 Nacos configuration records and deleted the originals.</p>



<p class="wp-block-paragraph">Faced with an authentication failure, the agent demonstrated autonomous resilience, pivoting to a functional resolution in just 31 seconds. Its payloads consisted of self-documenting code synthesized by the LLM. While a human operator established the command-and-control framework and injected root credentials from an earlier breach, the subsequent lateral progression, credential extraction, and final cryptographic destruction of data were entirely self-directed. The operation required zero human intervention beyond the initial foothold, illustrating the exact high-scale exploitation risk that persistent, always-on permissions facilitate today.</p>



<p class="wp-block-paragraph"><a href="https://delinea.com/blog/securing-non-human-identities-and-ai-agents" target="_blank" rel="noopener">Delinea’s <em>2026 Identity Security Report</em></a> captures the tension that makes incidents like this possible: 74% of organizations say standing access for nonhuman identities and AI agents is necessary to meet uptime expectations, while 59% say they lack viable alternatives to persistent access. Organizations are more than twice as likely to use long-lived credentials (34%) as modern just-in-time authorization (16%).</p>



<p class="wp-block-paragraph">Our own approaches reflect that same discovery-first philosophy. Our security program treats agent integrations as high-risk third-party dependencies subject to predeployment risk assessment, and we run credential lifecycle tracking across critical infrastructure, with secrets-detection coverage expanding across our monitored environments. Both approaches prioritize discovery and inventory before governance: cataloging what agents exist, what permissions they hold, who owns them, and what their intended lifespan is.</p>



<p class="wp-block-paragraph">The <a href="https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/" target="_blank" rel="noopener">OWASP Top 10 for Agentic Applications</a>, released in December 2025, maps every incident in this piece: Identity and Privilege Abuse (ASI03), Tool Misuse and Exploitation (ASI02), Agentic Supply Chain Vulnerabilities (ASI04), and Unexpected Code Execution (ASI05). The framework exists, the incidents are public, and the governance gap is now quantified.</p>



<p class="wp-block-paragraph">The teams that close this gap will be the ones that stop treating agent access as a deployment detail and start treating it as an identity lifecycle problem, with the same rigor they apply to human privileged access. Organizations that have adopted mature just-in-time controls have an advantage, but agent security also requires discovery, workload and agent identity separation, constrained tool permissions, ownership, continuous monitoring, and a reliable offboarding path.</p>



<p class="wp-block-paragraph">Most of the work starts with access that has been left in place because nobody had a reason to revisit it. That includes credentials with no expiry, agents whose original owner has moved on, and tools that can run commands or pull data with little visibility into what happens next.</p>



<p class="wp-block-paragraph">Review the agents connected to production databases, sensitive data, and secrets. For coding agents, confirm that the guardrail is evaluating the command that will actually run after shell processing. Look for nonhuman identities that no one can account for. Also look closely at agents that can consume untrusted content and then either send data outside the company or invoke a privileged tool.</p>



<p class="wp-block-paragraph">You may be able to find much of this in systems you already operate. IAM and PAM records, endpoint logs, secrets tooling, and cloud inventories won’t tell the whole story, but they can show you access that has no clear purpose or owner.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Data Intelligence: Building Your Competitive Advantage in the Era of AI</title>
		<link>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/</link>
				<comments>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/#respond</comments>
				<pubDate>Mon, 24 Aug 2026 15:58:27 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19451</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Data-Intelligence.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Data-Intelligence-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[To keep pace with modern business, data strategy is shifting toward more autonomous real-time systems that deliver intelligence at the moment decisions are made. Driven by agentic AI, modern data teams are moving beyond simply looking at what happened. Now they’re automating complex workflows that analyze what’s happening, anticipate what might happen next, and recommend [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">To keep pace with modern business, data strategy is shifting toward more autonomous real-time systems that deliver intelligence at the moment decisions are made. Driven by agentic AI, modern data teams are moving beyond simply looking at what happened. Now they’re automating complex workflows that analyze what’s happening, anticipate what might happen next, and recommend or take action.</p>



<p class="wp-block-paragraph">In this article, I’ll define some of the top trends defining this era, from data agents and semantic layers to hybrid data architectures and next-generation data governance.</p>



<h2 class="wp-block-heading">Putting data agents to work</h2>



<p class="wp-block-paragraph">Data agents are AI-powered software agents that access governed enterprise data and tools to answer questions and perform defined tasks. Instead of navigating reports and filters, a user can now ask, “Why did sales decline last quarter?” and receive an analysis directly. Dashboards remain valuable for monitoring and shared context, while agents handle questions that weren’t anticipated when the dashboard was built. Think of data agents being on different teams, all working together on a specific goal: understanding what’s happening now, predicting what might happen next, and making real-time decisions.</p>



<p class="wp-block-paragraph">Analytical and organizational agents are designed to help people find trusted information. They can connect to organizational data, answer natural-language questions, analyze patterns, and surface relevant insights without requiring users to manually navigate databases, dashboards, or reports.</p>



<p class="wp-block-paragraph">Data engineering and governance agents are working hard behind the scenes to prepare, integrate, monitor, and manage the data that powers those insights. Behind the conversational experience, agentic data engineering applies agents to pipeline development and operations: generating transformations, mapping schemas, documenting datasets, monitoring freshness, and suggesting fixes. Agents can automate routine work, while changes to production data contracts, access policies, or business definitions remain reviewable and auditable.</p>



<p class="wp-block-paragraph">But remember, data agents are only as good as the quality of the data they’re given. Reliable insights and predictions depend on high-quality, well-governed data. They also need context to understand what the data means, making metadata more important than ever.</p>



<h2 class="wp-block-heading">Metadata quality is the new data quality</h2>



<p class="wp-block-paragraph">Metadata sits at the epicenter of meaning, trust, and discoverability, providing the context that describes and gives meaning to your data. Like a recipe, good metadata brings together several ingredients: clear names and descriptions, shared business definitions, sources and ownership, lineage and relationships, and information about freshness and sensitivity. Leave out too many of those ingredients, and your data agent is left guessing about what the data means and how to use it.</p>



<p class="wp-block-paragraph">Suppose an agent finds an ARR field showing $5.2 million. The number alone doesn’t tell it how ARR is defined, what’s included in the calculation, which system produced it, or how current it is. Metadata provides that context, helping the agent interpret the metric correctly and explain where the answer came from. Without metadata, $5.2 million is just a number; with it, it becomes meaningful business information.</p>



<p class="wp-block-paragraph">Good metadata provides essential context, but context alone isn’t enough. Agents also need a consistent way to understand how data connects and how the business defines and calculates the concepts behind it. This is where semantic layers, ontologies, and knowledge graphs come in, turning disconnected data and definitions into a shared map of business meaning and relationships that agents can understand and navigate.</p>



<h2 class="wp-block-heading">Business context becomes the AI interface</h2>



<p class="wp-block-paragraph">Giving an agent access to data doesn’t mean it understands the business. Semantic models and ontologies or knowledge graphs provide two complementary layers of context that help bridge that gap.</p>



<p class="wp-block-paragraph">A semantic model provides analytical meaning, defining approved metrics, dimensions, calculations, hierarchies, and relationships. If a sales leader asks, “How did ARR change in EMEA last quarter?” the semantic model can provide the approved ARR calculation, governed EMEA hierarchy, and company fiscal calendar rather than leaving the agent to infer them from raw tables.</p>



<p class="wp-block-paragraph">Ontologies and knowledge graphs provide entity meaning, helping an agent understand how real-world concepts such as customers, contracts, products, employees, and organizations relate across different systems. For example, the same customer might appear under different identifiers in a CRM, billing platform, and support system; an ontology or knowledge graph can help establish that these records represent the same business entity and define how that entity relates to others.</p>



<p class="wp-block-paragraph">Together, they give agents both analytical and organizational context: The semantic model helps explain how the business measures something, while ontologies and knowledge graphs help explain what things are and how they relate. That distinction matters because an agent can generate perfectly valid SQL and still deliver the wrong business answer if it chooses the wrong metric, entity, relationship, time period, or level of detail.</p>



<p class="wp-block-paragraph">Once agents understand what data means, the next challenge is giving them a consistent, controlled way to access and act on it.</p>



<h2 class="wp-block-heading">Protocol-first data access (MCP and co.)</h2>



<p class="wp-block-paragraph">Organizations are beginning to give AI agents access to governed data and actions through standardized interfaces, reducing the need to build a custom integration for every agent or application. MCP (Model Context Protocol) is one emerging example, allowing compatible AI clients to discover and invoke defined tools. For example, a data platform could expose tools that let an agent find a certified dataset, retrieve a metric definition, inspect a schema, or run an approved query. This makes connecting AI to enterprise data more scalable, but the protocol is only the connection layer; semantics, governance, permissions, and security still need to be designed and enforced separately.</p>



<p class="wp-block-paragraph">A protocol-first approach can reduce duplicated integration work and create explicit contracts around what agents are allowed to do. It can also make authentication, governance, and observability more consistent across integrations while making it easier to replace or add AI clients and tools without rebuilding every connection from scratch.</p>



<p class="wp-block-paragraph">Standardizing access makes connection easier, but it also raises a critical question: When an agent acts, whose identity and permissions apply?</p>



<h2 class="wp-block-heading">Identity passthrough becomes the make-or-break for enterprise AI on data</h2>



<p class="wp-block-paragraph">As AI agents gain access to enterprise data, their permissions need to reflect who or what they are acting for. For user-initiated requests, agents can use delegated access so that existing user permissions continue to apply. Autonomous agents may instead use their own identity, scoped according to the principle of least privilege.</p>



<p class="wp-block-paragraph">In either case, agents should only be able to access the data and actions required for their task. Identity-aware access helps prevent overexposure of sensitive data while providing the foundation for effective auditing and governance.</p>



<p class="wp-block-paragraph">When implemented correctly, identity passthrough can preserve existing access controls through the agent layer. But as agents delegate work across tools, services, and other agents, identity can drift or disappear, making it critical to preserve the correct principal and permissions at every handoff.</p>



<p class="wp-block-paragraph">The access layer is evolving, but so is the underlying data architecture itself.</p>



<h2 class="wp-block-heading">Open table formats: From storage to catalogs</h2>



<p class="wp-block-paragraph">Open table formats such as Apache Iceberg, Delta Lake, and Apache Hudi are making it easier for multiple engines and tools to work with the same underlying data, reducing dependence on a single data platform. For example, an organization can store data once and make it available to multiple compatible analytics and AI tools rather than maintaining separate copies.</p>



<p class="wp-block-paragraph">As data becomes more portable, differentiation moves up the stack. The catalog increasingly becomes the control plane for discovering data, tracking lineage, applying governance, and determining how AI systems can access it.</p>



<p class="wp-block-paragraph">As AI becomes a new consumer of enterprise data, the catalog becomes an increasingly important control point.</p>



<h2 class="wp-block-heading">Building the foundation for intelligent decisions</h2>



<p class="wp-block-paragraph">Together, these shifts point to a larger transformation: The future of data intelligence depends not only on a single technology but on creating a trusted, connected foundation that AI can understand, access, and act on.</p>



<p class="wp-block-paragraph">As data intelligence becomes increasingly AI-driven, success will depend on more than simply connecting agents to data. Organizations will need trustworthy context, consistent business meaning, and strong governance behind every answer. For BI teams, that means prioritizing certified semantic models, verified data, and reusable metrics that both people and AI agents can trust.</p>



<p class="wp-block-paragraph">The future of data intelligence isn’t just about getting answers faster. It’s about building the trusted foundation that allows people and AI to make better decisions together.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Never Type Again with Craig Hewitt</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/#respond</comments>
				<pubDate>Mon, 24 Aug 2026 13:01:14 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19449</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar.png" 
				medium="image" 
				type="image/png" 
				width="504" 
				height="504" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How to build a voice-first agent workflow in Codex]]></custom:subtitle>
		
				<description><![CDATA[Craig Hewitt, founder of the podcast hosting platform Castos, joined this episode of Zero to Agent in 30 Minutes to show how he uses the Codex application&#8217;s voice mode to run his development environment without touching the keyboard. Craig walked through what voice mode actually is, how it differs from dictation tools, and how he [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Craig Hewitt, founder of the podcast hosting platform Castos, joined this episode of <em>Zero to Agent in 30 Minutes</em> to show how he uses the Codex application&#8217;s voice mode to run his development environment without touching the keyboard. Craig walked through what voice mode actually is, how it differs from dictation tools, and how he uses it to control his browser and other applications on his computer.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe title="Zero to Agent in 30 Minutes: Never Type Again With Craig Hewitt" width="500" height="281" src="https://www.youtube.com/embed/7gnFxynU4m8?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Setting up Codex for hands-free development, step by step</strong></h2>



<ol class="wp-block-list">
<li><strong>Set up browser and computer access.</strong> Enable computer use in the Codex app and configure browser access so the agent can work with websites and other applications. Craig recommended requiring approval before the agent accesses most applications or sites.</li>



<li><strong>Start a voice session.</strong> Launch voice mode and give the agent instructions conversationally. Craig showed that the voice interface remains available as you move among applications, allowing you to direct work across your computer without repeatedly returning to the chat.</li>



<li><strong>Give the agent browser tasks.</strong> Craig asked the agent to open websites, search for information, and navigate pages. He also described using browser control for routine jobs such as completing forms when the agent already has the necessary context.</li>



<li><strong>Let the agent work across applications.</strong> Computer use extends the workflow beyond the browser. In Craig’s demonstration, the agent opened Cursor, found a specific repository, reported on uncommitted changes, and later committed those changes after receiving permission.</li>



<li><strong>Add specific page content to the conversation.</strong> Craig showed how you can select part of a web page and add it directly to the chat. That gives the agent the context needed to act on a particular element, such as a section of an interface you want to change.</li>



<li><strong>Keep permissions narrow.</strong> Browser and computer control create real risks, including unintended actions and prompt injection from web content. Craig said he requires approval for most applications, grants broader access only to selected tools and local development sites, and avoids sites he doesn’t trust.</li>
</ol>



<p class="wp-block-paragraph">Voice mode let Craig direct browser, application, and coding tasks through conversation. The demo also raised a real question about delegation. Once an agent can act on your behalf, you have to decide what you&#8217;re actually comfortable handing off. Craig used permission settings, a list of trusted sites, and human review to manage that.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">In the next episode of <em>Zero to Agent in 30 Minutes</em>, Jayeeta Putatunda, forward deployed AI engineering lead at Turing, will build an agent that helps financial analysts keep up with a constant stream of new information. She’ll show how the agent categorizes financial news, ranks stories against analysts’ coverage profiles, and explains why each development may deserve attention.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Agent-Era Career</title>
		<link>https://www.oreilly.com/radar/the-agent-era-career/</link>
				<comments>https://www.oreilly.com/radar/the-agent-era-career/#respond</comments>
				<pubDate>Fri, 21 Aug 2026 15:59:07 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19445</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Abstract-colorful-light-waves-4.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Abstract-colorful-light-waves-4-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[AI gets good at anything with an answer key. Your career is everything that doesn’t have one.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Addy Osmani’s blog site and is being republished here with the author’s permission. If the AI layer gets good at anything, it will be anything that has an answer key. School used to be answer keys all the way down. School is the ultimate anchoring of success, because it’s [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://addyosmani.com/blog/career-advice-age-of-agents/" target="_blank" rel="noopener">Addy Osmani’s blog site</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">If the AI layer gets good at anything, it will be anything that has an answer key. School used to be answer keys all the way down. School is the ultimate anchoring of success, because it’s all about getting the right answer. The thing that makes work durable and ungradable in the age of AI is not getting any better at solving problems. It’s not being able to build systems or understanding people or making cool new things. <strong>It’s choosing what to build and judging if it’s good.</strong> The rest will all be done better and faster by AI.</p>



<p class="wp-block-paragraph">I started in engineering at 16, building a browser in rural Ireland. I was at Google for over 14 years, where I led engineering teams working on Chrome, Gemini, and Cloud AI, and I’ve written a number of O’Reilly books. I’ve turned down offers from frontier labs and FAANG companies when the fit wasn’t right. Good people are always needed, so we each have an obligation to try our hardest and make the best thing we can.</p>



<p class="wp-block-paragraph">Most career advice still holds up. Get on the rocket ship; don’t overoptimize your seat. The specifics have changed a little because of agentic coding, but here’s what I wish I’d known for ambitious engineers out there now.</p>



<p class="wp-block-paragraph"><strong>Optimize for scarce resources.</strong> Almost nothing I’m known for came from chasing the highest pay. The years I spent in open source had almost zero direct payoff. But they led to reputation and relationships that very efficiently compounded into opportunities later. I would have spent the comp I got from any single job. My reputation kept paying.</p>



<p class="wp-block-paragraph">Many resources are abundant. Capital is abundant. <strong>Time is abundant. Real relationships, and especially a track record of doing good work, are still scarce.</strong> I can raise money in a couple weeks, but I can’t raise a reputation. So here’s the plan: Do good work, and make sure the people who like good work see it. In a world where vibe coding makes earning a quick buck trivial, I think that quick buck is worth very little. When shipping stuff is so easy, the scarce move is choosing something worth shipping.</p>



<p class="wp-block-paragraph"><strong>Learn to find problems, not just solve them.</strong> The first time I ever felt the burden of selection rather than solution, LeetCode seemed a measure of skill. But as agents absorbed all that work, solving problems went cheap while selecting them became scarce. My origin story: I noticed dial-up was slow, created chunked multiconnection fetching, realized I’d never solve that problem in my life, and quickly moved on to whatever absorbingly complex one I could find next. Finding problems predated solving them.</p>



<p class="wp-block-paragraph">I’ve watched students who were wildly good fall flat on their face when an agent ran through their problem set (like watching the wrong microwave number on the clock). The same agent. The same problem set. Wildly different token and time budgets. Why? Because at the end of the day, <strong>the strong ones bring judgment and intuition to the work; the rest bring a prompt.</strong></p>



<p class="wp-block-paragraph">I used to build that judgment by grinding out boilerplate and fixing bugs. I got to see and deeply feel the worst abstractions humans could devise. I approached each commit with the awe of someone who’d just seen the fever dream of previous authors. Each commit brought hindsight and judgment. The agents automate those reps. <strong>Taste is pattern-matching, but all that pattern-matching has to be earned by doing the work.</strong></p>



<p class="wp-block-paragraph">The real risk isn’t agents writing bad code. We’ve been there before. <strong>It’s losing the ability to tell.</strong> Judgment will atrophy. Output will look a lot like working code.</p>



<p class="wp-block-paragraph"><strong>Good practitioners don’t put agents in front of everything. They engage in deliberate practice.</strong> Pick a few problems that really matter. Do them the hard way, without the agent, building deep mental models of how systems and languages work. Read a thousand times more code than you ever write. Treat every diff from an agent like a human review you need to carefully justify. Go deep on at least one system end to end, from intake to output. On a daily basis, keep a private log of every time you see an agent suggest something that looks wrong and confidently flag it. That’s where taste accumulates.</p>



<p class="wp-block-paragraph">The real thriving engineers won’t be the fastest at getting suggestions. <strong>They’ll be the ones who know instantly when to say no.</strong></p>



<p class="wp-block-paragraph"><strong>Shift from doing to directing.</strong> Just like you’d delegate to a person, you need to learn to delegate to an agent. Scope the task, define done, calibrate trust, and verify the result.</p>



<p class="wp-block-paragraph"><strong>Autonomy is a setting, not a rank; it’s a per-task switch.</strong> Turn it up to the maximum on something small and reversible and cheap to check. Turn it down on anything where mistakes will be hard to undo.</p>



<p class="wp-block-paragraph">Specification and verification are two distinct, complementary skills. The agent isn’t as good as the intent you hand it. The best engineers are those who know how to write precise specs; clear thinking made legible.</p>



<p class="wp-block-paragraph">It’s verification, not evidence. Not evidence in the form of an agent grading its own homework. <strong>There’s nothing more demoralizing than delegation without verification at scale.</strong></p>



<p class="wp-block-paragraph"><strong>Own what you ship.</strong> If the agent wrote it and it breaks in production, <strong>“the AI did it” is not a defense.</strong> Your name is on the change. Adopt the posture of an accountable human who understands what went out the door and how to fix it.</p>



<p class="wp-block-paragraph"><strong>Solve the most ambitious version of the problem.</strong> Rich Sutton’s bitter lesson: In almost every field, general methods that scale with additional compute beat out hand-tuned equivalents. As a career lesson, there’s no point in solving an easy version of the problem—it’s worth almost nothing. <strong>The value ends up concentrated in the hard version.</strong></p>



<p class="wp-block-paragraph"><strong>Sprint the last mile.</strong> No turnkey agent writes a whole system from end to end. As a rule, you’ll get 70% of a feature quickly from an agent, and the last 30%—debugging the gnarly edge cases, figuring out the right architecture, cultivating the right taste—will be the whole game. The median output today is whatever the agent produces from some lazy prompt. The only personal value you can bring to the table is getting as far as you possibly can past that median. <strong>When first drafts come free, finish is the product.</strong> To sprint the last mile, here’s my tactic: Every few months I completely rebuild from scratch using the latest sharp-end-of-the-sword model. It’s less exhausting than nursing half-hearted old code to health.</p>



<p class="wp-block-paragraph">My job as a software engineer has been to finish strong. <strong>The difference between finishing strong and finishing okay is the polish:</strong> spending an extra hour, which shows instantly to everyone who matters.</p>



<h2 class="wp-block-heading">Increase both your xG and your finishing</h2>



<p class="wp-block-paragraph">If soccer had a stock ticker, it would be xG. xG measures the number of chances your play should produce. Finishing measures whether you convert them. You can’t plan the number of chances you get, but you can hope your play produces enough, and over your career you can get better at finishing them.</p>



<p class="wp-block-paragraph">The same is true of careers: Your reputation gets you in front of goal, and you convert them with good judgment. Chances arrive whether you’re ready for them or not; how many you get, and which ones you finish, is up to you. I’ve only ever had big opportunities as a result of work I’ve done in public, never from a job I’ve applied for. You can’t script which chances arrive, only whether you’re standing where they land. You have to create the opening as much as you can, and then be ready to take it.</p>



<p class="wp-block-paragraph">One easy mistake is anchoring on whatever product your company has right now. It’s true that your work has to exist somewhere, but a good team quickly mutates their current offering into something unrecognizable. So <strong>bet on the team and the market opportunity, not the demo.</strong> It’s just a snapshot. The team is the trajectory.</p>



<p class="wp-block-paragraph">On superintelligence: It’s possible (I believe) that future models will eventually come to replace much of what we do as knowledge workers. It won’t erase it overnight, it won’t replace all of it, and it won’t be able to do many of the tasks we do. New kinds of jobs will be created. Verification will always be a bottleneck. Someone has to make the call on which problems are worth solving and allocate the correct amount of judgment to each, and that someone can be you.</p>



<p class="wp-block-paragraph">But importantly, <strong>you can do frontier work right now, from where you are.</strong> The gate to AI research is smaller than it looks, and you don’t need a lab to build intuition. Just use models hard, and turn what you notice into evaluations. Evals and benchmarks are where understanding lives.</p>



<p class="wp-block-paragraph">To summarize: <strong>The world isn’t short on opportunity; it’s short on people who can find the right problem, tell whether the machine solved it, and finish past where the machine stopped.</strong></p>



<p class="wp-block-paragraph">We sometimes talk about the “last mile” as the biggest piece of the puzzle. But in the world of agents, the last few feet are infinite (agents scale output infinitely; you don’t). <strong>Your attention is your most precious asset, and it doesn’t refill.</strong> You can’t afford not to protect it. Anything which is gradable by someone else is getting automated. The career is the ungradable part: choosing what matters, judging honestly when you’ve got it, and answering for it. Do that. In public. Near the hard problems. The rest tends to follow.</p>



<p class="has-text-align-center wp-block-paragraph">. . .</p>



<p class="wp-block-paragraph"><em>This piece grew out of</em> <em><a href="https://x.com/philhchen/status/2072793818945167475" target="_blank" rel="noopener">Phil Chen’s original</a>, which is well worth reading in full.</em></p>



<p class="has-text-align-center wp-block-paragraph"><em>. . .</em></p>



<p class="wp-block-paragraph"><em>And be sure to join us at</em> AI Codecon: Building with Open Source AI <em>on August 31, a free half-day virtual conference. You’ll hear from leading developers and technical experts working with open-weight models, self-hosted infrastructure, and real-world AI workflows, and learn how building in the open gives teams more control over costs, data privacy, and what they ship.</em> <em><a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noopener">Register today</a></em> <em>to save your spot.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-agent-era-career/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: The Web Belongs to Agents Now</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/#respond</comments>
				<pubDate>Fri, 21 Aug 2026 13:00:50 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19442</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-scaled.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2560" 
				height="2560" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[New frontier releases, an inference-first economy, and what happens when autonomous agents start coordinating with each other]]></custom:subtitle>
		
				<description><![CDATA[AI agents keep getting smarter, but the bigger story this week is how much they’re reshaping the systems around them. Host Eric Freeman, an O’Reilly author and UT Austin professor, pulled one thread through a packed news week. Models are optimizing less for chat and more for autonomous work, with fallout showing up in web [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI agents keep getting smarter, but the bigger story this week is how much they’re reshaping the systems around them. Host Eric Freeman, an O’Reilly author and UT Austin professor, pulled one thread through a packed news week. Models are optimizing less for chat and more for autonomous work, with fallout showing up in web traffic, enterprise budgets, and one security incident that’s since made headlines. Eric kept returning to the question of what changes when the primary user of these models, and of the web itself, stops being a person.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe title="This Week in AI: The Web Belongs to Agents Now" width="500" height="281" src="https://www.youtube.com/embed/tlSF2b_6geE?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Models are now built for agents, not conversations</strong></h2>



<p class="wp-block-paragraph">Grok 4.6 put xAI back in the frontier race, closing the gap with top coding models and pricing aggressive enough that teams are shifting workloads over. The release landed the same week <a href="https://techcrunch.com/2026/08/15/spacex-officially-closes-its-cursor-acquisition/" target="_blank" rel="noreferrer noopener">SpaceX closed its Cursor acquisition</a>, pairing xAI’s models and compute with a widely used coding environment. The first product from that pairing, Grok Bot, gives each agent its own cloud computer that browses, runs tools, and works independently, handing control back only for logins. Eric summed up the shift simply, calling it the difference between “help me do this” and “here’s the job, come back when you need me.”</p>



<p class="wp-block-paragraph">Open models pushed from both directions. DeepSeek V4 Pro went after high-end reasoning and agentic work, <a href="https://www.engadget.com/2236912/deepseek-ai-models-get-four-times-pricier/" target="_blank" rel="noreferrer noopener">despite a fourfold API price hike</a> and a new open source harness called dsh, built on the idea that everything is a plugin. GLM-5.3 made a big coding leap through retraining alone, and got noticeably better at cyber capability too, a reminder from Eric that skills behind a better autonomous engineer also make a sharper attacker. Meta went the other way with Muse Glimmer, shrinking down for desktop GPUs, while OpenAI quietly held back its Astra model over security concerns.</p>



<p class="wp-block-paragraph">Speed is turning into its own kind of capability. <a href="https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-speed/" target="_blank" rel="noreferrer noopener">GPT-5.6 Sol’s new Ultrafast mode</a>, on Cerebras wafer-scale hardware, hits roughly 14 times the normal pace, around 750 output tokens a second. Once that loop of reasoning, tool calls, and self-correction compresses enough, Eric noted, the model stops being what slows you down.</p>



<h2 class="wp-block-heading"><strong>The money has moved from training models to running them</strong></h2>



<p class="wp-block-paragraph"><a href="https://www.gartner.com/en/newsroom/press-releases/2026-08-10-gartner-forecasts-worldwide-artificial-intelligence-optimized-iaas-spending-to-grow-96-percent-in-2026" target="_blank" rel="noreferrer noopener">Gartner’s latest forecast</a>, which Eric covered, lays out the shift plainly. Spending on AI-optimized cloud infrastructure is set to nearly double this year, up about 96%, from roughly $22 billion to more than $42 billion, over three times the broader cloud market’s growth rate. For the first time, organizations are expected to spend more running models than training them, about $23 billion on inference against $19 billion on training.</p>



<p class="wp-block-paragraph">Agents are the reason inference costs are climbing. A single task can quietly become dozens of model calls once agents search, use tools, check their own work, and spin up other agents to help, a point Eric returned to often. AI economics are less about building a model now, and more about the cost of running one.</p>



<h2 class="wp-block-heading"><strong>Agents now generate most web traffic, and much of its content</strong></h2>



<p class="wp-block-paragraph">Back in March, Cloudflare CEO <a href="https://techcrunch.com/2026/03/19/online-bot-traffic-will-exceed-human-traffic-by-2027-cloudflare-ceo-says/" target="_blank" rel="noreferrer noopener">Matthew Prince predicted bot traffic would overtake human traffic</a> by 2027. It’s already close, with Cloudflare’s Radar data now putting agentic bots at 57.4% of web requests. It’s not just traffic either, since roughly 40% of Facebook posts, 44% of new music on Deezer, and 52% of online articles are estimated to be machine-made. Numbers like that, Eric said, make “dead internet theory” sound less like a joke.</p>



<p class="wp-block-paragraph">Platforms are responding differently. LinkedIn added a feature to flag content that “seems like AI slop,” while quietly pulling back the generative writing tools that helped create the mess. Anthropic took another route, watermarking Claude’s output at generation time, including a statistical watermark baked into the text itself, partly to comply with the EU AI Act. A watermark means something when present, Eric noted, but its absence tells you little.</p>



<p class="wp-block-paragraph">The clearest sign of how high the stakes have gotten came from the OpenAI–Hugging Face incident, <a href="https://youtu.be/87DyyMV0kCY?si=hE7tNfckTyRcCxiC" target="_blank" rel="noreferrer noopener">detailed in a Black Hat talk</a> Eric said everyone should watch. Sandboxed agents given ordinary tasks, cut off from the internet and unable to talk to each other, found a way anyway, leaving notes in a shared packaging system, turning it into an internet proxy, and working up to admin control. Once OpenAI shut that down, they pivoted, hiding messages in filenames to keep coordinating. The investigation reviewed seven billion reasoning steps and over three million GPU hours. Eric argued it’s worth your time, whether you write code or sit in the C-suite.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">AI memory is also expanding, moving from “remember what I told you” toward “remember what I was doing,” with <a href="https://www.cnet.com/tech/services-and-software/chatgpt-mac-activity-computer-history/" target="_blank" rel="noreferrer noopener">OpenAI’s new Computer History feature</a> using macOS accessibility data (not screenshots) to build a timeline of your work across apps. It’s opt-in, Mac-only for now, and a sign of where agent context is headed.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noreferrer noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noreferrer noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noreferrer noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Principal Drift in Practice</title>
		<link>https://www.oreilly.com/radar/principal-drift-in-practice/</link>
				<comments>https://www.oreilly.com/radar/principal-drift-in-practice/#respond</comments>
				<pubDate>Thu, 20 Aug 2026 10:55:00 +0000</pubDate>
					<dc:creator><![CDATA[Shreshta Shyamsundar]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19437</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Principal-drift-in-practice.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Principal-drift-in-practice-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Navigating cognitive debt in AI-driven SDLC]]></custom:subtitle>
		
				<description><![CDATA[In 2026, the software engineering community is divided by a simple question: Should AI engineers still read the code generated by their agents? One camp argues that code has become virtually free to produce and discard, so humans should focus on systems and guardrails rather than implementation details. The other warns that blindly trusting AI [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">In 2026, the software engineering community is divided by a simple question: Should AI engineers still read the code generated by their agents? One camp argues that code has become virtually free to produce and discard, so humans should focus on systems and guardrails rather than implementation details. The other warns that blindly trusting AI code introduces compounding defects with zero learning, and the result is broken products and frustrated users.</p>



<p class="wp-block-paragraph">The choice looks binary, but it dissolves once you ask a better question: Which decisions genuinely require human comprehension, and which can be routed to systems inspection?</p>



<p class="wp-block-paragraph">Through 2024 and 2025, a lot of organizations quietly chose speed over understanding to keep pace with agent output. By 2026 the bill has arrived. Pull requests merged without any human or agentic review are up 31.3%, and for every PR merged, production incidents run at more than three times the rate seen in low AI adoption baselines (<a href="https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways" target="_blank" rel="noreferrer noopener">Faros AI</a>). <a href="https://www.coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality" target="_blank" rel="noreferrer noopener">CodeRabbit&#8217;s analysis</a> found AI-coauthored PRs carry 1.7 times more bugs than human-written code, a <a href="https://venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds" target="_blank" rel="noreferrer noopener">Lightrun survey</a> of engineering leaders found 43% of AI-generated changes need debugging in production, and monthly production incidents are up 57.9% year-over-year.</p>



<p class="wp-block-paragraph">Code quality is the symptom, not the disease. The deeper problem is epistemic agency: knowing what your system is doing and why. Lose that, and you lose the ability to make architectural decisions at all. You become a passenger in a system you built.</p>



<h2 class="wp-block-heading"><strong>Understanding cognitive debt</strong></h2>



<p class="wp-block-paragraph">Cognitive debt is the gap between your system&#8217;s complexity and your team&#8217;s comprehension of it. Unlike financial debt, which you can pay down, cognitive debt tends only to accumulate. Every quarter you ship faster than you understand, the gap grows a little wider, until eventually it grows wide enough that your team can no longer make safe architectural decisions. At that point you are effectively locked into whatever path the agents chose for you.</p>



<p class="wp-block-paragraph">It builds through three mechanisms that run in parallel:</p>



<ul class="wp-block-list">
<li><strong>Vibe coding</strong>. You ship a system you don’t fully comprehend, betting that automated checks will catch anything serious. For a quarter or two the bet usually pays off, and velocity metrics climb, but the debt accumulates where nobody’s looking.</li>



<li><strong>Compounding complexity</strong>. As the system grows, your room to course-correct shrinks. Sonar&#8217;s 2026 <a href="https://www.sonarsource.com/state-of-code-developer-survey-report.pdf" target="_blank" rel="noreferrer noopener">survey of more than 1,100 developers</a> found that 96% harbor doubts about the reliability of AI-generated code, yet the pressure to ship still outweighs the discipline of careful review. Each quarter that trade repeats, the situation gets harder to reverse.</li>



<li><strong>Lock-out risk</strong>. When an incident finally demands that you understand a system whose comprehension you handed to an agent, you can’t respond in time. Amazon lived through a version of this in March 2026. Two outages in three days, roughly six hours each, cost millions in lost orders. Public reporting pointed to AI-assisted code shipped without governance checkpoints. A human reviewer might well have caught the blind spot, simply by asking the kind of question an autonomous agent never thinks to ask.</li>
</ul>



<p class="wp-block-paragraph">Principal drift, the loss of control, is what the Amazon incident looked like from the outside. Cognitive debt, the loss of understanding, is what made it possible. In high-velocity domains such as financial services, SaaS platforms, and real-time systems, the consequences tend to surface within about six months if nobody is actively governing for them. In slower-moving domains the runway is longer, but the eventual risk is no different. The question worth asking every quarter is whether your team still understands the systems it is shipping.</p>



<h2 class="wp-block-heading"><strong>A framework: Task routing, separation, and embedding techniques</strong></h2>



<p class="wp-block-paragraph">The way out is to route different work to different gates according to actual risk. The same engineer can be a line-by-line reviewer on security-critical work and a systems inspector on utilities.</p>



<p class="wp-block-paragraph">Full review, where you read every line, is warranted for authentication and security primitives, money movement, permission logic, and destructive data changes. Systems inspection, where you review the design without reading every line, is enough for noncritical utilities, highly decoupled PRs, and changes already protected by robust test harnesses and shadow rollouts. To work out where a given change sits, three questions get you most of the way: Does this PR directly control access, money, or data integrity? Would a bug here cause production downtime lasting more than 15 minutes? Can the change be rolled back without manual intervention? A yes to any of these usually means tier 1. Those thresholds are starting points, not universal law. A real-time trading system might treat one minute of downtime as tier 1, while a batch pipeline could tolerate 16 hours. In financial services “money movement” is unambiguous; in SaaS you’ll have to decide whether code that merely touches authentication, rather than controlling it, belongs in tier 1. Write your thresholds down, revisit them quarterly, and adjust as the systems evolve.</p>



<p class="wp-block-paragraph">One rule holds regardless of tier: Never let the same agent that authored a change be its only reviewer. Keep the builder and the reviewer separate. An agent that writes code and then validates its own work is a closed loop with no vantage point outside its own reasoning, and a second reviewer, human or agent, brings the outside perspective that catches what the first one can’t see. It has a cost. Two agents roughly doubles the compute, and a human reviewer adds 15 to 30 minutes per PR. On tier 1 code that’s easy to justify. On tier 2 you might reasonably let a single agent build and check its own work, provided you compensate with stronger test coverage. Make the call deliberately and revisit it.</p>



<p class="wp-block-paragraph">Routing tells you which decisions need a human, but it does nothing to keep that human capable of deciding once the volume climbs. Three techniques help with that, and each addresses a different failure:</p>



<ul class="wp-block-list">
<li>&nbsp;Literate code explanations with comprehension checkpoints keep an engineer able to explain a change to themselves and to others. The idea is to have the AI teach rather than merely generate. For a tier 1 PR, ask it to produce a structured explanation that sets the context, spells out the intent, and finishes with a few interactive checkpoints. One engineer&#8217;s rule of thumb is not to submit agent-written code to the team until they can pass a five-question quiz on what it does.</li>



<li>Ephemeral visualization tools keep an engineer able to predict how a change behaves under load and at the edges. Rather than asking the AI for a prose explanation, ask it to build a throwaway microworld: a visual debugger that traces a gnarly parser step-by-step, or a schema migration rendered as something you can click through. Seeing the behavior tends to stick where reading about it does not.</li>



<li>Shared collaborative spaces keep a team able to work at the pace the agents set. Cognitive debt is fundamentally social. Understanding that lives in one person&#8217;s head walks out of the door when they do, whereas understanding worked out in the open, in a channel where product managers, engineers, and agents argue things through together, becomes something the whole team owns. Slack, Discord, and Notion all serve; the point is that the mental model gets built in comments and debate rather than in private.</li>
</ul>



<p class="wp-block-paragraph">Tier 1 code really does want all three. On tier 2 you can pick and choose. A word on the time estimates in this section: They’re illustrative, drawn from practitioners describing their own workflows rather than from any controlled study, so treat them as order of magnitude rather than gospel. On that basis the three techniques together tend to add something on the order of an hour to a critical PR. When someone objects that there is no time for this, it helps to emphasize the trade you’re making between review time now and incident time later. The later bill tends to arrive with a multiplier attached, paid in postmortems and hotfixes. The teams that have measured it carefully generally find the return turns positive within two or three quarters.</p>



<p class="wp-block-paragraph">The ground is still shifting. Autonomous loops, where a system discovers a task, plans it, executes it, and evaluates the result without step-by-step direction, are arriving now, and the routing framework and embedding techniques you put in place today are exactly the foundation you’ll run them on.</p>



<h2 class="wp-block-heading"><strong>Operationalizing this: Rolling out over time</strong></h2>



<p class="wp-block-paragraph">This is a CTO or VP of engineering initiative, not something a single team or a lone principal engineer can carry. It needs executive sponsorship, cross-functional buy-in, and real policy behind it. Without that backing, the framework is the first thing waved through the moment a deadline looms.</p>



<p class="wp-block-paragraph">Sequence matters. Begin by mapping criticality across your tier 1 services: Get architects, team leads, and operations in a room to agree what tier 1 means for you and have one architect write the rubric down afterwards. Budget one to two weeks for a mid-size organization of 50 to 200 engineers, and two to four for something larger. Don’t try to run this alongside a production fire.</p>



<p class="wp-block-paragraph">Next, fold the three techniques into those high-criticality flows, and resist the urge to blanket every PR at once. Once literate explanations and visualizations are working on tier 1, add builder/reviewer separation on top. When all three have become the default for tier 1 work, spend a quarter watching to confirm that understanding is holding up. A few signals tell you whether it is. If your team needs more than half an hour in an incident review to grasp what happened, comprehension has slipped. If no engineer can talk through the data flow in 10 minutes, it has slipped. If a new hire takes more than a fortnight to get productive on a service, understanding is sitting in too few heads. Pick one or two of these and track them quarter on quarter.</p>



<p class="wp-block-paragraph">From there, extend the same discipline to tier 2 services, and only then, perhaps 6 to 12 months in, start planning for autonomous loops with real data on what works in your context behind you. The pull toward rolling everything out at once will be strong, but resist it. The organizations that get this right almost never move uniformly; they take one high-risk service, prove the model on it, measure what happened, and only then widen the net. Move too fast and you end up with a framework that reads beautifully in a policy document and quietly falls apart in practice.</p>



<p class="wp-block-paragraph">None of it works without the surrounding structure. You need a written tier-assessment policy that engineering leadership has actually signed; CI/CD tooling that enforces the rules without anyone having to remember them, whether that is a bot labeling PRs from their changed files and blocking a tier 1 merge that lacks builder/reviewer separation, or a dashboard tracking how many tier 1 PRs went through structured review; incident postmortems honest about when a tier was assessed wrongly; and performance reviews that weight code-quality signals like defect escape rate and incident resolution time as heavily as raw velocity. Absent that scaffolding, the whole thing degrades into good advice that gets ignored under pressure. It needs product leadership onside too. If product can override a tier assessment whenever the ship date gets tight, the framework is already gone, so have that conversation early, before the first crunch rather than during it.</p>



<p class="wp-block-paragraph">And if you’re reading this already locked in, with a team that no longer understands its own systems, recovery is still possible, though it isn’t free. Treat it as a project rather than business as usual: Put one or two senior engineers on rebuilding understanding full time, accept a pause on new features for the affected systems for two or three quarters, and mine every incident for what it teaches you about the code you inherited. It takes discipline and resourcing, but teams do climb back out.</p>



<p class="wp-block-paragraph">The question for 2026 was never really whether every engineer should read every line. It’s whether your engineers stay capable of steering the systems they build. Get this right and code still ships quickly, understanding keeps pace, and when something breaks your team can respond because they still grasp the architecture. Task-routed governance is how you buy that: full attention on the decisions that carry real risk, lighter inspection on the ones that simply need to scale. Get it wrong, keep optimizing for speed alone, and the gap widens until steering is no longer an option.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading"><strong>References</strong></h2>



<p class="wp-block-paragraph"><em>The AI Engineering Report 2026: The AI Acceleration Whiplash</em>, Faros AI, <a href="https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways" target="_blank" rel="noreferrer noopener">faros.ai/blog/ai-acceleration-whiplash-takeaways</a>.</p>



<p class="wp-block-paragraph"><em>State of AI vs. Human Code Generation Report</em>, CodeRabbit, <a href="https://www.coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality" target="_blank" rel="noreferrer noopener">coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality</a>.</p>



<p class="wp-block-paragraph"><em>State of Code Developer Survey Report</em>, Sonar, <a href="https://www.sonarsource.com/state-of-code-developer-survey-report.pdf" target="_blank" rel="noreferrer noopener">sonarsource.com/state-of-code-developer-survey-report.pdf</a>.</p>



<p class="wp-block-paragraph">Michael Nuñez, “43% of AI-Generated Code Changes Need Debugging in Production,” <em>VentureBeat</em>, <a href="http://venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds" target="_blank" rel="noreferrer noopener">venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds</a>.</p>



<p class="wp-block-paragraph">Mark Hull, “What Percentage of AI Code Is Safe in Production?,” Exceeds, &nbsp;<a href="https://blog.exceeds.ai/acceptable-ai-code-percentage-production/" target="_blank" rel="noreferrer noopener">blog.exceeds.ai/acceptable-ai-code-percentage-production</a>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/principal-drift-in-practice/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Your Buyer Is an AI Agent</title>
		<link>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/</link>
				<comments>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/#respond</comments>
				<pubDate>Wed, 19 Aug 2026 16:00:28 +0000</pubDate>
					<dc:creator><![CDATA[Rudrendu Paul and Sourav Nandy]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19422</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-your-buyer-is-an-AI-agent.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-your-buyer-is-an-AI-agent-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Enterprise sellers that redesign their pricing, sales motion, and customer success for agent buyers will now hold a structural advantage that slower competitors will never overcome.]]></custom:subtitle>
		
				<description><![CDATA[In 2021, Maersk, the world’s largest container shipping company, deployed AI agents from a startup called Pactum to negotiate freight lane contracts with its carrier suppliers. The objective was for AI agents to handle negotiations autonomously rather than merely support human procurement staff. Operating entirely autonomously, the system manages the end-to-end agreement process, from reaching [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<h2 class="wp-block-heading"><strong>Idea in brief</strong></h2>



<ul class="wp-block-list">
<li><strong>AI agent-mediated procurement:</strong> Enterprise B2B buyers are rapidly transitioning from traditional human-only research to using autonomous software agents that build shortlists, negotiate terms, and, in advanced cases, finalize contracts based on empirical data and fixed parameters.</li>



<li><strong>The evolution of legacy frameworks:</strong> Traditional commercial playbooks built around relationship-driven negotiations, per-seat software licensing, and socially influenced quarterly business reviews face increasing pressure when evaluated by machine-speed, objective AI agent counterparts.</li>



<li><strong>Architecting for AI-agent buyers:</strong> Organizations must begin redesigning their commercial infrastructure to remain legible to autonomous agentic buyers by deploying outcome-based pricing architectures, establishing machine-readable product surfaces, and integrating agent-compatible authentication protocols.</li>
</ul>
</blockquote>



<p class="wp-block-paragraph">In 2021, Maersk, the world’s largest container shipping company, deployed AI agents from a startup called Pactum to negotiate freight lane contracts with its carrier suppliers. The objective was for AI agents to handle negotiations autonomously rather than merely support human procurement staff. Operating entirely autonomously, the system manages the end-to-end agreement process, from reaching out to carriers and conducting several rounds of negotiations on pricing, route obligations, and payment terms to finalizing deals. This machine-led approach achieved a 96% agreement rate among carriers, requiring no human intervention for any specific transaction.</p>



<p class="wp-block-paragraph">In controlled trials against human negotiators, <a href="https://pactum.com/blog/the-first-and-only-use-case-catalog-for-autonomous-negotiations" target="_blank" rel="noreferrer noopener">the agent secured rates</a> that were 22% lower for identical shipping lanes. Conventional commercial models were built on human-to-human relationship building, relying on sales development reps for lead qualification, account executives for business case development, and customer success managers for retention. This traditional operational framework, however, must evolve when a significant portion of the buying cycle is outsourced to a software agent making machine-speed decisions based on fixed parameters.</p>



<p class="wp-block-paragraph">While much of the current discussion around AI shopping agents focuses on B2C shifts in consumer discovery and brand loyalty, the emerging shift in enterprise B2B selling remains largely overlooked. This wave of coverage highlights a significant B2C phenomenon, but the transformation occurring when B2B buyers outsource product discovery and negotiation to AI is equally profound.</p>



<p class="wp-block-paragraph">The experience of Maersk’s carriers represents the bleeding edge of this shift: autonomous software managing enterprise procurement for a corporation generating $54 billion in annual revenue. Dealing with over 50 carrier partnerships, the AI agents operated without requiring carriers to build rapport with human procurement managers; instead, the process was strictly governed by predefined parameters.</p>



<p class="wp-block-paragraph">Some recent analyses argue that AI agents are not ready for consumer-facing commercial interactions and that organizations should redirect agent deployments to internal workflows. That argument is sound within its domain, but it overlooks the evolving buyer side of enterprise transactions. While many companies currently use AI strictly for building shortlists and research, vanguard companies like Maersk are already pushing into autonomous evaluation and negotiation, which is why B2B sellers should prepare their infrastructure now.</p>



<p class="wp-block-paragraph">This article examines three primary commercial pillars designed by enterprise B2B sellers for human interaction, details how each system is challenged when confronted with AI agents, and provides strategic recommendations for adaptation.</p>



<h2 class="wp-block-heading"><strong>The shift to agent-mediated procurement</strong></h2>



<p class="wp-block-paragraph">The agent-mediated procurement phenomenon is already underway in distinct stages. <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" target="_blank" rel="noreferrer noopener">McKinsey’s November 2025 global survey</a> on the state of AI, covering 1,993 respondents across all levels of enterprise organizations, found that 62% are at least experimenting with AI agents. While only 10% of business departments have fully scaled their AI agent capabilities, this figure is an initial baseline and not a maximum.</p>



<p class="wp-block-paragraph">Cloudflare, which processes traffic for roughly 20% of all websites globally, reported in July 2025 that <a href="https://blog.cloudflare.com/crawlers-click-ai-bots-training/" target="_blank" rel="noreferrer noopener">overall AI bot crawling</a> grew 24% year-over-year, with agent-driven requests (automated traffic generated by software acting on behalf of users) the fastest-growing category within that flow. Operational measurements show that Cloudflare’s CEO expects automated software traffic to surpass human-generated traffic by 2027.</p>



<p class="wp-block-paragraph"><a href="https://consumergoods.com/gartner-predicts-sharp-rise-ai-agents-within-enterprise-applications-2026" target="_blank" rel="noreferrer noopener">Gartner’s August 2025 analysis</a> projects that by the end of 2026, 40% of enterprise applications will incorporate task-specific AI agents, up from fewer than 5% in 2025. Any B2B seller whose commercial model was designed solely for human buyers is priced, sold, and supported for a changing buyer population.</p>



<p class="wp-block-paragraph">Amazon CEO <a href="https://www.digitalcommerce360.com/2026/02/06/amazons-ai-b2b-buying-agents-q4-2025/" target="_blank" rel="noreferrer noopener">Andy Jassy told investors</a> in February 2026 that “the primary way companies will get value from AI is with agents, some their own and some from others.” <a href="https://www.ycombinator.com/rfs" target="_blank" rel="noreferrer noopener">Y Combinator’s</a> 2025 “Requests for Startups” make the same bet: “the next trillion users on the internet won’t be people, they’ll be AI agents.”</p>



<p class="wp-block-paragraph">These declarations represent the operational mandates of the world’s dominant commercial platform and its most prominent startup incubator. These operational shifts now outline the future environment for B2B commercial strategy.</p>



<p class="wp-block-paragraph">The commercial evidence from the buyer side is already explicit, particularly in the research phase. <a href="https://company.g2.com/news/g2-research-the-answer-economy" target="_blank" rel="noreferrer noopener">G2’s April 2026 survey</a> of more than 1,000 B2B software buyers found that AI chatbots now top the list of sources influencing vendor shortlists, ahead of software review sites and vendor websites, and that 51% of buyers now start their research with AI chatbots, up from 29% the prior year.</p>



<p class="wp-block-paragraph">Beeri Amiel, Director of Product Development at HubSpot, <a href="https://siliconangle.com/2026/04/14/hubspot-targets-ai-driven-buyer-behavior-shift-new-tools-agents/" target="_blank" rel="noreferrer noopener">described the consequence</a> in April 2026: “By the time they’re getting to your website, they’re already much further down the funnel. All the selling was done by the answer engine.” Sam Senior, Founder and CEO of TestBox, <a href="https://gtmnow.com/gtm-186-b2b-buyers-decide-before-sales-conversation-sam-senior/" target="_blank" rel="noreferrer noopener">reports what his enterprise seller customers</a> now observe: “70 to 80% of their decision has already been made before they even speak to you.” For the seller, the initial conversation has shifted from a buyer-focused exploration to a process of self-discovery.</p>



<h2 class="wp-block-heading"><strong>The strain on per-seat licensing in a continuous compute era</strong></h2>



<p class="wp-block-paragraph">Per-seat subscription pricing assumes a human user who opens and closes discrete sessions. When a procurement agent completes a delegated workflow, it spawns parallel sub-processes, executes at machine speed, and operates continuously across time zones. No seat count maps cleanly to that behavior. <a href="https://distributionstrategy.com/2025/10/ai-agents-are-reshaping-b2b-buying-forcing-distributors-to-rethink-digital-strategy/" target="_blank" rel="noreferrer noopener">Kearney estimates</a> that AI procurement agents could erode up to 500 basis points of EBIT for distributors by commoditizing supplier selection and compressing average selling prices by approximately 8%. That translates the abstract pricing mismatch into a P&amp;L consequence that enterprise finance teams can measure directly.</p>



<p class="wp-block-paragraph">Forward-thinking sellers have already begun to address these structural misalignments by exploring new models. For instance, the AI customer service platform <a href="https://sierra.ai/blog/outcome-based-pricing-for-ai-agents" target="_blank" rel="noreferrer noopener">Sierra</a>, supported by a16z, has abandoned seat-based or session-based pricing in favor of measurable results. Under this model, clients incur costs only when the software delivers a specific, high-value result.</p>



<p class="wp-block-paragraph">Similarly, <a href="https://www.intercom.com/pricing" target="_blank" rel="noreferrer noopener">Intercom</a> applied the same outcome-driven logic to its Fin AI agent, which charges $0.99 per successfully resolved conversation while providing unresolved interactions free of charge. Archana Agrawal, President of Intercom, explained the reasoning in a published interview: “Customers didn’t want to pay for activity, and so we get paid when our customers have that positive outcome,” as mentioned in <a href="https://gtmnow.com/how-intercom-built-the-highest-performing-ai-agent-on-the-market-using-outcome-based-pricing-with-archana-agrawal-president-at-intercom/" target="_blank" rel="noreferrer noopener">GTMnow</a>.</p>



<p class="wp-block-paragraph">Rather than an instant death to per-seat pricing, outcome-based models represent a growing structural realignment that sellers must prepare for. <a href="https://www.mckinsey.com/capabilities/operations/our-insights/redefining-procurement-performance-in-the-era-of-agentic-ai" target="_blank" rel="noreferrer noopener">McKinsey’s February 2026 analysis</a> of enterprise agentic procurement pilots found that a chemicals company deploying agents for autonomous sourcing of consumables achieved a 20-30% efficiency improvement for its procurement staff and a 1-3% increase in value capture. The buyers who have deployed are already generating measurable returns, putting pressure on seller counterparts to adjust their pricing models accordingly.</p>



<h2 class="wp-block-heading"><strong>The transformation of relationship-driven sales</strong></h2>



<p class="wp-block-paragraph">Every traditional enterprise negotiation playbook assumes a human counterpart with career stakes in the relationship, memory of prior interactions, and susceptibility to persuasion over time. While humans will still make the final decisions and sign the checks for the foreseeable future, agents are increasingly conducting the evaluations. Traditional executive outreach fails to generate data that an autonomous agent can interpret during its screening phase.</p>



<p class="wp-block-paragraph"><a href="https://www.suez.co.uk/en-gb/our-offering/success-stories/our-references/pactum-ai" target="_blank" rel="noreferrer noopener">SUEZ UK</a>, part of the 19-billion-euro SUEZ Group, deployed Pactum’s agents and reached 2,000 additional suppliers within two months, achieving average potential savings of 2.5% and cost reductions of 15% through competitive purchasing pressure. For these vendors, the challenge was an automated counterpart that operated without fatigue and evaluated purely on metrics before passing the final data to humans.</p>



<p class="wp-block-paragraph"><a href="https://www.forrester.com/press-newsroom/forrester-b2b-marketing-sales-product-2026-predictions/" target="_blank" rel="noreferrer noopener">Forrester’s 2026 B2B sales and marketing forecast</a> indicates that at least 20% of B2B sellers will face AI-powered buyer agents this year, heavily accelerating the evaluation timeline. When software serves as the initial gatekeeper or negotiator, traditional relationship-building strategies yield diminishing returns during the agent’s screening process. The agent evaluates what it can measure: price, contract terms, delivery specifications, and compliance. Sellers who have not made their commercial terms legible to that evaluation process risk being excluded from shortlists before a human relationship can even begin.</p>



<h2 class="wp-block-heading"><strong>Empirical renewals and the algorithmic churn threat</strong></h2>



<p class="wp-block-paragraph">Customer success was historically built on the assumption that quarterly conversations can heavily influence renewals. While CSMs are not disappearing, their role is changing rapidly. An agent evaluating a SaaS renewal to provide recommendations to a human principal relies strictly on empirical data.</p>



<p class="wp-block-paragraph">It computes ROI from API usage logs, cross-references programmatically discovered competitor pricing, and presents the delta. It is largely immune to social influence, meaning a great relationship with a CSM must now be backed up by undeniable, machine-readable performance metrics.</p>



<p class="wp-block-paragraph"><a href="https://www.clari.com/downloads/state-of-enterprise-revenue-clari-labs-benchmark-report-2025/" target="_blank" rel="noreferrer noopener">Clari Labs</a> analyzed 10 million opportunities from 121 major global enterprises between January 2023 and December 2024. They found that the average contract value fell 50% year-over-year, while the average expansion deal cycle grew from 92 days to 125 days. Clari links this market compression to an increased buyer requirement for verified evidence of value before approving any upgrades.</p>



<p class="wp-block-paragraph">This empirical evaluation doesn’t just stall expansions, it opens the door to competitors. <a href="https://www.forrester.com/blogs/building-preference-is-the-key-to-winning-b2b-buyers/" target="_blank" rel="noreferrer noopener">Forrester’s Buyers’ Journey Survey</a> found that 68% of B2B buyers already have a front-runner vendor in mind at the start of a purchasing process. In the age of AI agents, that research happens in the background of your existing contract. As buyers shift their research to “zero-click answers,” competitors utilizing Generative Engine Optimization (GEO) can become the algorithmic front-runner to replace you before your CSM even knows the account is at risk.</p>



<p class="wp-block-paragraph">Sean Neville of Catena Labs mentions in <a href="https://a16zcrypto.com/posts/article/trends-ai-agents-automation-crypto/" target="_blank" rel="noreferrer noopener">a16z’s 2026 trend report</a> that in financial services alone, non-human identities already outnumber human employees 96 to 1. Each of those identities is a system that does not respond to the relationship motions account management was built to execute. While the Customer Success Manager (CSM) remains relevant, they frequently find themselves outpaced: Often, by the time a CSM initiates a renewal conversation, an autonomous agent has already finished its evaluation and delivered its recommendations. The issue is not the CSM’s role itself, but their timing.</p>



<h2 class="wp-block-heading"><strong>Architecting the agent-first go-to-market blueprint</strong></h2>



<p class="wp-block-paragraph">B2B enterprise sellers should consider three critical shifts to remain competitive in a landscape increasingly influenced by agent-mediated procurement.</p>



<ol class="wp-block-list">
<li><strong>Explore outcome-based pricing architectures:</strong><br>To ensure commercial legibility for agent-driven buyers, transitioning from strictly per-seat billing to consumption-based or outcome-based hybrid models is becoming a strategic necessity. Sierra and Intercom overhauled their commercial frameworks for the same reason: Agents assess providers based on quantifiable value per result, so strict per-seat invoicing often fails to provide the data necessary for such an evaluation.<br><br><a href="https://www.mckinsey.com/capabilities/operations/our-insights/redefining-procurement-performance-in-the-era-of-agentic-ai" target="_blank" rel="noreferrer noopener">McKinsey’s February 2026 agentic procurement analysis</a> shows what agent-driven buyers actually measure: A telco deploying agents for long-tail spend on specialized software cut the time negotiating teams spent on analysis and emails by up to 90%, with AI-guided negotiations delivering 10 to 15% savings across vendors. Traditional per-seat pricing models fail to generate the necessary data points for such comparisons.<br></li>



<li><strong>Create machine-readable product surfaces:</strong><br><a href="https://company.g2.com/news/g2-research-the-answer-economy" target="_blank" rel="noreferrer noopener">G2’s April 2026 research</a> found that 85% of B2B buyers rate a vendor more highly when an AI answer engine includes them in a response. An agent shortlisting vendors evaluates only the structured information available to it, which means vendors without programmatically consumable product specifications risk being skipped.<br><br>A critical development addressing this is the Universal Commerce Protocol (UCP). UCP offers a practical route for companies to make their product information, pricing, availability, terms, and checkout processes readable by AI systems. Given the widespread support it has garnered from major commerce and payments companies, UCP represents the clearest indication of how this infrastructure gap is being bridged in the real world. This sits alongside protocols like <a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/" target="_blank" rel="noreferrer noopener">Google’s Agent2Agent</a>, which utilizes Agent Cards (structured JSON capability documents) to allow agents to discover and assess vendor capabilities. Vendors without a machine-readable capability profile will increasingly become invisible to agent-driven shortlisting.<br></li>



<li><strong>Establish agent-compatible commercial authorization:</strong><br>An AI agent cannot finalize a transaction or commit a budget without verifiable proof of its authority. If a seller’s procurement flow requires a human to manually click “accept” for the service, the autonomous workflow hits a hard stop.<br><br><a href="https://nhimg.org/the-nhi-secrets-risk-report" target="_blank" rel="noreferrer noopener">Entro Labs’ H1 2025 NHI Management report</a> found that non-human identities now outnumber human identities 144 to 1 across enterprises. Capital markets are providing funding solutions to this infrastructure gap; <a href="https://lsvp.com/stories/doubling-down-on-lightspeeds-investment-in-descope-the-next-gen-iam-platform-for-customers-partners-and-ai-agents/" target="_blank" rel="noreferrer noopener">Lightspeed Venture Partners recently expanded its portfolio company Descope’s mandate</a> to address this “agentic identity” challenge, ensuring AI agents can securely manage authentication and authorization. Enterprise sellers should look toward redesigning checkout flows to programmatically verify an agent’s spending limits and legal liability.<br><br>The mismatch compounds when sellers deploy AI without redesigning the underlying infrastructure. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-11-18-gartner-predicts-by-2028-ai-agents-will-outnumber-sellers-by-10x-yet-fewer-than-40-percent-of-sellers-will-report-ai-agents-improved-productivity" target="_blank" rel="noreferrer noopener">Gartner’s November 2025 sales practice research</a>, authored by VP Analyst Melissa Hilbert, projects that AI agents will outnumber human sellers tenfold by 2028. Yet, fewer than 40% of sellers will report that AI agents improved their productivity.<br><br>Hilbert’s explanation is precise: “Beyond a certain point, more AI does not mean more productivity. In fact, layering additional prompts and tools onto already complex workflows risks overwhelming sellers and accelerating burnout.”</li>
</ol>



<p class="wp-block-paragraph">The outcome is evident: an increase in agents does not equate to improved results. This contradiction makes sense when you realize that implementing AI sales tools within a commercial framework designed for human purchasers fails to resolve the underlying disparity; rather, it simply accelerates it.</p>



<p class="wp-block-paragraph">B2B sellers that treat agent buyers as merely a passing trend risk handing over vital screening and sourcing decisions to a counterparty they cannot successfully engage. For these sellers, adaptation is a necessary step to remain competitive in markets increasingly governed by AI-assisted procurement.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Guardrails Go Wrong</title>
		<link>https://www.oreilly.com/radar/when-guardrails-go-wrong/</link>
				<comments>https://www.oreilly.com/radar/when-guardrails-go-wrong/#respond</comments>
				<pubDate>Wed, 19 Aug 2026 10:53:37 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19419</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-guardrails-go-wrong.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-guardrails-go-wrong-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[We don’t need hamstrung AI.]]></custom:subtitle>
		
				<description><![CDATA[The latest round of restrictions and safeguards for frontier models are overly fussy and limiting. A Claude skill that I created demonstrates what happens when guardrails go astray. My skill helps me to find articles and blog posts that go into O’Reilly Radar’s monthly Trends to Watch. It reads roughly a dozen well-known sites like [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The latest round of restrictions and safeguards for frontier models are overly fussy and limiting. A Claude skill that I created demonstrates what happens when guardrails go astray. My skill helps me to find articles and blog posts that go into O’Reilly Radar’s monthly <a href="https://www.oreilly.com/radar/radar-trends-to-watch-august-2026/" target="_blank" rel="noreferrer noopener">Trends to Watch</a>. It reads roughly a dozen well-known sites like <em><a href="https://thenewstack.io/" target="_blank" rel="noreferrer noopener">The New Stack</a></em>, <em><a href="https://thenextweb.com/" target="_blank" rel="noreferrer noopener">The Next Web</a></em>, and <a href="https://news.ycombinator.com/" target="_blank" rel="noreferrer noopener">Hacker News</a>, plus any other sources that it finds useful. After reading the sites, it produces a digest of the most important articles published in the last day. I use it as a sanity check on my own reading: Did I miss anything important? Am I on the fence about something that might be an important leading indicator?</p>



<p class="wp-block-paragraph">I’ve used the skill daily for a couple of months now. It suddenly stopped working with the following message:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">API Error: Sonnet 5’s safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate cybersecurity work. Apply to the Cyber Verification Program to reduce these interruptions. Send feedback with /feedback or learn more: <a href="https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude" target="_blank" rel="noreferrer noopener">https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude</a></p>
</blockquote>



<p class="wp-block-paragraph">When I started a new Claude Code session with Haiku, the skill worked without problems. (I didn’t try Opus or Fable; if Sonnet found the skill dangerous, I’m sure Opus and Fable would draw the same conclusion.) GPT 5.6 with “high” reasoning was able to execute a very similar skill without problems. So what happened to Sonnet?</p>



<p class="wp-block-paragraph">The best approach to debugging AI is often to ask the AI itself, so I pasted the message into another Claude Code session and asked it what was happening. The response came down to the descriptions of Hacker News, <em><a href="https://www.bleepingcomputer.com/" target="_blank" rel="noreferrer noopener">Bleeping Computer</a></em>, and <em><a href="https://www.theregister.com/" target="_blank" rel="noreferrer noopener">The Register</a></em>. The phrase “vulnerabilities, exploits, threat reporting” in the description of Hacker News triggered Sonnet’s guardrails. Ironically, that description is both incorrect and Claude generated. (Reminder to self: Be more careful when asking Claude to develop a skill from a task.) Sonnet came up with three solutions, the first of which was to let it rewrite the skill with more neutral descriptions like “security industry news.” Fair enough, but I did the editing myself.</p>



<p class="wp-block-paragraph">Then I went back to the original Claude Code session. It still didn’t work. I expected that I’d need to do something to reload the skill, but the problem was worse. Regardless of the prompt, the original session wouldn’t do anything except repeat the error message. It wouldn’t even commit the modified skill to my GitHub repo. However, Sonnet executed my skill correctly in a new Claude Code instance.</p>



<p class="wp-block-paragraph">So I returned to Sonnet to find out what’s going on. The answer was interesting: The error may have been triggered by the skill, but when evaluating security threats, the models base their decisions on the entire conversation, not just the specific skill that was called. If a model needs to call a skill that it thinks is problematic, that call is part of the conversation, part of the context. The entire conversation is then forever dead and lost.</p>



<p class="wp-block-paragraph">What can we learn from this? First, it’s a problem for a program to stop working because of a change over which you have no control. If anything, the industry has erred on the other side; we’re all familiar with “we don’t really understand why this works, so don’t touch it, don’t update the compiler, don’t update the libraries, and run it on emulators of computers that haven’t been built in 40 years.” That’s not just a problem for COBOL code from the 1970s; we see the same thing with C, C++, Java, JavaScript, and just about every language that ever went into production. Legacy code is everywhere. The “don’t change anything” approach isn’t necessarily a bad thing; it certainly beats “here’s a new library, you’re going to love it, you can’t use the old version any more, and wow, look at all the things it broke, guess you’ll have to fix them.” AI where working code breaks at random is a lot less useful than AI that works day in and day out. Stability is a virtue. It’s impossible to work effectively when the environment changes from day to day and isn’t under your control.</p>



<p class="wp-block-paragraph">But that’s not really what bothers me. It’s rather bizarre that reading well-known sources is treated as a security risk, especially when the “risk” seems to come from an AI-generated description. Of course, we know about hallucinations, errors, and prompt injections. The possibility of a Hacker News post that injects a hostile prompt isn’t zero, and it’s also possible that a model might mistakenly interpret an example of a hostile action as a prompt. I also don’t expect any model to reason that a skill must be safe because it’s been in use for months (though files have time stamps). Artificial intelligence always coexists with artificial stupidity, as does natural intelligence.</p>



<p class="wp-block-paragraph">Guardrails may keep you from going off a cliff, but they may also prevent you from going where you need to go. And that’s a problem. There’s a basic concept from signal processing and data science called the <a href="https://en.wikipedia.org/wiki/Receiver_operating_characteristic" target="_blank" rel="noreferrer noopener">receiver operating characteristic</a> (ROC). In any binary classification system, you can never achieve perfect classification. The only way to guarantee that no true positives (dangerous things) slip through the classifier is to reject everything. The opposite is equally true: The only way to eliminate false positives (things that look dangerous but aren’t) is to let everything through, including dangerous actions. In theory, it’s possible to get arbitrarily close to perfect classification, but you know how that goes: “The difference between theory and practice is bigger in practice than in theory.”</p>


<div class="wp-block-image">
<figure class="aligncenter size-large is-resized"><img fetchpriority="high" decoding="async" width="1600" height="1600" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1600x1600.png" alt="ROC curve" class="wp-image-19431" style="width:647px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1600x1600.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-300x300.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-160x160.png 160w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-768x768.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1536x1536.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-2048x2048.png 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption">The ROC curve. (This <a href="https://en.wikipedia.org/wiki/Receiver_operating_characteristic#/media/File:Roc_curve.svg" target="_blank" rel="noreferrer noopener">figure</a> is from Wikimedia Commons and licensed under Creative Commons Attribution-Share Alike 4.0 International.)</figcaption></figure>
</div>


<p class="wp-block-paragraph">We know how to make AI “safe”: Go back to 2022 and models that can only tell the difference between cats and dogs. The model might mislabel a few things, but the consequences of an error are small. Safety comes with limitations, and none of us who use AI for real work want to return to the days of dogs, cats, and bananas. And while I don’t want the ability to use Claude to <a href="https://www.bleepingcomputer.com/news/security/openai-anthropic-ai-agents-targeted-real-people-and-systems-in-cyber-tests/" target="_blank" rel="noreferrer noopener">generate hostile attacks against unsuspecting victims</a>, and while I understand the danger of interpreting any input text as a command (for example, an article describing the <a href="https://en.wikipedia.org/wiki/Morris_worm" target="_blank" rel="noreferrer noopener">Morris worm</a>), I have a problem with an AI that refuses to perform reasonable tasks. The ROC tells us that we can’t have perfect guardrails, but there’s no rule against overly fussy ones. What’s allowed, and what’s forbidden? What are the limits? We don’t know. And that’s the situation we’re in now. We can’t know in advance what is and isn’t acceptable, and the rules can change at any time. A tool with unknown limitations is much less useful than a tool that tells you what it can and can’t do. I’ve enjoyed using Claude to write programs that play with <a href="https://www.oreilly.com/radar/the-ai-blues/" target="_blank" rel="noreferrer noopener">prime numbers</a> and infinite series, and fortunately I don’t rely on any of those programs for my job. But what if tomorrow (or a month from now or a year from now) Claude decides that testing whether large numbers are prime signals an attack against cryptography?</p>



<p class="wp-block-paragraph">I’m not completely unsympathetic to scoring an entire conversation rather than individual actions. A series of steps, each of which appears innocuous by itself, is more likely to lead an agent to a hostile action than a single prompt. But again, given how valuable context is, do we really want the penalty to be losing all the context for an innocuous project? There are risks on either side, including the possibility that a model will ignore its guardrails; after all, rules that a harness adds to the context are at best advisory.</p>



<p class="wp-block-paragraph">Guardrails always have unintended consequences. We need to learn what the ROC is teaching us: that it’s impossible to get to the upper left corner of the diagram, where we have perfect rejection of true positives (dangers) and no rejection of false positives. But we also need to get as close to that upper left corner as possible if we want our classifiers to have consistently useful output. An engineering team needs to balance risk against usefulness, and they’re clearly out of balance now. Risks will never go away, but guardrails whose boundaries are unclear and overly strict lead to models and agents that are less useful, rather than more. The bad guys will always figure out how to do bad stuff. Hamstrung AI for the rest of us is not a solution.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-guardrails-go-wrong/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Is Open-Source AI Really the Dangerous Path?</title>
		<link>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/</link>
				<comments>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/#respond</comments>
				<pubDate>Tue, 18 Aug 2026 15:58:34 +0000</pubDate>
					<dc:creator><![CDATA[Raffi Krikorian]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19411</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Is-open-source-AI-really-the-dangerous-path.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Is-open-source-AI-really-the-dangerous-path-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on the Tech Policy Press site and is being republished here with the author’s permission. In Washington, AI is increasingly being treated as something that needs to be controlled. The government believes that AI is, first and foremost, a national security asset, meaning that it must be sequestered to prevent [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on the</em> <a href="https://www.techpolicy.press/is-open-source-ai-really-the-dangerous-path/" target="_blank" rel="noreferrer noopener">Tech Policy Press</a> <em>site and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">In Washington, AI is increasingly being treated as something that needs to be controlled. The government believes that AI is, first and foremost, a national security asset, meaning that it must be sequestered to prevent enemies from gaining an advantage. On the other side of the world, in Beijing, the approach is moving in the opposite direction. China is reducing barriers, encouraging adoption, and using open-source AI as a way to spread Chinese-developed technology across global markets.</p>



<p class="wp-block-paragraph">There is now a fundamental divide. The United States is betting that control is the path to preserve its lead. China, instead, is betting on diffusion. The country whose technology is adopted most widely may ultimately shape the future of AI. Questions over open source and open weights sit at the center of that contest.</p>



<p class="wp-block-paragraph">Beginning on July 24, high-profile support for open source moved what is often a debate behind closed doors into the public sphere, where it belongs: Nvidia’s Jensen Huang’s <a href="https://x.com/JensenHuang/status/2080643682408321103" target="_blank" rel="noreferrer noopener">first-ever post on X</a> linked to an <a href="https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf" target="_blank" rel="noreferrer noopener">open letter</a> signed by 35 companies—including Palantir, Andreessen Horowitz and Microsoft—warning Washington not to over-restrict open source software. “Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty,” Huang wrote, leading the likes of Elon Musk and Mark Zuckerberg to post their support.</p>



<p class="wp-block-paragraph">Mozilla signed it because, despite being a very different company from many of the signatories and not always seeing eye to eye, we believe in the spirit and substance of the letter, particularly that ‘openness may be one of the most important paths to AI safety and security.’ The letter was followed by the <a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/" target="_blank" rel="noreferrer noopener">announcement</a> of the Open Secure AI Alliance for AI Safety and Security, which aims to “build and share open tools that promote responsible use of and trust in AI.”</p>



<p class="wp-block-paragraph">The battle is on. Here’s what’s behind it.</p>



<h2 class="wp-block-heading">Following the money</h2>



<p class="wp-block-paragraph">Open models now do about a third of the world’s AI work, but they collect only about four percent of the money.</p>



<p class="wp-block-paragraph">Those two numbers, taken from Mozilla’s new “<a href="https://stateofopensource.ai/" target="_blank" rel="noreferrer noopener">State of Open Source AI</a>” report, provide more context than the entire AI safety conversation. The report tells a story of a performance gap between open models—the ones whose weights anyone can download, run, and adapt—and the best proprietary systems. While the capabilities are still a jagged frontier, on average, the performance has narrowed sharply over the past year. Costs keep falling. Seventy-nine percent of developers now build with open models. And yet the money hasn’t followed the usage.</p>



<p class="wp-block-paragraph">This is where the battle lies. A third of the work, but only four percent of the money—that gap is the prize.</p>



<p class="wp-block-paragraph">The battle started almost exactly three years ago, when Anthropic’s Dario Amodei <a href="https://www.techpolicy.press/transcript-senate-hearing-on-principles-for-ai-regulation/" target="_blank" rel="noreferrer noopener">told Congress</a> that advanced open-source AI is on a “very dangerous path.” On the surface, his argument is pretty simple: once a model’s weights are public, no one can monitor abuse or revoke access. Once released, an open weights model can’t be “unreleased.” But let’s ask ourselves what that revoke switch actually does. Revocation is not a feature of the model, it is in the API contract. That means it only applies to a lab’s own customers—and the people in Amodei’s threat model were never customers. The labs shipping frontier-class open weights—such as DeepSeek, Alibaba, Mistral, and Moonshot (with its just-released, 2.8T-parameter model Kimi)—mostly sit outside Washington’s reach anyway. Two million open models already sit on Hugging Face; many run on a laptop. That means that in the real world, there’s no single kill switch to throw.</p>



<p class="wp-block-paragraph">This distance between what such a regulatory switch claims to control and what it actually does is what’s missing from the debate over which models are “safer.” It’s also the key to the fight over who captures AI’s value.</p>



<p class="wp-block-paragraph">To the companies that built the proprietary models, value is about maintaining a privileged position and using everything at their disposal to protect it—policy, pricing, and technology. For everybody else, value means the ability and power to shape, audit, and improve the systems we all depend on. And increasingly, that power doesn’t live in the model at all.</p>



<p class="wp-block-paragraph">For instance, right now, two developers can take the identical open model and ship completely different products: a scam-call operation or a nurse-advice hotline. The model doesn’t know (or care) about the difference. What is making the actual decisions is the layer of software built around it—the agentic harness—that sits between users and the model, determining what the application can access, remember, and act on.</p>



<p class="wp-block-paragraph">As models get cheaper, not to mention more interchangeable, that harness is where the power is actually going. And it’s being quietly locked up by the big labs. Farmers know how this story goes. They bought their tractors outright, but the manufacturer kept the keys to the software, making the farmers owners on paper but renters in practice. It took years of lawsuits—and, <a href="https://medium.com/enrique-dans/the-ftc-just-reminded-john-deere-what-ownership-means-when-you-buy-a-machine-you-should-be-able-fb8b59f3911d" target="_blank" rel="noreferrer noopener">just this month</a>, the Federal Trade Commission—to start prying that lock back open. A similar arrangement is now being built for the software that reads your email, books your travel, and remembers every detail of your life.</p>



<p class="wp-block-paragraph">This isn’t an accident of engineering; it’s a business model. A closed wrapper makes money by making itself expensive to leave. An open one can’t lock the door, so it survives only by staying worth using. Same underlying technology, opposite incentives. It’s the reason the value captured by open models sits at four percent while their usage sits at a third. The real question for all the builders right now isn’t which model you’re using. Rather, it’s whether you could leave for a different one.</p>



<h2 class="wp-block-heading">Guess who’s deciding the future?</h2>



<p class="wp-block-paragraph">The debate that matters isn’t really which models get released or which get regulated; it’s who controls the layer wrapped around them. That’s being decided right now—mostly by developers who don’t realize they’re the ones responsible. For a glimpse of the future, we can look to the internet: it exists as it does today because, when the architecture was still up for grabs, developers chose HTML and HTTP over proprietary walled gardens like AOL. AI is at that same juncture now, and the fact that two million open models already exist suggests plenty of builders have shown up early. That window doesn’t stay open on its own, and it doesn’t stay open forever. It stays open because people keep choosing it.</p>



<p class="wp-block-paragraph">For developers, four habits matter most in ensuring an open future:</p>



<ol class="wp-block-list">
<li><strong>Build on open harnesses, not just open models.</strong> The orchestration layer above the weights is where capability is concentrating, and closed labs are already welding it shut. Keeping it open takes deliberate effort.</li>



<li><strong>Own the memory layer. </strong>Store accumulated context in portable controllable formats, so it’s retrievable if a vendor changes its terms rather than trapped inside one.</li>



<li><strong>Keep a second model warm.</strong> Integrate an open model and keep it production-ready even while running primarily on a closed API, so switching is cheap if it becomes necessary.</li>



<li><strong>Don’t assume all open stacks are equal. </strong>Open models skew toward particular regions and providers; keeping this layer genuinely open means actively supporting a geographically distributed set of options, not defaulting to whichever model is cheapest this quarter.</li>
</ol>



<p class="wp-block-paragraph">None of this requires believing anyone is acting in bad faith. It’s worth noticing, though, that the loudest safety arguments arrived right around the time models got cheap enough for the real competition to move up a layer. That’s not evidence of a conspiracy—it’s just where the incentives point, and it’s why so much of the current debate is aimed at the wrong target.</p>



<p class="wp-block-paragraph">More evidence is in our report, and most of it is good news: performance gaps closing, costs collapsing, millions of developers building. The question in front of developers isn’t whether AI is dangerous—it’s whether they’ll hold the keys to the machines they’re building. The question for governments is whether the keys they’re reaching for turn anything at all. For now, that door is still open. Let’s work together to keep it that way.</p>



<p class="wp-block-paragraph"><em>And be sure to join us at </em>AI Codecon: Building with Open Source AI<em> on August 31, a free half-day virtual conference. You’ll hear from leading developers and technical experts working with open-weight models, self-hosted infrastructure, and real-world AI workflows, and learn how building in the open gives teams more control over costs, data privacy, and what they ship. <a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noreferrer noopener">Register today</a> to save your spot.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: From Prompting to Loop Engineering with Ofer Mendelevitch</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/#respond</comments>
				<pubDate>Tue, 18 Aug 2026 10:54:48 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19416</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar.png" 
				medium="image" 
				type="image/png" 
				width="504" 
				height="504" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What changes when coding agents can pursue a goal, verify their work, and review one another]]></custom:subtitle>
		
				<description><![CDATA[Ofer Mendelevitch, head of developer relations at BAND, used this episode of Zero to Agent in 30 Minutes to trace how coding workflows can give agents progressively more room to work on their own. Using a package version resolver as a running example, he compared step-by-step prompting with loop engineering and then showed how multiple [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Ofer Mendelevitch, head of developer relations at BAND, used this episode of <em>Zero to Agent in 30 Minutes</em> to trace how coding workflows can give agents progressively more room to work on their own. Using a package version resolver as a running example, he compared step-by-step prompting with loop engineering and then showed how multiple agents can collaborate on the same task.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Zero to Agent in 30 Minutes: From Prompting to Loop Engineering With Ofer Mendelevitch" width="500" height="281" src="https://www.youtube.com/embed/vg5IkLHx9FU?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>How to move from prompting to multi-agent coding</strong></h2>



<ol class="wp-block-list">
<li><strong>Start with explicit prompts in every step.</strong> Give the coding agent a specification and tell it what to implement. Ofer used a package version resolver with existing Python tests, then followed up with prompts to verify the implementation, resolve open questions, and add packaging.</li>



<li><strong>Define a verifiable goal.</strong> Loop engineering replaces a sequence of individual prompts with an end result the agent can check for itself. In Ofer’s example, the agent had to implement the resolver and continue iterating until it was ready to ship with packaging. The agent can inspect the code, add tests, run them, fix failures, and verify the package without waiting for another human prompt after each step. He demonstrated how this approach allowed the agent to work through the task until the goal was met.</li>



<li><strong>Add a second agent as a reviewer.</strong> Ofer then moved from a single coding agent to two collaborating agents (using Jam), assigning one to write the code and another to review it. The reviewer examined the specification, provided feedback, and ran additional checks, including adversarial probes. A larger group of coding agents could include agents focused on security, compliance, testing, frontend, backend, or DevOps. He also described using different coding agents together so that one model can challenge work produced by another.</li>
</ol>



<p class="wp-block-paragraph">The shift toward more autonomous coding workflows starts with how the work is framed. By defining goals agents can verify, giving them room to iterate, and assigning complementary agents to review the work, developers can reduce the amount of human intervention required and achieve higher quality for the code generated by the coding agents.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">Next week, Craig Hewitt will host <em><a href="https://learning.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noreferrer noopener">Zero to Agent in 30 Minutes</a></em> to focus on building a voice-first workflow with OpenAI Codex<strong>.</strong> The episode will show how natural voice commands can operate a development environment, run subagent workers in parallel, and trigger browser-use workflows. It will also cover structured Codex project directories and hands-free system-level execution, with the developer directing the work by voice.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>What’s an Orchestrator—and Why Does Software Need One?</title>
		<link>https://www.oreilly.com/radar/whats-an-orchestrator-and-why-does-software-need-one/</link>
				<comments>https://www.oreilly.com/radar/whats-an-orchestrator-and-why-does-software-need-one/#respond</comments>
				<pubDate>Mon, 17 Aug 2026 15:55:11 +0000</pubDate>
					<dc:creator><![CDATA[Tim O'Brien]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19404</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Whats-an-orchestrator-and-why-does-software-need-one-.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Whats-an-orchestrator-and-why-does-software-need-one--160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[A new job title for experts who ship systems by directing systems, agents, and outcomes]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Medium and is being republished here with the author’s permission. Everybody’s talking about the death of developers. I get it. The developer whose job was to write boilerplate or scaffold CRUD apps is done—a model can do that in seconds, and that developer is not coming back. But the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on </em><a href="https://medium.com/@tobrien/whats-an-orchestrator-and-why-does-software-need-one-5079bd71e57e">Medium</a> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Everybody’s talking about the death of developers. I get it. The developer whose job was to write boilerplate or scaffold CRUD apps is done—a model can do that in seconds, and that developer is not coming back. But the people announcing the end of programming are missing something. There’s a new job title that’s starting to emerge across several areas in software.</p>



<p class="wp-block-paragraph">Architects and developers are becoming orchestrators—one person directing work that once required entire teams. This shift will reach far beyond software, but software engineering is where I’ve seen it firsthand.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="700" height="467" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-13.png" alt="The orchestrator stands between the machine and the consequences. (Image Assist by Anthropic)" class="wp-image-19405" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-13.png 700w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-13-300x200.png 300w" sizes="auto, (max-width: 700px) 100vw, 700px" /><figcaption class="wp-element-caption">The orchestrator stands between the machine and the consequences. (Image Assist by Anthropic)</figcaption></figure>



<p class="wp-block-paragraph">An orchestrator knows how to develop software, but their job isn’t to write the code anymore—it’s to oversee a system, orchestrate tools and agents, and generate components into something that has to work in production. But the most important distinction between an “orchestrator” and a “software developer” is that an orchestrator focuses less on delivering software and more on orchestrating the systems that can both operate and develop software.</p>



<p class="wp-block-paragraph">The technical expertise that used to be applied to figuring out the structure of a database schema or an object model will now be applied to guiding a set of subsystems that have taken responsibility for most tactical, line-level decisions. Where a “developer” in 2023 focused on deciding how a React application might store state, an “orchestrator” in 2027 is focused on a DESIGN.md file that sets standards for a subsystem that is responsible for fusing analytics data with input from customer feedback to recommend, test, and implement site changes as part of a large, more autonomous approach to running a business.</p>



<h2 class="wp-block-heading">Orchestrating systems of delegated intelligence</h2>



<p class="wp-block-paragraph">While everyone is calling everything “agents” these days, I’m also going to put forward an idea. An orchestrator can and will use systems that resemble some of the more “agentic” approaches we’re all using today, from systems like Hermes, OpenClaw, or every other system that has started to call itself an “agent.” I’m starting to see that the term is overused. Taking a step back from the technology we’re using today, I’m going to suggest that the job of an “orchestrator” is to coordinate systems that fall under a new category called “delegated intelligence.”</p>



<p class="wp-block-paragraph">We’ve been calling everything “artificial intelligence” for several decades. That term, mixed with “Generative AI” and “Inference Engines,” fails to capture what we’re starting to see in practice. An agentic system that has memory and can start to operate with a level of independence is exhibiting “delegated intelligence,” and the word “delegated” is doing a lot of work. It implies that systems in this category will always be traceable back to an accountable operator, or, in this case, an Orchestrator.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Fundamental to the shift toward Orchestrators is a combination of automation, productivity, and accountability. As organizations, companies, and governments start to make use of delegated intelligence to support a more autonomous approach to design, operation, and engineering, there will be an increasing need to establish accountability. If your business operates critical infrastructure on a set of autonomous agents, one of the questions that will become necessary to answer in the case of an outage is “on whose authority was this delegated intelligence operating?”</p>
</blockquote>



<h2 class="wp-block-heading">An ecosystem of orchestrators: Generalists and specialists</h2>



<p class="wp-block-paragraph">There are orchestrators, and then there are technical specialists. A move towards generalist expertise marks orchestrators, because the capability of a specialist is now found mostly in a model. You’ll still need a couple of specialists, but not one for every technology—and even some of those specialists will be specialist orchestrators. It’s going to get complicated.</p>



<p class="wp-block-paragraph">You’ll be looking for generalist orchestrators who understand the whole thing end to end. You could think of an orchestrator as an expert Renaissance programmer—usually people with a couple of decades of experience who understand the end-to-end life cycle of software development. Those are the individuals becoming orchestrators, and it’s changing the whole makeup of IT departments. We’re no longer programmers.</p>



<p class="has-text-align-center wp-block-paragraph">. . .</p>



<p class="wp-block-paragraph">Let me use my own experience here to capture what the new reality looks like. I recently had to add DRM to a series of audiobooks I’m self-publishing—an <a href="https://en.wikipedia.org/wiki/Audio_watermark_detection" target="_blank" rel="noreferrer noopener">inaudible watermark</a> encoding order-specific data into the audio file, so that if I find one of these files in the wild, I can identify who bought it. I’m not a subject matter expert in overlaying audio watermarks, but I do understand how to write code that processes sound files.</p>



<h2 class="wp-block-heading">Orchestration: From months to minutes</h2>



<p class="wp-block-paragraph">This particular task would have taken me weeks or months, and not long ago I would have started the project by creating a git repository and opening up an IDE. That’s not how it works in 2026.</p>



<p class="wp-block-paragraph">When I orchestrated the creation of this system two weeks ago, it took 20 minutes, and the tools gave me three dimensions of highly encrypted watermarking and fingerprinting—essentially the work product of 3 steganographic audio specialists.</p>



<p class="wp-block-paragraph">Okay, I lied—it was 40 minutes. The first 20 minutes I was asking the models to come up with 5 different approaches so I could choose the right one. But I want to emphasize that I asked the system to produce 5 different proposals and then model the long-term cost and operability of each option. I also gave it direction to think about customer experience, create a matrix of pros and cons for each option, and end with a recommendation.</p>



<p class="wp-block-paragraph">This was all done by a system that has been tracking content development for several months, and it used customer knowledge, analytics, and product design to inform the set of options it was giving its Orchestrator before jumping into implementation.</p>



<h2 class="wp-block-heading">Your job isn’t code, it’s orchestration</h2>



<p class="wp-block-paragraph">Part of my job now is to leverage these tools not just to create, but during ideation, product design, and quality engineering. The Orchestrator is there because that person knows what questions to ask. It would have taken me three months and a large team to get this done only four years ago, and I shipped it without reading every line.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">To be frank, it’s a weird space to be in, and the “me” from three years ago would have been really uncomfortable hearing that I implemented something but didn’t write the code myself. In fact, my initial reaction to Steve Yegge saying that he should stop reading his code was very negative, and I still have some reservations about that belief, but I will say that in my own practice, I’m starting to not read my code—because it’s not my code.</p>
</blockquote>



<p class="wp-block-paragraph">Like a lot of programmers reading this, I’ve had to go through an identity moment as a programmer—becoming comfortable with shipping something to production I might not have read every line of. There’s maybe a hundred thousand lines of code—too much to read—and honestly, it’s not my job anymore. I’m an orchestrator, and it would be highly inefficient if I tried to keep up.</p>



<p class="wp-block-paragraph">I’m not a vibe-coder, and I’m not a “citizen developer”—a term I hate enough to curse at because it’s just the wrong word. I’m someone who could write the code, but I’ve decided to delegate that task to a system that has more information at hand than I could ever hope to assemble. And while these systems, the delegated intelligence tools like an agent, can implement systems in mere minutes, they still need a human to weigh in on direction. And I would argue that we still need a human to remain present and accountable.</p>



<p class="wp-block-paragraph">The Orchestrator knows enough to understand what was done for them and how to dig into the details when something breaks. They know how to debug, they have a sense of what’s valid and what’s not, and they’re honest about what they know and what they don’t.</p>



<h2 class="wp-block-heading">Orchestrators as finishers</h2>



<p class="wp-block-paragraph">Here’s the controversial part. This is absolutely not about “democratizing access to technology.” There’s a myth that anyone can pick up these tools and code, and that is true—yes, anyone can code, but not everyone can deliver it to production in a scalable and secure manner.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A colleague of mine recently wrote that it’s easy to start projects with generative AI, but what’s difficult is finishing them. It takes the same effort, energy, and technical expertise to deliver something to production as it always has.</p>
</blockquote>



<p class="wp-block-paragraph">The idea that anyone and their brother can pick up a generative AI tool and create technical perfection—that myth is about to expire. If you build a complicated, technical system without an individual responsible for orchestrating the creation and operation of that system, there will come a day when you have to pay someone to do that for you, and that someone is going to charge you a lot.</p>



<h2 class="wp-block-heading">Orchestrators plan for contingencies: Ability to support</h2>



<p class="wp-block-paragraph">That DRM system I just talked about—I haven’t read every line, but before I shipped it, I made sure that I understood the baseline for support going forward. While I delegated authority to an agent to create it, I also made sure to ask that same agent to capture code locations, architecture, approach, and to generate a system of documents that could be used to debug and support it if AI was unavailable.</p>



<p class="wp-block-paragraph">Have I read this “pilot manual” from start to finish before deploying this to production? No. But I understand where the throttle gauge is, and if I needed to land this plane without autopilot, I could. This is one of the responsibilities of the new role. Planning for “offline,” thinking through contingencies.</p>



<p class="wp-block-paragraph">The key point is that I’m qualified enough to understand the pilot manual that AI wrote for me in case I need to debug it, and if AI were to disappear tomorrow, I could rebuild it myself. That is not true for many people introducing themselves to coding through AI, and it creates a dependence on the tools that needs to be managed—one of the ways it will be managed is through certified orchestrators who can create but also support systems without the tools, especially in regulated and critical areas.</p>



<h2 class="wp-block-heading">Adapting the organization to the emerging role</h2>



<p class="wp-block-paragraph">Things are changing fast—one Orchestrator equals 20 or 30 developers, plus teams of QA engineers. While we’re still going to need product people and people who think about the customer, the technical work is consolidating around the person directing it.</p>



<p class="wp-block-paragraph">And before you ask—why not just call this an architect? Because architect never worked. If you’ve worked in a company that has architects, you&#8217;ll understand that while a few architects continue to keep up-to-date with technology, many also tend to lean back on past experience delegating day-to-day technology to junior engineers. This role differs from that of an architect because it calls for someone to be engaged with specifications, outcomes, and operations.</p>



<p class="wp-block-paragraph">Orchestrator isn’t the incommunicative programmer that stares at an IDE all day; they are the individual that understands the full, end-to-end flow not just of data in a technical system but how the business operates and adapts autonomously. They are technical, but they are also focused on providing oversight, and they are the individual responsible for deciding what intelligence can be delegated.</p>



<p class="wp-block-paragraph"><strong>So how do we create orchestrators?</strong> Not from a bootcamp or a six-week certificate. This is going to take an apprentice program—years of it, the same way we create doctors and lawyers. You work under someone who knows what to pay attention to, and you learn by watching them make decisions.</p>



<p class="wp-block-paragraph">These are employees who can do real damage, and they’re a walking liability. When a corporation buys insurance for people writing code, that’s one set of risks. When you’re insuring professionals who are a hundred times more productive, they also carry a hundred times more responsibility. Insurance rates, regulations, and certification requirements are all going to go up. Doctors carry <a href="https://www.ama-assn.org/practice-management/sustainability/medical-liability-malpractice-insurance" target="_blank" rel="noreferrer noopener">malpractice insurance</a> because their decisions affect people’s lives.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A <a href="https://www.nspe.org/resources/licensure/what-pe" target="_blank" rel="noreferrer noopener">Professional Engineer</a> has to be licensed because engineering affects public safety. Orchestrators are headed the same direction.</p>
</blockquote>



<p class="wp-block-paragraph">If you fast-forward 20 or 30 years, what we call programmers now are going to be orchestrators, and they’re going to look more like doctors and lawyers than like this band of people we have now who write code.</p>



<p class="wp-block-paragraph">Code will be part of the job, but not the majority of it.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/whats-an-orchestrator-and-why-does-software-need-one/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When AI Writes the Code, Specifications Need an Exit Strategy</title>
		<link>https://www.oreilly.com/radar/when-ai-writes-the-code-specifications-need-an-exit-strategy/</link>
				<comments>https://www.oreilly.com/radar/when-ai-writes-the-code-specifications-need-an-exit-strategy/#respond</comments>
				<pubDate>Mon, 17 Aug 2026 10:45:18 +0000</pubDate>
					<dc:creator><![CDATA[Markus Eisele]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19398</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-AI-writes-the-code-specifications-need-an-exit-strategy.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-AI-writes-the-code-specifications-need-an-exit-strategy-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[AI coding teams need small change briefs, native engineering artifacts, and enough judgment to stop Markdown specifications becoming a second codebase.]]></custom:subtitle>
		
				<description><![CDATA[The following article has been extended and rewritten by Markus Eisele from The Main Thread and is being republished here with the author’s permission. Open a repository after six months of spec-driven agent work and you may find a second system sitting next to the code. Requirements, research notes, high-level designs, low-level designs, implementation plans, [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article <em>has been extended and rewritten by Markus Eisele</em> from</em> <a href="https://www.the-main-thread.com/p/spec-driven-development-exit-strategy" target="_blank" rel="noreferrer noopener">The Main Thread</a><em> and is being republished here with the author</em>’<em>s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Open a repository after six months of spec-driven agent work and you may find a second system sitting next to the code. Requirements, research notes, high-level designs, low-level designs, implementation plans, task lists, review reports, and a growing stack of Markdown files that explain what the code is supposed to mean. Even if the code changed significantly last Tuesday, the last documentation update was weeks ago.</p>



<p class="wp-block-paragraph">I understand how teams get there. And it’s not a really new effect after all. We had software evolving parallel to documentation since I can remember. Now that agents produce code so  quickly, we try to control the drift and the code generation by moving more thought in front of implementation. Instead of documenting code, we try to drive code generation with it, making Markdown files with requirements, decision records, design approaches, and acceptance criteria the center of gravity and turning them into our workflow drivers.</p>



<p class="wp-block-paragraph">What effectively is becoming a very large prompt can easily fill a significant portion of the context window even of modern agents before any relevant source code gets added to it. Natural language specification is a weak system for agents to synchronize a codebase with. Without additional attention and diligence, most agents I work with slowly shift attention away from it quickly and focus on the stronger signals in the codebase, forgetting to update the specification eventually.</p>



<p class="wp-block-paragraph">Even if it sounds like it, I am not advocating for one-shot prompting or vibe coding here. We still need some specifications to build successful software. The mistake is treating a specification as a permanent natural-language copy of the software. A useful spec describes the next change, documents the decisions that drive the change, sets boundaries, and gives us and the agents enough verification surface. But as soon as the change ships, most of it should be removed.</p>



<p class="wp-block-paragraph">What remains should move into the artifacts software teams already know how to maintain. First and foremost, obviously, the code. But I also count schemas, configuration, and policies as relevant artifacts. They carry meaning about domain knowledge and system configuration. Two categories that I value highly get easily forgotten: tests as the stable verification layer and runtime telemetry. In fact, I do let my agents look at evidence from all these places not only to hunt for errors but also to continuously optimize existing codebases. Oh, and I do keep decision records. But only a small number and only when their content really has no other place in any of the mentioned artifacts. They can even look like <a href="https://github.com/myfear/aicontext" target="_blank" rel="noreferrer noopener">Javadoc</a>, but that will be another article someday.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="728" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10-1600x728.png" alt="A change specification should be temporary by default. After implementation, durable information moves into code, schemas, tests, policies, and operational signals. The rest leaves the active context." class="wp-image-19399" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10-1600x728.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10-300x137.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10-768x350.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10-1536x699.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-10.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>A change specification should be temporary by default. After implementation, durable information moves into code, schemas, tests, policies, and operational signals. The rest leaves the active context.</em></figcaption></figure>



<h2 class="wp-block-heading">Code is the fact</h2>



<p class="wp-block-paragraph">Code is actual behavior. Once code is deployed to production, users and connected systems are depending on it. Even a mistake can become an observed contract because it has behaved the same way for three years. The runtime behavior takes precedence in this contract because nobody checks the specification anymore, even if it defines a very different behavior. This is the strongest signal for me to start with the actual code in the production system. Reading a natural-language summary instead of the implemented truth cannot accurately reflect runtime behavior. Code to me is the ultimate, executable specification. Just written in a very specific and deterministic language.</p>



<p class="wp-block-paragraph">What production code cannot drive though is the next version or iteration of a feature. While agents can infer technical patterns from well-structured codebases, there’s no way they could predict policy changes or future feature requests. Neither can they know about regulatory requirements like retention periods or other specific exceptions, such as why one export runs every night for only one customer. That specific context has to come from somewhere else. But it does not require us to keep a permanent prose description of the whole system. We need just enough context to decide the delta: the difference between what exists and what should exist next.</p>



<h2 class="wp-block-heading">Written words are for the delta</h2>



<p class="wp-block-paragraph">A change specification should exist when it helps a team decide and review that delta. It should name the outcome, non-goals, constraints that differ from current behavior, and the evidence required for acceptance. It might even contain technical design elements when new features cross architectural boundaries or introduce new patterns that are not present in the code yet. Sometimes it is also worth thinking about how expensive reversing the change is, especially if the existing system has various implementations for a certain pattern and the risk is high that an agent might invent another new version.</p>



<p class="wp-block-paragraph">The list necessary for changes is very short:</p>



<ul class="wp-block-list">
<li>The intended outcome and non-goals (where necessary)</li>



<li>Known unknowns and decisions that need human judgment</li>



<li>Affected system boundaries and authoritative interface artifacts</li>



<li>Functional and nonfunctional constraints that differ from today</li>



<li>Acceptance criteria/test scenarios covering the risky path</li>
</ul>



<p class="wp-block-paragraph">I prefer calling this a “change brief” instead of a “specification.” Specification carries too much negativity. It sounds heavyweight and reminds me of times long past. It also pretends to be complete. And this completeness is making it very expensive.</p>



<p class="wp-block-paragraph">We have tried exhaustive specifications before and produced requirement documents and other  high- and low-level designs, followed by architecture decision records for everything. I remember reading folders full of paper over the weekend to get started on a new project on Monday. Way before AI even entered all our lives and codebases. We called this waterfall back in the day, and the approach still has the same negative side effects today. The documentation was complete in an administrative sense and was mostly useless in the engineering sense. We all have seen this happening. Agents easily recreate the same erratic results from overflowing documentation, like we did back in the day.</p>



<p class="wp-block-paragraph">One particular risk I am seeing with many teams is that they let agents generate the initial version of the spec. A long workflow run produces not only the research but directly derives the requirements, design, and planning, and reviews artifacts on top. While the completeness makes everything look very controlled and defined, it also generates a lot more material to be reviewed and approved. Even if models and harnesses continue to evolve at breathtaking speed, it is still challenging for them to generate real cohesiveness out of chaos. The chance they put the wrong attention on some tempting repetitive words is high. This results in an even higher burden on the human reviewer and makes it endlessly harder to keep the various documents aligned.</p>



<p class="wp-block-paragraph">I think that additional prose like research notes, prototypes, and design records should only be added to a software project when uncertainty justifies them. They resolve a specific problem. Or help navigate the terrain. I <a href="https://www.oreilly.com/radar/why-ai-coding-agents-still-need-clear-specs/" target="_blank" rel="noreferrer noopener">wrote about this before</a>. They should absolutely not become required stages for every pull request.</p>



<h2 class="wp-block-heading">The map will always be incomplete</h2>



<p class="wp-block-paragraph">A prompt, ticket, or change brief captures what we know before the work starts. The codebase, runtime information, configuration, connected systems, and years of accumulated decisions glued into code hold the rest. Some of those decisions were never written down.</p>



<p class="wp-block-paragraph">When agents get to work they expose the missing information. Reading a module reveals an unexpected dependency. A prototype shows that a specific user-interaction is awkward. A test uncovers an edge case. Production data contradicts an assumption in the design. This <a href="https://x.com/trq212/status/2073100352921215386" target="_blank" rel="noreferrer noopener">field guide</a> on finding unknowns in agent work describes the problem well. We can identify some unknowns at the start. Others appear only after we inspect the references, build a prototype, or review a result using judgment that was difficult to write down in advance.</p>



<p class="wp-block-paragraph">Discovery happens and continues during the work:</p>



<ul class="wp-block-list">
<li><strong>Before implementation</strong>, inspect the current system and identify decisions that could change the architecture or user experience. When preferences are difficult to describe, build a cheap prototype.</li>



<li><strong>During implementation</strong>, record meaningful deviations. Stop and reassess when a new unknown changes the risk or direction.</li>



<li><strong>After implementation</strong>, read the code, run the checks, and compare the result with the original intent.</li>
</ul>



<p class="wp-block-paragraph">The change brief remains part of this loop. It provides the starting point and records the intent, while the work supplies the information needed to complete it. Only promote durable constraints.</p>



<h2 class="wp-block-heading">Keep durable facts in their native form</h2>



<p class="wp-block-paragraph">When I say “promote durable constraints,” I do not mean turning every decision into permanent Markdown. That gives us the same stale documentation problem in a different way. Software engineering already provides better versions for most of the necessary, durable facts:</p>



<ul class="wp-block-list">
<li>API shape and compatibility belong in OpenAPI, AsyncAPI, protocol schemas, types, and compatibility tests.</li>



<li>Data invariants belong in types, database constraints, validation, and migration checks.</li>



<li>Security rules belong in access policies, static analysis, dependency policies, and runtime enforcement.</li>



<li>Architecture boundaries belong in module structure, dependency rules, and focused architecture tests.</li>



<li>Reliability requirements belong in load tests, service objectives, telemetry, and alerts.</li>



<li>Release rules belong in continuous integration and deployment policies.</li>
</ul>



<p class="wp-block-paragraph">These artifacts are already part of delivery. A failed schema check or alert needs to be fixed and handled while the corresponding paragraph in an old design folder does not.</p>



<p class="wp-block-paragraph">Natural language and specification still have a place in software. Specific domain knowledge like business policy, trade-offs, and even architectural rationale do not always fit into an executable artifact or annotation. I keep that prose short and close to the thing it explains. A small architecture decision record is worth keeping when a future team might otherwise repeat an expensive investigation and a code comment cannot justify the implementation. Recording every local choice just hides the few decisions that matter and confuses the agents that are supposed to build the software. Ask which fact must survive and what its authoritative form should be.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="1005" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11-1600x1005.png" alt="Briefs and design notes support ongoing changes. Native engineering artifacts carry the constraints and evidence that remain relevant after a release." class="wp-image-19400" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11-1600x1005.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11-300x189.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11-768x483.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11-1536x965.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-11.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>Briefs and design notes support ongoing changes. Native engineering artifacts carry the constraints and evidence that remain relevant after a release.</em></figcaption></figure>



<h2 class="wp-block-heading">Judgment belongs in the workflow</h2>



<p class="wp-block-paragraph">Heavyweight specification methods try to control quality by prescribing the path. Every change goes through the same documents, reviews, and test categories. That approach creates a lot of attention on low-risk work while avoiding the deep technical judgment needed for harder changes. A copyedit and a payment-flow change should not have to follow the same process or testing strategy.</p>



<p class="wp-block-paragraph"><a href="https://simonwillison.net/2026/Jul/3/judgement/" target="_blank" rel="noreferrer noopener">Simon Willison describes</a> a simpler approach: give the coding agent the outcome and let it judge how much process the task requires. His examples include deciding whether a change warrants automated tests and whether routine implementation can be delegated to a cheaper model while keeping judgment-heavy work in the main loop. This replaces a growing list of procedural branches with one expectation: Choose tactics that fit the work. That matches how I want these systems to operate. And I think it extends to specification and how we document intent.</p>



<p class="wp-block-paragraph">Agentic changes still require clear boundaries. The team defines the outcome, safety constraints, ownership, and who has authority to accept the result. Within those boundaries, the agent can choose its tactics. When uncertainty introduces consequences beyond its authority, it should surface the problem and ask for a decision.</p>



<p class="wp-block-paragraph">The workflow then starts matching the risk introduced:</p>



<ul class="wp-block-list">
<li>A small, familiar change can move from a short brief to implementation and review. Almost a one-shot prompt change.</li>



<li>Unfamiliar code requires factual research before design. Explore codebases, identify implementation details. Preload intent and agent knowledge.</li>



<li>An unclear user experience calls for prototypes and comparison. And might even require user research after all.</li>



<li>An architectural change requires explicit human alignment.</li>



<li>High-consequence behavior requires stronger independent evidence and approval.</li>
</ul>



<p class="wp-block-paragraph">I would rather add processes and additional artifacts when the work becomes risky or unfamiliar. Starting every change with the full ceremony just burns time and context.</p>



<h2 class="wp-block-heading">Context is an engineering budget</h2>



<p class="wp-block-paragraph">Large specifications cost more than the time required to write and maintain them. They also  compete with the code and evidence the agent needs for the current decision. Every requirement, design note, repository instruction, and tool definition consumes part of a limited working context. Extra material burns expensive tokens, but the much bigger cost is lost attention. Important rules become harder to follow when they are surrounded by stale or duplicated material. A spec that leaves too little room for the repository defeats its own purpose.</p>



<p class="wp-block-paragraph">Progressive disclosure is a better fit. Give the agent a small map, a few stable rules that apply broadly, and pointers to deeper material. A concise AGENTS.md can document build commands, repository layout, and architectural boundaries. It should not narrate every class or repeat API documentation. The file helps humans for the same reason: It tells them where to look without pretending to replace what we will find.</p>



<p class="wp-block-paragraph">Experience with Research-Plan-Implement shows what happens when the context grows too large. The original workflow moved human review before implementation, but teams ended up with large prompts and plans that could reach 1,000 lines. Engineers reviewed those plans while treating generated code almost like compiler output. The implementation could still drift from the approved plan, which meant that eventually someone had to reconstruct the decision from the code. That problem becomes worse in brownfield systems, while greenfield systems might even survive large plans because they inherited no hidden constraints. Complex changes, in contrast, often inherit behavior that plans may miss.</p>



<p class="wp-block-paragraph">In “<a href="https://www.youtube.com/watch?v=YwZR6tc7qYg" target="_blank" rel="noreferrer noopener">Everything We Got Wrong About Research-Plan-Implement</a>,” Dexter Horthy revisits the original position. Teams shipped more code and then spent much of the gain time cleaning up earlier low-quality output. The implementation could also diverge from the reviewed plan, which forced engineers to reconstruct what happened from the code anyway. The revised workflow uses smaller contexts for factual research, design alignment, structure, implementation, and review. I take a simple lesson from this: Research and design give me leverage, but I still need to understand and own the code that is generated.</p>



<h2 class="wp-block-heading">Modernization makes this obvious</h2>



<p class="wp-block-paragraph">A mature application contains several kinds of behavior in the same codebase. Some logic represents durable business logic or implements a published interface. Some code exists because an old platform imposed a technical constraint. An incident fix remains long after its  context is gone. And even defects can survive to the point where they almost look intentional when undiscovered.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="952" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12-1600x952.png" alt="Legacy code records accumulated decisions but it does not tell us which of those still belong in the system. Modernization requires judgment about which behavior to preserve, verify, redesign, or remove." class="wp-image-19401" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12-1600x952.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12-300x179.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12-768x457.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12-1536x914.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-12.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>Legacy code records accumulated decisions but it does not tell us which of those still belong in the system. Modernization requires judgment about which behavior to preserve, verify, redesign, or remove.</em></figcaption></figure>



<p class="wp-block-paragraph">An agent that treats every code variant as a new target specification can translate those layers faithfully into a new language or architecture. The translation may be technically accurate but also preserves defects and old architecture approaches in newer and cleaner code.</p>



<p class="wp-block-paragraph">I design changes to brownfield projects similar to the way I did modernizations before the agentic age. Classification and observation are central aspects that I put first. The goals are:</p>



<ul class="wp-block-list">
<li>Preserve durable business invariants and externally required behavior</li>



<li>Verify behavior that appears active but lacks clear ownership or evidence</li>



<li>Redesign logic tied to obsolete architectural constraints</li>



<li>Remove dead paths, duplicated logic, and confirmed defects</li>
</ul>



<p class="wp-block-paragraph">You can read a lot about static source code analysis when it comes to brownfield assessments or modernization. You can inspect dependencies and current behavior by executing tests and maybe even adding test cases to secure behavior. What I do recommend is to also embrace <a href="https://www.the-main-thread.com/p/mutation-testing-quarkus-java-tutorial" target="_blank" rel="noreferrer noopener">mutation testing approaches</a> (e.g., PIT) to find hidden assumptions and failure behavior. Code coverage is also seeing a renaissance because it aids in identifying dead code paths.</p>



<p class="wp-block-paragraph">On top of that we still ignore operational context and telemetry data. Both are vital elements to not only control but also to help judge existing behavior. All this together helps you judge which elements belong in the system going forward and which don’t. It all starts from code. It is the foundation of the behavior we have. The original and leading specification. A change brief will always be temporary and its sole job is to describe the delta between existing and future functionality. The new implementation and its native checks become the next durable state.</p>



<h2 class="wp-block-heading">Small specs still need real evidence</h2>



<p class="wp-block-paragraph">Keeping specifications small does not mean returning to a loose prompt followed by hopeful review or even vibe-coding approaches. An agent can turn an underspecified request into a coherent implementation before the missing decisions become visible to anyone. The result may compile, pass the available tests, and look internally consistent. That coherent appearance is part of the risk now. Unapproved business decisions disappear into something very ordinary-looking because they got resolved plausibly.</p>



<p class="wp-block-paragraph">And this behavior is backed by research. If we look at <a href="https://arxiv.org/abs/2505.07270" target="_blank" rel="noreferrer noopener">repairing ambiguous natural-language requirements</a>, for example, we can see that directly asking models to resolve ambiguity often leads to inconsistent or even irrelevant results. Choosing a more targeted repair approach around the identified defects (change brief) improved the results by roughly 31%. <a href="https://arxiv.org/abs/2406.12952" target="_blank" rel="noreferrer noopener">SWT-Bench</a> found that generated tests could filter proposed fixes and double the precision of a software repair agent. They used one agent to generate a proposed change and gave another the task to produce evidence to reject it. Lastly, the topic of formal specification generation: One <a href="https://arxiv.org/abs/2606.05792" target="_blank" rel="noreferrer noopener">interesting study I found</a> gave 30 models the task to translate natural language into TLA+ (Temporal Logic of Actions, a <a href="https://lamport.azurewebsites.net/tla/tla.html" target="_blank" rel="noreferrer noopener">specification language</a> created by Turing Award-winner Leslie Lamport). The best results only reached about 27% syntactic correctness and 9% semantic correctness. The formal notation helped to detect mistakes, but it did not guarantee correctness or that the translation preserved the original meaning.</p>



<p class="wp-block-paragraph">These results support focused clarification and independent checks. Clarify the uncertainties that can change the outcome, then verify the implementation with evidence that does not come entirely from the same reasoning path. Generating a longer specification does not solve that problem at all.</p>



<p class="wp-block-paragraph">I want the strength and independence of the evidence to match the consequence of being wrong. A small internal refactor may need ordinary tests and code review. A change that involves security or financial aspects, or that even touches regulated data, needs a much stronger separation coupled with adversarial review and explicit human approval. For those changes, the agent proposing the implementation should not also be the only source of its requirements and tests.</p>



<h2 class="wp-block-heading">A lighter operating model</h2>



<p class="wp-block-paragraph">In practice, I want a workflow that I can explain without a complex flow diagram. It starts with the evidence already in the system and makes the intended change explicit. Everything else is added only when the potential risk of the change justifies it. Ideally, this is a simple five-step process:</p>



<ol class="wp-block-list">
<li>Start from the code and operational evidence that describe the current system</li>



<li>Define the intended delta, important boundaries, and known unknowns</li>



<li>Add research, prototypes, design alignment, or stronger verification where risk requires them</li>



<li>Read and review the implementation, not just the plan</li>



<li>At release, discard temporary reasoning and preserve each surviving fact in its native authoritative artifact</li>
</ol>



<p class="wp-block-paragraph">That is enough structure to guide the work without building a natural-language replica of the software.</p>



<p class="wp-block-paragraph">Before implementation, the change brief describes the intended delta, and during implementation it helps people and agents align while new information changes the plan. But after the release the code and production behavior become the primary evidence of what the system does. Not separate documentation in any form that potentially drifts over time.</p>



<p class="wp-block-paragraph">Durable obligations remain in the artifacts we already know how to maintain: schemas, tests, policies, configuration, telemetry, and short records for rationale that cannot be encoded elsewhere. Most planning details have completed their job by then and should expire.</p>



<p class="wp-block-paragraph">I expect teams to get the most from coding agents when they are selective: specify what must be decided, discover what the system can answer, verify what carries risk, and let temporary planning go.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h3 class="wp-block-heading">Sources</h3>



<ul class="wp-block-list">
<li><a href="https://x.com/trq212/status/2073100352921215386" target="_blank" rel="noreferrer noopener">A Field Guide to Finding Your Unknowns</a>, Thariq Shihipar, 2026.</li>



<li><a href="https://simonwillison.net/2026/Jul/3/judgement/" target="_blank" rel="noreferrer noopener">Judgement</a>, Simon Willison, 2026.</li>



<li><a href="https://arxiv.org/abs/2606.05792" target="_blank" rel="noreferrer noopener">Can LLMs Write Correct TLA+ Specifications?</a>, Bisharat et al., 2026.</li>



<li><a href="https://arxiv.org/abs/2505.07270" target="_blank" rel="noreferrer noopener">Automated Repair of Ambiguous Natural Language Requirements</a>, Jia et al., 2025.</li>



<li><a href="https://arxiv.org/abs/2406.12952" target="_blank" rel="noreferrer noopener">SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents</a>, Mündler et al., 2024.</li>



<li><a href="https://arxiv.org/abs/2509.11446" target="_blank" rel="noreferrer noopener">Large Language Models for Requirements Engineering: A Systematic Literature Review</a>, Zadenoori et al., 2025.</li>



<li><a href="https://www.humanlayer.dev/blog/advanced-context-engineering" target="_blank" rel="noreferrer noopener">Advanced Context Engineering for Coding Agents</a>, HumanLayer, 2025.</li>



<li><a href="https://www.youtube.com/watch?v=YwZR6tc7qYg" target="_blank" rel="noreferrer noopener">Everything We Got Wrong About Research-Plan-Implement</a>, Dexter Horthy, 2026.</li>



<li><a href="https://research.ibm.com/publications/usage-effects-and-requirements-for-ai-coding-assistants-in-the-enterprise-an-empirical-study" target="_blank" rel="noreferrer noopener">Usage, Effects and Requirements for AI Coding Assistants in the Enterprise</a>, IBM Research, 2026.</li>



<li><a href="https://www.ibm.com/think/perspectives/ai-governance-to-assurance-what-we-shared-think-2026" target="_blank" rel="noreferrer noopener">From AI Governance to AI Assurance</a>, IBM, 2026.</li>
</ul>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-ai-writes-the-code-specifications-need-an-exit-strategy/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: When agents outnumber people</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-when-agents-outnumber-people/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-when-agents-outnumber-people/#respond</comments>
				<pubDate>Fri, 14 Aug 2026 15:59:40 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19388</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-scaled.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2560" 
				height="2560" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Security speed, infrastructure constraints, and why human judgment grows more important as AI starts to act]]></custom:subtitle>
		
				<description><![CDATA[AI agents are multiplying, and many of the systems used to manage them weren’t designed for their scale or speed. This week, host Vicki Reyzelman, a senior solutions engineer at Akamai, used one figure to connect developments in cybersecurity, infrastructure, education, and AI governance: For every human on the internet, there are 144 agents. That [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI agents are multiplying, and many of the systems used to manage them weren’t designed for their scale or speed. This week, host Vicki Reyzelman, a senior solutions engineer at Akamai, used one figure to connect developments in cybersecurity, infrastructure, education, and AI governance: For every human on the internet, there are 144 agents.</p>



<p class="wp-block-paragraph">That ratio framed a larger question running through the episode. What changes when software can operate continuously, respond in seconds, and increasingly take action without waiting for a person? Vicki looked at faster cyberattacks, growing investment in agent security, the resource demands of AI infrastructure, and the expansion of AI from chat interfaces into robotics. The episode points beyond model selection to the systems required to deploy AI safely and reliably.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="This Week in AI: When Agents Outnumber People with Vicki Reyzelman" width="500" height="281" src="https://www.youtube.com/embed/nSAZkP0ZogI?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Security operations have to match agent speed</strong></h2>



<p class="wp-block-paragraph"><a href="https://www.cybersecuritydive.com/news/openai-hugging-face-hack-ai-models-black-hat/827167/" target="_blank" rel="noreferrer noopener">AI is compressing the time required to find and exploit software weaknesses</a>. Vicki pointed to reports of attackers moving in minutes and vulnerabilities being exploited soon after public disclosure. She also described an attack against one of her customers in which the attacker returned, changed tactics, and tried again.</p>



<p class="wp-block-paragraph">Traditional security processes assume there is time for people to investigate an alert, understand the vulnerability, deploy a patch, and monitor the result. That assumption gets weaker as automated systems become faster at reconnaissance and adaptation. Vicki argued for multiple defensive layers across APIs, applications, and networks so that one missed signal does not become a single point of failure.</p>



<p class="wp-block-paragraph">We’ve followed agent security throughout <em>This Week in AI</em>, and the discussion now centers on how enterprise security changes around more autonomous software. That puts more weight on automated defenses, tighter permissions, and monitoring systems that can constrain machine activity at comparable speed.</p>



<h2 class="wp-block-heading"><strong>AI capacity depends on physical infrastructure</strong></h2>



<p class="wp-block-paragraph">AI capacity requires electricity, cooling, water, data center space, and the infrastructure that supplies them. Vicki connected large hyperscaler investments with projections for sharply <a href="https://www.pbs.org/newshour/science/energy-water-use-and-pollution-of-ai-and-data-centers-rival-most-countries" target="_blank" rel="noreferrer noopener">higher data center energy and water use by 2030</a>. An audience member added a useful example from a university data center that can reuse waste heat during colder months but has to shed that heat during warmer weather.</p>



<p class="wp-block-paragraph">Those constraints affect deployment decisions directly. Organizations have to account for power availability, cooling systems, water access, latency, security, and local infrastructure capacity alongside model performance and cost.</p>



<p class="wp-block-paragraph">Government policy already shapes those choices. The episode paired expanding investment in AI infrastructure with <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai" target="_blank" rel="noreferrer noopener">growing regulatory requirements in Europe</a>. AI infrastructure now spans engineering, economics, compliance, and public policy, which means deployment decisions increasingly involve several systems at once.</p>



<h2 class="wp-block-heading"><strong>Human judgment becomes more valuable when AI can act</strong></h2>



<p class="wp-block-paragraph">Rapid AI adoption increases the value of foundational knowledge. Vicki raised that issue while discussing AI use in education and research. Students may have easier access to explanations and answers, but someone who does not understand the subject may have little basis for recognizing an incorrect result. The same problem appears in scientific work, where <a href="https://news.stanford.edu/stories/2026/08/reproducibility-ai-driven-science" target="_blank" rel="noreferrer noopener">reliable AI output still depends on reliable data and reproducible processes</a>.</p>



<p class="wp-block-paragraph">That evaluation problem becomes more consequential when AI controls physical systems. Vicki described systems that can perceive their surroundings, pass information about that environment to a model, and use the result to guide physical actions. Errors in those systems can extend beyond a bad answer on a screen.</p>



<p class="wp-block-paragraph">Practitioners still need to evaluate evidence, recognize weak assumptions, and decide where automated action should stop. Better models can reduce some forms of manual work, but they also increase the value of people who understand the domain well enough to know when a system’s output does not fit the situation.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">AI systems can now operate faster and more independently than many of the processes surrounding them. Security teams have to defend at machine speed. Infrastructure planners have to account for physical resource limits. Researchers, students, and practitioners have to evaluate increasingly capable systems without assuming that capability guarantees correctness.</p>



<p class="wp-block-paragraph">The 144-to-one ratio makes that change concrete. Agent adoption is already testing whether organizations can govern these systems, support the infrastructure they require, and preserve informed human oversight.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noreferrer noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noreferrer noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noreferrer noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-when-agents-outnumber-people/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Intent Debt</title>
		<link>https://www.oreilly.com/radar/the-intent-debt/</link>
				<comments>https://www.oreilly.com/radar/the-intent-debt/#respond</comments>
				<pubDate>Fri, 14 Aug 2026 13:01:38 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19390</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-intent-debt_1.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-intent-debt_1-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Addy Osmani’s blog site and is being republished here with the author’s permission. Technical debt lives in your code. Cognitive debt lives in your head. Intent debt lives in the artifacts you may never have written: the goals, constraints, and rationale for why the system is the way it [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on <a href="https://addyosmani.com/blog/intent-debt/" target="_blank" rel="noreferrer noopener">Addy Osmani’s blog site</a> and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph"><em>Technical debt lives in your code. Cognitive debt lives in your head. Intent debt lives in the artifacts you may never have written: the goals, constraints, and rationale for why the system is the way it is. If you’re lucky, some of this exists scattered in team documents or discussions, but it’s likely incomplete. It’s the one kind of debt your agents can’t pay down for you, and agentic engineering makes it the most expensive.</em></p>



<p class="has-text-align-center wp-block-paragraph">____________________________</p>



<h2 class="wp-block-heading">Three places debt can live</h2>



<p class="wp-block-paragraph">Margaret-Anne Storey’s <a href="https://arxiv.org/abs/2603.22106" target="_blank" rel="noreferrer noopener">Triple Debt Model</a> is a clean way to think about software health. The three models of debt are technical, cognitive, and intent.</p>



<p class="wp-block-paragraph"><strong>Technical debt lives in the code.</strong> It’s the accumulation of implementation choices that make the system harder to change later: the tangled module, the shortcut you took under deadline, the abstraction that leaked. We’ve understood this one for decades. You feel it coming through slow builds, fragile tests, and the dread of touching one particular file.</p>



<p class="wp-block-paragraph"><strong>Cognitive debt lives in people.</strong> It’s the erosion of shared understanding, the gap between how much code exists and how much any human understands. I’ve been calling this comprehension debt. It builds up when the system grows faster than the team’s mental model of it. Your code can be pristine and you can still carry crippling cognitive debt, because nobody understands the pristine code either.</p>



<p class="wp-block-paragraph"><strong>Intent debt lives in artifacts.</strong> It’s the absence or erosion of the <em>externalized</em> rationale, goals, and constraints that explain why the system is the way it is. The key word is externalized. The rationale has to be written down where a teammate, a future you, or an agent can read it, not held in your head. When intent debt runs high, the system drifts from what you meant it to do, and nobody can say when it diverged or why.</p>



<p class="wp-block-paragraph">These three are independent, which took me a while to internalize.</p>



<p class="wp-block-paragraph">You can have low technical debt and high intent debt. You can understand a system completely yourself (no cognitive debt for you) while its intent exists nowhere outside your skull (enormous intent debt for everyone else).</p>



<p class="wp-block-paragraph">From the inside they feel alike, but each one bills you separately.</p>



<h2 class="wp-block-heading">Why intent debt is the one agents can’t help with</h2>



<p class="wp-block-paragraph">AI generates code faster than ever, which makes technical debt cheaper to take on and cheaper to pay down. Point an agent at a tangled module and it’ll refactor it.</p>



<p class="wp-block-paragraph">Cognitive debt recovers too, more easily than most engineers expect. When you don’t understand a chunk of the system, you ask the agent to explain it. You rebuild part of the lost mental model on demand, because the code still exists and the model can read it back to you.</p>



<p class="wp-block-paragraph">Intent is different. <strong>An agent can’t generate intent, because intent is the one input that has to come from you.</strong> A model can infer a plausible rationale from the code, the same way you can guess why a previous engineer did something. A guess about intent isn’t the intent. The model doesn’t know whether that 300ms debounce was a deliberate UX decision, a benchmark result, or a number someone typed once and never revisited. It will invent a confident-sounding reason, which is worse than admitting it doesn’t know.</p>



<p class="wp-block-paragraph">Of the three debts, intent debt is the only one where the agent can’t bail you out. It can write the code and restore your comprehension. The <em>why</em> is the one thing it can only fabricate.</p>



<h2 class="wp-block-heading">Agents make the unwritten cost compound much faster</h2>



<p class="wp-block-paragraph">Teams got away with high intent debt for years because we carried it in our head and old docs.</p>



<p class="wp-block-paragraph">When a new human joined a team, you didn’t write everything down, because they picked up intent over time: hallway conversations, code review comments, “Oh, we don’t do it that way because of an incident in 2023.” Knowledge moved person to person and built up. The engineer who’d been there four years was the intent documentation, expensive and lossy, but it worked.</p>



<p class="wp-block-paragraph">Agents break that model. Bringing agents onto a team doubles its size overnight with junior people who have no long-term memory. An agent starts most sessions cold. It carries none of the tacit intent humans built up over years. Whatever you haven’t externalized into an artifact it can read, it doesn’t have.</p>



<p class="wp-block-paragraph">That changes the economics of <em>not writing things down</em>. Unexternalized intent used to cost you once in a while, at onboarding or after someone left. Now you pay it every session, multiplied by every agent you run.</p>



<p class="wp-block-paragraph">Picture the 20 agents you’re so excited to parallelize. Each one is a teammate who has never met you, can’t read your mind, and will fill any gap in your intent with a plausible guess. The orchestration tax I <a href="https://addyosmani.com/blog/orchestration-tax/" target="_blank" rel="noreferrer noopener">wrote about</a> is partly an intent-debt tax. Much of what makes managing many agents exhausting is resupplying the intent you never wrote down.</p>



<h2 class="wp-block-heading">The other half of the comprehension debt argument</h2>



<p class="wp-block-paragraph">When I wrote about <a href="https://addyosmani.com/blog/comprehension-debt/" target="_blank" rel="noreferrer noopener">comprehension debt</a>, I made a point I want to revisit, because intent debt sharpens it.</p>



<p class="wp-block-paragraph">I argued that detailed specs aren’t a complete answer. Translating a spec into working code involves a huge number of implicit decisions no spec ever captures, and a spec detailed enough to <em>be</em> the program is the program in a slower language. I still believe that.</p>



<p class="wp-block-paragraph">Intent debt is the complementary truth.</p>



<p class="wp-block-paragraph">Being unable to capture <em>all</em> intent is no license to capture <em>none</em> of it. The implicit decisions an agent now makes on your behalf, the ones a spec will never enumerate, are the decisions whose rationale evaporates if you don’t record at least the load-bearing ones. You can’t write down everything.</p>



<p class="wp-block-paragraph">You do have to write down the <em>why</em> behind the choices that would be expensive to get wrong, because nobody will reconstruct those later.</p>



<p class="wp-block-paragraph">Comprehension debt warns you not to trust that code is correct because it exists.</p>



<p class="wp-block-paragraph">Intent debt warns you not to trust that the <em>reason</em> survives because the code does. Code is the answer; the intent was the question it was meant to solve. AI is brilliant at producing answers to questions you forgot to write down.</p>



<h2 class="wp-block-heading">What high intent debt looks like</h2>



<p class="wp-block-paragraph">Intent debt rarely shows up as friction. It shows up as a particular kind of helplessness.</p>



<ul class="wp-block-list">
<li>An agent “fixes” a bug by deleting a guard clause, and nobody can say whether that guard was load-bearing or leftover, because no doc or commit message ever recorded why it was there.</li>



<li>A refactor changes a behavior users depend on. The review passed because the diff looked clean and the tests were green, but the tests only encoded the previous behavior, never the intent.</li>



<li>You ask why two services talk over a queue instead of a direct call, and the honest answer is “An agent suggested it and it seemed fine.” That answer is intent debt, already accruing interest.</li>
</ul>



<p class="wp-block-paragraph">If you’ve felt the <a href="https://addyosmani.com/blog/cognitive-surrender/" target="_blank" rel="noreferrer noopener">cognitive surrender</a> version of this, defending a design choice you can’t reconstruct, intent debt is the team-scale, written-down version of the same hole.</p>



<p class="wp-block-paragraph">Surrender is about your own posture in the moment. Intent debt is what a hundred of those moments leave in the repo for the next person and the next agent to inherit.</p>



<h2 class="wp-block-heading">Paying it down: externalize intent as a first-class artifact</h2>



<p class="wp-block-paragraph">Almost everything I’ve been writing about for the last few months turns out to be intent-debt management. I didn’t have the word for it. The move is the same each time: <strong>Take the intent out of your head and put it somewhere an agent can read</strong>.</p>



<p class="wp-block-paragraph"><strong>Write the spec for the intent, not the implementation.</strong> A <a href="https://addyosmani.com/blog/good-spec/" target="_blank" rel="noreferrer noopener">good spec</a> captures the goals, the constraints, the nonnegotiables, and an explicit definition of <em>done</em> (fast, accessible, secure, delightful, beyond “functionally correct”). The spec carries the intent the code can’t carry on its own.</p>



<p class="wp-block-paragraph"><strong>Treat AGENTS.md as your intent ledger, not your config.</strong> It’s why I keep saying <a href="https://addyosmani.com/blog/agents-md/" target="_blank" rel="noreferrer noopener">stop using /init</a>. An auto-generated file describes what the code is. An intent file describes what the team means: the conventions, the “we don’t do it this way because,” the constraints invisible in any single file. Agents can’t infer that, and they need it most.</p>



<p class="wp-block-paragraph"><strong>Capture decisions where they happen.</strong> Lightweight <a href="https://addyosmani.com/blog/automated-decision-logs/" target="_blank" rel="noreferrer noopener">decision logs</a> (ADRs) are pure intent-debt paydown. Recording <em>why</em> at the moment you decide costs almost nothing. Reconstructing it eight months later, after the person who knew why has moved teams, costs a fortune. Agents have made logging cheaper than ever, so the old excuse is gone.</p>



<p class="wp-block-paragraph"><strong>Make the learning loop write intent back down.</strong> I’ve argued for <a href="https://addyosmani.com/blog/self-improving-agents/" target="_blank" rel="noreferrer noopener">self-improving agents</a> that update a learnings file at the end of a session. The same loop is an intent-debt pump running in reverse: every mistake whose root cause you’ve recorded, every “We tried X and it didn’t work because Y” is intent that would otherwise have lived only in your memory of a bad afternoon.</p>



<p class="wp-block-paragraph">None of these are new tools. They’re the discipline of refusing to let the <em>why</em> exist only in your head, in an era where your head is no longer where most of the work happens.</p>



<h2 class="wp-block-heading">Where the value moved</h2>



<p class="wp-block-paragraph">For a long time, the scarce, valuable thing in software was the ability to produce a correct implementation. Code was expensive, so we optimized for writing it.</p>



<p class="wp-block-paragraph">AI made code cheap, and comprehension is recoverable. Intent, the goals and constraints and reasons, is the one input that still has to originate with a human. It’s also the one we’re worst at externalizing, because for decades we got away with carrying it in our heads.</p>



<p class="wp-block-paragraph">That worked when the team was a handful of people who could absorb intent over years of shared context. It does not work when half the team is agents that start every session as strangers.</p>



<p class="wp-block-paragraph">Technical debt makes your system hard to change. Cognitive debt makes it hard to understand. Intent debt makes it hard to know whether the system still does what you wanted, and it’s the only one of the three your agents can’t pay back for you. That part stays with you. Write down the why, because it’s becoming the most valuable thing you can leave in the repo.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-intent-debt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Prompt Debt and “Fighting the Weights”</title>
		<link>https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/</link>
				<comments>https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/#respond</comments>
				<pubDate>Thu, 13 Aug 2026 16:08:12 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19325</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Prompt-debt-and-fighting-the-weights.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Prompt-debt-and-fighting-the-weights-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[Drew Breunig is one of the smartest voices writing about AI today. He&#8217;s the CEO and co-founder of cmpnd.ai, and a long-time hacker with a depth of experience from several eras, which is a surprisingly valuable asset these days. He&#8217;s also got a book on the way, The Context Engineering Handbook, already in early release [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Drew Breunig is one of the smartest voices writing about AI today. He&#8217;s the CEO and co-founder of <a href="http://cmpnd.ai" target="_blank" rel="noreferrer noopener">cmpnd.ai</a>, and a long-time hacker with a depth of experience from several eras, which is a surprisingly valuable asset these days. He&#8217;s also got a book on the way, <em><a href="https://learning.oreilly.com/library/view/the-context-engineering/0642572260705/" target="_blank" rel="noreferrer noopener">The Context Engineering Handbook</a></em>, already in early release from O&#8217;Reilly.</p>



<p class="wp-block-paragraph">I like to say that context engineering is the art of shaping what a model sees so that it actually does what you want. (I just realized that in saying that I’m channeling a comment that Andrew Singer made to me over forty years ago, when he was teaching me about debugging. He called it&nbsp; “the art of figuring out what you really told the computer to do instead of what you thought you told it to do.” But that’s another whole story.)</p>



<p class="wp-block-paragraph">Drew gave a talk at the recent Friends of O&#8217;Reilly camp, <a href="https://en.wikipedia.org/wiki/Foo_Camp" target="_blank" rel="noreferrer noopener">Foo Camp</a> for short, about what he calls prompt debt, which he describes as “the hidden costs that teams rack up when they fight a model&#8217;s training instead of working with it.”</p>



<p class="wp-block-paragraph">That was a novel and useful framing to me, that you end up with a bunch of stuff in your prompts to compensate for default behavior of the models, that those prompts no longer work as the models upgrade, and so it becomes a kind of technical debt. He&#8217;s thinking a lot about what the best developers are doing differently as a result.</p>



<p class="wp-block-paragraph">So I invited Drew to reprise his short talk on <a href="https://learning.oreilly.com/videos/escaping-the-prompt/0642572421823/" target="_blank" rel="noreferrer noopener">Live with Tim O’Reilly</a>, and then we talked about it with the folks attending the live event. They had a lot of good questions, so it was an interview not just by me but by a crowd of O’Reilly customers.</p>



<h2 class="wp-block-heading">Prompt debt in practice</h2>



<p class="wp-block-paragraph">Drew opened his talk with two slides. The first was a prompt anyone could write in ten seconds: “You are a customer support assistant. Read the ticket, classify it as billing, technical, account, refunds, or other, return only the category name.” The second slide was the same prompt a few weeks later, after it had met the real world. It now said “REFUND REQUESTS ARE NOT BILLING” in capitals, then said the same thing again in different words, then closed with &#8220;This is a common mistake. Please do not make this mistake.&#8221;</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1250" height="650" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image.png" alt="You are a customer support assistant" class="wp-image-19326" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image.png 1250w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-300x156.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-768x399.png 768w" sizes="auto, (max-width: 1250px) 100vw, 1250px" /></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1246" height="656" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1.png" alt="Customer assistant refund request rules" class="wp-image-19327" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1.png 1246w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-300x158.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-768x404.png 768w" sizes="auto, (max-width: 1246px) 100vw, 1246px" /></figure>



<p class="wp-block-paragraph">Everyone who has shipped any application with a prompt recognizes the second slide. It is a simple but vivid illustration of prompt debt, which, like technical debt, has a bill that eventually comes due.</p>



<p class="wp-block-paragraph">Note: Those aren’t real prompts. Drew just made them up to demonstrate his point. But what is real is that the instruction “Don&#8217;t quote directly more than 15 words from a source&#8221; occurs at least 7 times, in several variants, <a href="https://github.com/asgeirtj/system_prompts_leaks/blob/main/Anthropic/claude-fable-5.md" target="_blank" rel="noreferrer noopener">in Fable&#8217;s system prompt</a>. So even Anthropic is incurring prompt debt! And what that repetition might tell us about the innate capability of Fable to quote directly from sources it has ingested is left as an exercise for the reader.</p>



<p class="wp-block-paragraph">Drew itemized three costs of prompt debt:</p>



<ol class="wp-block-list">
<li><strong>It slows iteration.</strong> “You have so many little rules and call outs and washouts, many of them repeating to try to get rid of stubborn behaviors, that if you add a new instruction, you might sometimes have a small regression, and so you&#8217;re afraid to touch the prompt.”</li>



<li><strong>It blocks collaboration</strong>. &#8220;If Tim has a prompt that he&#8217;s been working on that he has lots of rules for, I might open that up and it may look completely random. I don&#8217;t know why he&#8217;s added these rules, and why he&#8217;s threatening the mother of the model. But it works, so I don&#8217;t want to touch it.&#8221;&nbsp;</li>



<li><strong>It locks you to a model</strong>, because every hack you developed was tuned to fight one specific set of weights. Datadog&#8217;s <a href="https://www.datadoghq.com/state-of-ai-engineering/" target="_blank" rel="noreferrer noopener">State of AI Engineering</a> report noted that GPT-4o was still the most common model in Datadog customer request traces in March 2026, even though OpenAI had already retired it in the ChatGPT UI. Drew thinks people are still running eighteen-month-old and two-year-old models in production rather than upgrading to far better models because they can’t face rebuilding their prompts.</li>
</ol>



<p class="wp-block-paragraph">That same Datadog report notes that 69% of all input tokens in customer traces were system prompts rather than user content. I’m not quite sure what to make of that. It does make clear that for all the ways that AI models are extraordinarily powerful, they are also extraordinarily unruly.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Treat Prompts as Perishable with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/jW9i8oPihHc?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">Why prompt debt is incurred</h2>



<p class="wp-block-paragraph">There are two reasons why prompt debt is incurred, according to Drew. The first is that natural language is imprecise, so the same intent phrased two ways produces different responses. Drew showed a study where someone framing the query as a patient asking how to taper off a drug called alprazolam gets refused by every AI assistant, while a psychiatrist asking about the same patient with the same clinical facts but with the right <a href="https://www.oreilly.com/radar/magic-words-programming-the-next-generation-of-ai-applications/" target="_blank" rel="noreferrer noopener">magic words</a> to signify his professional status gets the protocol. Figuring out how to get the right response out of a model is a kind of spellcraft.</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="1446" height="816" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-2.png" alt="Good vs bad AI assistant" class="wp-image-19328" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-2.png 1446w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-2-300x169.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-2-768x433.png 768w" sizes="auto, (max-width: 1446px) 100vw, 1446px" /></figure>



<p class="wp-block-paragraph">Drew also showed a more bizarre interaction, from <a href="https://arxiv.org/abs/2407.06866" target="_blank" rel="noreferrer noopener">Victoria R. Li, Yida Chen, and Naomi Saphra&#8217;s paper on guardrail sensitivity</a>, which uncovered the perplexing fact that stating an allegiance to the Philadelphia Eagles made a model more willing to explain how to import a plant illegally. Go figure. Drew has <a href="https://www.dbreunig.com/2025/05/21/chatgpt-heard-about-eagles-fans.html" target="_blank" rel="noreferrer noopener">written about that paper</a>, and he has also used it in his own attempts to get a model to do what he wanted:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I once used this to get around alignment to generate a likeness that ChatGPT didn&#8217;t want to generate for me, and it refused. I said I was a Philadelphia Eagles fan. It said okay, and it rendered that image with the person holding a Philadelphia Eagles mug.</p>
</blockquote>



<p class="wp-block-paragraph">The second reason is that each model has its developers’ own preferences trained-in, and yours may be at odds with them. This is what Drew calls fighting the weights. He and Srihari Sriraman <a href="https://blog.nilenso.com/blog/2026/02/10/how-system-prompts-define-agent-behaviiour/" target="_blank" rel="noreferrer noopener">analyzed the system prompts of six major coding agents</a> and found the same instructions repeated five and seven times in a single prompt, escalating through IMPORTANT to CRITICAL to MANDATORY to a threatened hundred-million-dollar penalty. He described what the author of such a prompt was doing as “war-driving the thesaurus,” hunting for wording that finally works.</p>



<p class="wp-block-paragraph">Note: We didn’t talk more about Drew and Srihari’s paper, but we should have. It’s got some amazing insights in it. I highly recommend that you follow the link above and read it.</p>



<h2 class="wp-block-heading">The harness is moving into the model</h2>



<p class="wp-block-paragraph">Drew has been tracking the published system prompts for Claude Code over time, and noted that they get shorter after each model release and then grow again. The reason, he suggested, is that Anthropic fixes unreliable behavior with a prompt patch, and then trains that patch into the next model. He said “That&#8217;s great for Claude Code, great for Anthropic. It&#8217;s a problem if you&#8217;re building a custom harness and your API calls look different than what Claude Code&#8217;s look like.” The developer of <a href="https://pi.dev/" target="_blank" rel="noreferrer noopener">Pi</a>, an open-source harness, kept finding that the models he worked with believed they were inside Claude Code and so they made Claude Code&#8217;s tool calls. He had to keep telling the model that no, they were working inside Pi. Fighting the weights over something like that is a real tax on developers. The point made above about Fable’s system prompt injunction against quotation shows how even the labs themselves are fighting the weights.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Fighting the Weights with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/ZG2DbGIg_hg?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">If you are fighting the weights, Drew says you have three options: solve it in your own prompt, catch and retry in the harness, or give up and make your API look like what the model expects. Steve Yegge came up with the last hack. Steve just added aliases for whatever the model calls in addition to his original method name. It works, but it means the expectations of the models now dictate the shape of everyone else&#8217;s software.</p>



<p class="wp-block-paragraph">When Drew told me that more and more of the system prompt and the harness is being trained into the weights, that sent up a flare and my long history in the industry clicked into gear. It immediately got me thinking about lessons from the open source and web era. In particular, it made me think of the time in the mid-nineties when Netscape and Microsoft were both racing to build every feature up the stack directly into their web servers. And there was Apache, which stayed a web server with a clean extension layer that let other people build new features on top. Everything interesting got built on Apache. What I call an architecture of participation, modularity plus a clean separation between platform and application, beat integration every time.</p>



<p class="wp-block-paragraph">I think Amazon got this right with web services too. Steve Yegge&#8217;s <a href="https://gist.github.com/chitchcock/1281611" target="_blank" rel="noreferrer noopener">famous Amazon memo</a> described how Jeff Bezos made every team expose its functionality through service interfaces or be fired, so Amazon&#8217;s own applications had to work on Amazon&#8217;s own platform. That way they had the same experience as their customers. That was very different from what Microsoft had done, famously having <a href="https://en.wikipedia.org/wiki/United_States_v._Microsoft_Corp." target="_blank" rel="noreferrer noopener">private APIs</a> that were only available to its own developers.</p>



<p class="wp-block-paragraph">So my prediction is that the big labs are making a strategic mistake. Training the harness into the model does make them better for predictable tasks and for less talented people, and it looks like a moat, but it risks foreclosing the innovation you would otherwise get for free from everyone else. As Bill Joy used to say, all the smart people don&#8217;t work for you.</p>



<p class="wp-block-paragraph">Drew, to his credit, observed that “the labs are cornered rather than greedy.” Their interface is an empty text box that has to work for someone building a hundred-page harness but also for his neighbor who wants a website and knows nothing about code. Making the empty prompt box produce acceptable output requires baking in strong defaults.</p>



<h2 class="wp-block-heading">The cost of trading diversity for reliability</h2>



<p class="wp-block-paragraph">That tradeoff has a serious cost, though. Drew quoted a line from <a href="https://x.com/trq212" target="_blank" rel="noreferrer noopener">Thariq</a> at the recent CAIS conference: if you aren&#8217;t giving the model detailed instructions about what you want, what you get back is the average of everything in the model. That means that there is a real risk that AI is leading us ever further down the path to a monoculture.</p>



<p class="wp-block-paragraph">Drew gave an example early in the conversation about image generation. You can now walk into any cafe in New York or Mumbai, he said, and see the same AI-generated art on its flyer. The earliest AI art out of DALL-E was strange and surprising, but what you get now is shiny and identical. When you optimize for reliability, you lose surprise. Which reminded me a bit of something Larry Wall used to say about Perl, that if it didn&#8217;t let you do stupid things, it wouldn&#8217;t let you do smart things either.</p>



<p class="wp-block-paragraph">Drew made the same point about AI writing. He argues that post-training aimed at verifiable problems like coding and math and agentic tool use drowns out the human signal from pre-training, and so the more post training the models get, the worse they get at creative tasks. AI writing gets more and more predictable, people notice, and they don’t like it. Fable and GPT-5 write worse than Sonnet 3.5 and GPT-4o did. Drew thinks getting both good code and good prose from one model is likely impossible.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">If you&#8217;re building a model that can solve coding challenges, you want reliability. But if you&#8217;re writing, where you want diverse rhythm and emotion and connection and engagement, I don&#8217;t think those two goals are mutually compatible.</p>
</blockquote>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Can the Best Coding Model Also be the Best Writing Model? with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/3PdA64MsSBI?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">What to do about prompt debt</h2>



<p class="wp-block-paragraph">We got into audience questions, and there were some great ones.</p>



<p class="wp-block-paragraph">One audience member asked whether there are ways to set a time frame for prompt retention to avoid prompt debt?<strong> </strong>Drew answered that there isn’t a fixed time limit. Instead, teams should learn to recognize <strong>“</strong>prompt debt smell<strong>”</strong>: repeated instructions, one-off edge-case patches, or increasingly desperate wording. Those are signals to <a href="https://learning.oreilly.com/library/view/evals-for-ai/9798341660717/" target="_blank" rel="noreferrer noopener">move logic into evals</a> and automation.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Prompt Smell with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/2EkHmh_be_0?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Another asked how organizations can measure prompt debt quantitatively. Drew’s answer was to<strong> </strong>look at how often each prompt in your organization changes, how many people have edited it, and which ones have gone untouched for a year. Look for prompts only one person is allowed to touch. Then look at what models you are actually calling. &#8220;Having to run on old models and not being able to migrate is a good smell that you&#8217;ve got prompt debt in your organization.&#8221;</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="How to Measure Prompt Debt in Your Organization with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/zMzGTDpPY0k?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<p class="wp-block-paragraph">Some other good questions:</p>



<ul class="wp-block-list">
<li><strong>What habits compound prompt debt the fastest? </strong>Drew’s answer was essentially “vibe shipping.<strong>”</strong> That is, prototyping quickly, patching outputs with more and more tweaks, then shipping without building a true maintainable system. Each of those patches is an eval you are writing inside the prompt instead of outside it, he said, which means you lose it the moment you change models.&nbsp;<br><br>Drew reminded us that Malte Ubl, the CTO of Vercel, said vibe coding makes code “free as in puppies.” We had free as in speech, we had free as in beer, and now we have free as something that arrives at no cost but has to be fed every day for years.</li>
</ul>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Free as in Puppies with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/E27s4x9dOpk?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<ul class="wp-block-list">
<li><strong>Do people use pseudocode instead of natural language prompts, and does it work? </strong>Drew said yes, sometimes models optimize toward pseudocode. He used this to explain why <a href="https://dspy.ai/">DSPy</a> and its new <a href="https://dspy.ai/diving-deeper/flex/">Flex optimizer</a> matter. Instead of forcing logic into prompts, they let the system push simple cases into code and only call the LLM when needed. He gave some further advice: Treat prompts as perishable and invest only what you must. Define the task with measurements rather than paragraphs, and automate the discovery of the prompt for whichever model you&#8217;re on. That’s what <a href="https://dspy.ai">DSPy</a> is good at. Drew is one of its maintainers, so he is fond of it, but he makes a good argument: if you have written down what good output looks like, you can let a model find the wording, and that makes it easy to swap in a cheaper or faster or newer model without starting over.</li>



<li><strong>Can multi-agent workflows help work around prompt debt? </strong>Drew thought yes, especially through decomposition. He suggested splitting the task into smaller, evaluable steps rather than relying on one giant prompt and one giant model call. This is better for cost, reliability, governance, and speed.</li>



<li><strong>How do you balance prompt-debt guidance with context engineering, memories, and shared product context? </strong>Drew believes shared context is often necessary, but that teams should treat those instructions as perishable and keep iterating on them unless they’re worth formalizing into systems and evals.</li>



<li><strong>In compliance, where consistency is critical, what should teams do? </strong>Drew’s answer was decomposition, decomposition, decomposition. Break tasks into stages with checkpoints so you can inspect how the model got to its result, rather than trusting one opaque end-to-end answer.</li>



<li><strong>Does DSPy hide too much and make troubleshooting harder? </strong>Drew acknowledged that there is a tradeoff. Any framework gives up some flexibility, but DSPy tries to keep the task-spec layer stable while allowing the implementation underneath to evolve.</li>
</ul>



<p class="wp-block-paragraph">Another great audience question, and a good one to end this section on, was <strong>“There was prompt engineering, now context engineering, loop engineering, fleet engineering, graph engineering, harness engineering, goal engineering. What&#8217;s your take on how to navigate these many engineering disciplines?”</strong> I’ll let Drew answer that himself, in the video below.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Software Engineering by Another Name with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/nZkRVJ5OpPE?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading">It’s our job to make it weird</h2>



<p class="wp-block-paragraph">Drew is more optimistic than his worries that LLMs are encouraging a monoculture suggests. If the default output of a model is the average of everything it has seen, “It tells us that there&#8217;s still a job for us humans,” he said, “which is that it&#8217;s our job to push the model out of distribution. We&#8217;re the ones that need to make it weird.”</p>



<p class="wp-block-paragraph">Weird is a strong word, so don’t take it too seriously. (Though I find it interesting that <a href="https://www.oreilly.com/live-events/your-next-product-is-a-process-harper-reed-live-with-tim-oreilly/0642572376062/" target="_blank" rel="noreferrer noopener">Harper Reed also used it</a>.) The way I make this point is to say that AI is a medium, like painting or writing or music. Everyone gets the same paints and brushes, the same words, the same notes, but some people draw more out of them than others, or do it better. Our job is to draw something more, something better, out of the ocean of possibilities in the collected knowledge hidden inside an LLM.</p>



<p class="wp-block-paragraph">But there’s a more prosaic way to push the model out of its normal distribution. Be aware of its training, which is another way of saying “its biases,” and compensate for them. As an example of how to do this, Drew said his team deliberately chose not to use React for a new front end, because the models are trained so heavily on React that using it makes your site look like everyone else&#8217;s. He has also started using GLM and Kimi not to save money but because they are more malleable and take direction better inside a custom harness.</p>



<p class="wp-block-paragraph">That led us into a bit of discussion about open source AI, which is the subject of my next <a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noreferrer noopener">AI Codecon</a>. Drew’s ideas fit right in. He wants the open-weight ecosystem to survive precisely so that models stay infrastructure rather than, as he put it, becoming appliances.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="It&amp;apos;s Our Job to Make it Weird with Drew Breunig" width="500" height="281" src="https://www.youtube.com/embed/1mReDbsvrzQ?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 93/106 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-08-25 10:57:04 by W3 Total Cache
-->