<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Fri, 09 Oct 2026 15:55:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>This Week in AI: More Capability, More Responsibility</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-more-capability-more-responsibility/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-more-capability-more-responsibility/#respond</comments>
				<pubDate>Fri, 09 Oct 2026 15:55:10 +0000</pubDate>
					<dc:creator><![CDATA[Christina Stathopoulos and Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19926</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Government oversight, persistent agents, world models, and practical AI in high-stakes settings]]></custom:subtitle>
		
				<description><![CDATA[AI systems are gaining more autonomy while governments, companies, and researchers are still working out how much oversight they need. On the latest episode of This Week in AI, we covered the US debate over AI governance, new frontier models, persistent agents, world models, and practical uses for AI in healthcare and disaster response. AI [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe width="800" height="450" src="https://www.youtube.com/embed/JFjaJky2Gvc?si=SvYPJunS0UtdZd_M" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<div style="height:100px" aria-hidden="true" class="wp-block-spacer"></div>



<p class="wp-block-paragraph">AI systems are gaining more autonomy while governments, companies, and researchers are still working out how much oversight they need. On the latest episode of <em>This Week in AI</em>, we covered the US debate over AI governance, new frontier models, persistent agents, world models, and practical uses for AI in healthcare and disaster response.</p>



<h3 class="wp-block-heading"><strong>AI oversight is moving beyond company promises</strong></h3>



<p class="wp-block-paragraph">The Trump administration announced a <a href="https://www.theguardian.com/us-news/2026/sep/29/trump-ai-deal-tech-ceos-superintelligence" target="_blank" rel="noopener">voluntary agreement with major AI companies</a> that calls for internal safety monitoring, external audits, and independent board reviews. Because the agreement carries no legal enforcement, it raises a familiar question about how far self-regulation can go when companies are developing increasingly powerful systems.</p>



<p class="wp-block-paragraph">Government agencies are also testing what existing law can do. <a href="https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investigation-into-anthropic-openai/" target="_blank" rel="noopener">The Federal Trade Commission launched an investigation</a> into OpenAI, Anthropic, and other AI companies focused on potential consumer risks. Cases like these could help establish whether current consumer protection laws are enough or whether AI will require a more specialized regulatory framework.</p>



<p class="wp-block-paragraph">For technical leaders, regulation can influence how organizations evaluate models, document risks, manage access, and choose vendors. As AI moves deeper into business processes, teams will need governance practices that can withstand outside review.</p>



<h3 class="wp-block-heading"><strong>AI systems are taking on longer, more autonomous work</strong></h3>



<p class="wp-block-paragraph">OpenAI just released <a href="https://openai.com/index/introducing-dots/" target="_blank" rel="noopener">Dots</a>, its “proactive assistant” designed to retain context, work across applications, pursue multiple goals, and act independently. To do that work without waiting for a prompt, Dots needs standing access to the apps and data it works across. That also increases the amount of personal data an agent can reach and raises the cost of mistakes or misuse.</p>



<p class="wp-block-paragraph">Google is also extending how long a model can work on a task. The company says <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" target="_blank" rel="noopener">Gemini 4 Argon</a> can generate up to a million output tokens in a single response, far beyond the typical output limits of current frontier models. The goal is to let a model stay with long multistep work such as extensive coding or financial and legal analysis. Longer-running models and persistent agents aren’t the same thing, but both let AI finish more work without handing control back to a person.</p>



<p class="wp-block-paragraph">As agents act more on their own, they need to anticipate the consequences of their actions. That becomes even more important when AI moves beyond software and begins acting in the physical world. World models aim to teach AI how these environments work, including cause and effect, which is why many researchers see them as building blocks for robotics and physical AI. World Labs (<a href="https://newsroom.amd.com/news/amd-acquire-world-labs/" target="_blank" rel="noopener">recently acquired by AMD</a>) is developing spatial intelligence models for interactive 3D environments, while British startup <a href="https://www.wired.com/story/the-next-evolution-of-ai-is-learning-from-your-dodgy-gaming-skills/" target="_blank" rel="noopener">Worldmodeldata has licensed nearly 1 million hours of video game data</a> paired with player actions. Researchers from NVIDIA, MIT, and Oxford also introduced <a href="https://physis-intelligence.github.io/physis-lang-web/" target="_blank" rel="noopener">Physis-Lang</a>, which uses descriptions of physical causes and effects to help video models learn why events happen. Researchers still don’t know which training approach will work best, so they’re testing several kinds of data and model design.</p>



<h3 class="wp-block-heading"><strong>AI support experts in high-stakes work</strong></h3>



<p class="wp-block-paragraph">Anthropic’s Claude is <a href="https://www.anthropic.com/features/ebola-response" target="_blank" rel="noopener">helping teams responding to an Ebola outbreak</a> in the Democratic Republic of the Congo organize daily reports, compare forecasting models, analyze genomic information, and support vaccine research. <a href="https://www.nature.com/articles/d41586-026-03106-y" target="_blank" rel="noopener">Researchers have also developed a Spanish-language speech model</a> that may eventually help identify accelerated biological aging and early signs of dementia. Mayo Clinic researchers built a <a href="https://www.newswise.com/articles/artificial-intelligence-model-predicts-pancreatic-cancer-risk-3-years-before-diagnosis" target="_blank" rel="noopener">model that identifies patterns associated with elevated pancreatic cancer risk</a> years before diagnosis. Following severe flooding in Nepal, local <a href="https://indianexpress.com/article/explained/explained-ai/disaster-management-india-nepal-using-artificial-intelligence-10896719/" target="_blank" rel="noopener">teams used AI to match reports of missing people with victim lists</a> and map damaged buildings using satellite imagery.</p>



<p class="wp-block-paragraph">Together, these examples show how AI can be used for good by helping people work through complex information faster and, ultimately, save lives, while keeping qualified experts responsible for the final decisions.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">As AI takes on more work and enters higher-stakes settings, organizations will need clearer answers about access and accountability. They’ll also need to decide where human judgment remains necessary as systems become more capable.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-more-capability-more-responsibility/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Intent, Not Identity</title>
		<link>https://www.oreilly.com/radar/intent-not-identity/</link>
				<comments>https://www.oreilly.com/radar/intent-not-identity/#respond</comments>
				<pubDate>Fri, 09 Oct 2026 10:54:17 +0000</pubDate>
					<dc:creator><![CDATA[Vicki Reyzelman]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19922</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Intent-not-identity.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Intent-not-identity-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[The web is building an identity layer for AI agents. It will tell you who’s knocking. It won’t tell you whether to open the door.]]></custom:subtitle>
		
				<description><![CDATA[On a recent episode of This Week in AI, we discussed that the number of nonhumans on the internet is greater than humans, 144 to 1. It’s an estimate, and the order of magnitude is more important than the number itself. What matters is what a ratio anywhere in that range does to a security [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">On a recent episode of <em>This Week in AI</em>, we discussed that the number of nonhumans on the internet is greater than humans, 144 to 1. It’s an estimate, and the order of magnitude is more important than the number itself. What matters is what a ratio anywhere in that range does to a security architecture. For more than two decades, systems were engineered around two foundational assumptions: either verifying a human user exercising discretionary judgment or validating the account executing hardcoded predictable logic. Autonomous AI agents shatter both sides of this legacy identity taxonomy because they possess broad, human-like operational scope across multiple internal and external applications while remaining inherently nondeterministic and probabilistic in their reasoning.</p>



<p class="wp-block-paragraph">Today, because the agent is interacting on behalf of a human through a genuine browser with standard extensions, its technical fingerprint looks identical to a normal human user’s device. Agents call tools, observe results, and revise. A person browses, hands off to an agent, and takes the session back. All the activity appears to be coming from the same user.</p>



<p class="wp-block-paragraph">Signed agents help the identity problem, and they’re arriving. <a href="https://developers.cloudflare.com/bots/reference/bot-verification/web-bot-auth/" target="_blank" rel="noopener">Web Bot Auth</a> is a thin layer over RFC 9421 HTTP Message Signatures where an agent declares its identity and purpose. But they only help with the identity problem: The signature proves who is making the request, not the intent behind it. Prompt injection and other techniques can be used to steer a verified, well-behaved agent into acting against your interest without its operator ever knowing.</p>



<p class="wp-block-paragraph">The organizations getting this right are treating agent policy as a commercial question with a security implementation. The following questions need to be answered:</p>



<ul class="wp-block-list">
<li>Which agents do we allow to use our service?</li>



<li>What endpoints do we need to protect? </li>



<li>How will agent management security policies impact revenue?</li>
</ul>



<p class="wp-block-paragraph">Here are the techniques successful companies are using to secure in the agentic era. A phased strategic roadmap implementing continuous execution-time verification. These are the steps:</p>



<ol class="wp-block-list">
<li><strong>Establish execution-time machine identity and scoped delegation.</strong> Replace static, long-lived API keys with machine single sign-on mechanisms that issue short-lived, context-bound credentials at the time of tool execution. Enforce strict delegation policies so that autonomous agents operate within verified “green zones” without inheriting broad human privileges.</li>



<li><strong>Deploy cryptographic ingress verification.</strong> Upgrade edge infrastructure to inspect Web Bot Auth HTTP Message Signatures, in order to separate cryptographically verified crawlers from anonymous traffic.</li>



<li><strong>Upgrade edge defenses to browser-layer intent classification.</strong> Transition legacy network-layer antibot web application firewall (WAF) to client-side browser-layer detection platforms. </li>



<li><strong>Harden agent execution pipelines against prompt injection.</strong> Instrument robust input sanitization, instruction hierarchy enforcement, and design-level safeguards around all LLM ingestion channels to neutralize indirect prompt injection attacks before external web content reaches execution tools.</li>



<li><strong>Adopt standardization for agentic commerce and admission.</strong> Integrate emerging trust frameworks such as the Framework for Agentic Commerce Trust (FACT), a neutral real-time trust layer for AI-driven transactions in addition to verification protocols like CAPTCHA to validate agentic capability vectors before delegating sensitive interagent tasks.</li>
</ol>



<p class="wp-block-paragraph">The rapid transition from human-directed software to autonomous AI agents breaks traditional enterprise security paradigms. AI agents operate with broad operational scope but nondeterministic, probabilistic reasoning. This is causing static perimeter defenses to fail. The teams that come through this well will be the ones that stopped asking “Who’s on the other end of the connection?” and started asking “What’s this request for, and what happens if it succeeds?”</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/intent-not-identity/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Preservation Gap in the AI Stack: Why Capability Is Advancing Faster Than Reconstructability</title>
		<link>https://www.oreilly.com/radar/the-preservation-gap-in-the-ai-stack-why-capability-is-advancing-faster-than-reconstructability/</link>
				<comments>https://www.oreilly.com/radar/the-preservation-gap-in-the-ai-stack-why-capability-is-advancing-faster-than-reconstructability/#respond</comments>
				<pubDate>Thu, 08 Oct 2026 10:54:41 +0000</pubDate>
					<dc:creator><![CDATA[Monika Dvorackova]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19915</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-preservation-gap-in-the-AI-stack.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-preservation-gap-in-the-AI-stack-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[Artificial intelligence systems are difficult to reproduce because their behavior depends on nondeterministic models and on data, configurations, policies, and external services that continue to change. But exact reproducibility is not always what operators, investigators, or auditors need. They often need something different: historical reconstructability. The central claim is that reconstructability is a system property [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Artificial intelligence systems are difficult to reproduce because their behavior depends on nondeterministic models and on data, configurations, policies, and external services that continue to change. But exact reproducibility is not always what operators, investigators, or auditors need. They often need something different: historical reconstructability. The central claim is that reconstructability is a system property established during execution, not inferred later from surviving artifacts. To make that property concrete, this article introduces Orrery, a reference architecture for preserving the runtime bindings and historical dependency states needed to reconstruct an AI system’s past execution.</p>



<p class="wp-block-paragraph">Imagine an AI-assisted decision that must be examined six months after it was made. Rerunning the system may produce a different answer, but an investigator may still need to determine which model endpoint, policy, retrieved data, tool contract, and runtime configuration were used at the time. That is the problem addressed here.</p>



<p class="wp-block-paragraph">The AI industry remains much better at evaluating capability than at preserving the knowledge required to understand past executions. Evaluation pipelines track benchmark scores, task success rates, retrieval precision and recall, tool-call success rates, latency, and cost, and they grow more elaborate with every release. In practice, this machinery often reduces to one acceptance question: <em>What can this system do?</em> That question captures capability, but not whether a past execution will remain understandable after the system changes.</p>



<p class="wp-block-paragraph">That missing property is historical reconstructability: the ability to establish which dependencies were used and under which conditions a particular execution occurred. It is related to, but distinct from, reproducibility. Reproducibility asks whether a result can be produced again under equivalent conditions. Reconstructability asks whether the conditions of the original execution can still be identified, even when repeating that execution would not produce an identical result.</p>



<p class="wp-block-paragraph">This distinction matters when a consequential AI execution completes successfully today and is challenged six months later. The model may still exist in a registry. The application revision may still exist in Git. The trace may still identify the request and the services it crossed. Yet none of these records necessarily reveals which model endpoint served the request, which input transformation and runtime configuration were applied, which data, feature state, or retrieved context entered the computation, or which policy state governed it. In an agentic system, the missing history may also include memory, tool contracts, delegation, and external effects. Even when a dependency identifier was recorded, the historical state to which it referred may no longer be available or resolvable.</p>



<p class="wp-block-paragraph">The problem is that an AI runtime may use dependencies whose historical identities and relationships are not preserved. This is an architectural problem, not merely a logging problem. The gap lies between what the runtime uses and what the surrounding infrastructure is required to preserve. I call this the preservation gap. The AI stack preserves models, code, deployments, and observations, but it does not require the effective historical configuration of a particular execution to remain available. The components may remain while the relations that made them one execution disappear. An AI runtime must assemble today’s execution; it is not necessarily required to remember exactly what it assembled yesterday. Reconstructability must therefore be established during execution, before later mutation erases the bindings that gave the execution its effective configuration.</p>



<p class="wp-block-paragraph"><a href="https://bazel.build/basics/hermeticity" target="_blank" rel="noopener">Hermetic builds</a> illustrate the opposite design pattern. They close the dependency graph before execution begins by pinning inputs, versions, and external dependencies. AI runtimes face the reverse situation: Part of the dependency graph remains open until execution, as model endpoints may resolve through mutable aliases, data and feature state evolve, runtime configuration and policy change, and retrieved context is selected only when a request is processed. Agentic systems extend the graph further through memory, tools, delegation, and external effects. In a hermetic build, dependency closure is an input. In a modern AI runtime, part of it may be an output of the run.</p>



<p class="wp-block-paragraph">Version control records source evolution, transaction logs record state transitions, and distributed tracing records execution relationships. These mechanisms preserve different views of a run but not its materially relevant execution-specific configuration. A trace may retain execution flow, a deployment manifest the declared application state, and a model registry the model revision. What is often missing is the record of which artifacts and states were bound together in a particular execution.</p>



<h2 class="wp-block-heading"><strong>The dependency graph no longer closes at deployment</strong></h2>



<p class="wp-block-paragraph">A deployment records the configuration known before execution begins. AI systems are harder to reconstruct because materially relevant dependencies are selected during execution rather than fixed at deployment. At inference time, the system may resolve a model endpoint, input transformation, data or feature state, retrieved context, runtime configuration, safety controls, and policy. Agentic systems extend this composition through memory, tools, delegation, and external effects. These execution-specific selections are runtime bindings: facts about what the execution actually used.</p>



<p class="wp-block-paragraph">The following example makes this distinction concrete. Consider a hypothetical agentic execution identified as E891. During this execution, the system retrieves two documents, evaluates a policy, invokes a tool, and produces an external effect. The deployment may identify application revision 8f31c2 and model v17, while the execution itself binds prompt h31, retrieval index r42, entities d182@17 and d761@4, tool contract h42, policy v8, and ultimately effect e3. The deployment system could not have known this entire set in advance, because some of these dependencies were selected only as the execution proceeded.</p>



<p class="wp-block-paragraph">The difference between what a deployment declared and what an execution actually used becomes consequential when dependencies change. A model alias may resolve differently, a feature or data source may change, preprocessing and runtime configuration may evolve, and policy v8 may become v9. In agentic systems, retrieved context, memory, and tool contracts introduce further independent mutation. Deployment identity describes what was declared, but the historical record of an execution must describe what was actually used. Unless those binding relations are captured when they occur, the effective configuration of E891 cannot be recovered reliably from later system state.</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecc8b91a&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecc8b91a" class="wp-block-image size-full wp-lightbox-container"><img fetchpriority="high" decoding="async" width="1024" height="538" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451883101-cba585e3-0ee3-488e-9c15-20e8269bcf02_1-1.png" alt="Figure 1. A deployment captures the initial configuration, while its execution-scoped dependency closure emerges as additional dependencies are bound during execution." class="wp-image-19917" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451883101-cba585e3-0ee3-488e-9c15-20e8269bcf02_1-1.png 1024w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451883101-cba585e3-0ee3-488e-9c15-20e8269bcf02_1-1-300x158.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451883101-cba585e3-0ee3-488e-9c15-20e8269bcf02_1-1-767x403.png 767w" sizes="(max-width: 1024px) 100vw, 1024px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption">Figure 1. A deployment captures the initial configuration, while its execution-scoped dependency closure emerges as additional dependencies are bound during execution. (<em>Diagram by the author.</em>)</figcaption></figure>



<h2 class="wp-block-heading"><strong>Current state is a lossy projection</strong></h2>



<p class="wp-block-paragraph">Once runtime bindings are treated as part of the execution, the limitation of current-state records becomes clear. The execution-scoped dependency closure is the combination of the declared state and the runtime bindings that formed the effective configuration of a particular execution.</p>



<p class="wp-block-paragraph">As dependencies mutate, different historical executions can leave behind the same observable evidence. The current state may show which artifacts still exist without revealing which combination a particular execution used. Once the distinguishing bindings are gone, the history cannot be reconstructed from what remains. A system can retain its artifacts while losing reconstructability: An identifier proves neither that a dependency was used nor that the historical state it denotes remains available. Reconstructability therefore requires durable binding relations and continued access to the historical states they identify.</p>



<h2 class="wp-block-heading"><strong>Existence is not usage</strong></h2>



<p class="wp-block-paragraph">The existence of an artifact and its use in a particular execution are different facts. Concurrent versions, caches, retries, and asynchronous changes make it unreliable to infer whether an execution used an artifact merely because that artifact existed at the relevant time.</p>



<p class="wp-block-paragraph"><a href="https://www.w3.org/TR/prov-o/" target="_blank" rel="noopener">W3C PROV</a> already provides the relevant semantic distinction: An activity can use an entity. The missing primitive is therefore not a new provenance vocabulary, but a runtime requirement to record which dependency a particular execution used and to preserve an identity that can be resolved later.</p>



<p class="wp-block-paragraph">This requirement also exposes the limit of observability. <a href="https://opentelemetry.io/docs/specs/otel/overview" target="_blank" rel="noopener">OpenTelemetry</a> can transport dependency identifiers through attributes, links, context, and baggage, but encoding an identifier does not preserve its historical meaning. A trace may retain execution topology while the identities that explain it disappear. Instrumentation can carry preservation metadata, but it cannot guarantee durable storage or continued resolution of the referenced states.</p>



<p class="wp-block-paragraph">Historical reconstructability therefore requires more than retained artifacts. For a defined class of executions and a specified retention period, the system must preserve two properties: binding integrity, which records which dependency versions and states an execution actually used, and resolution integrity, which keeps those recorded identities resolvable to the historical states they denote.</p>



<h2 class="wp-block-heading"><strong>Preservation requires failure semantics</strong></h2>



<p class="wp-block-paragraph">Once reconstructability is treated as a system property, the architecture must define what happens when the records required to reconstruct an execution fail to reach durable storage. Suppose policy v8 authorizes execution E891 to invoke tool contract h42, and the tool commits effect e3. If the usage record fails to reach durable storage, the action succeeds while the historical links among the effect, its policy, the tool contract, and the execution context are lost. The external effect remains, but the historical record no longer reliably explains which policy and tool contract authorized it. The external world and the historical record have diverged.</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecc8c367&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecc8c367" class="wp-block-image size-full wp-lightbox-container"><img decoding="async" width="1024" height="533" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451776273-ff6ef105-d939-4faf-b7e3-1efab5f173f6_1.png" alt="Figure 2. Execution E891 branches into a committed external effect and a failed usage record." class="wp-image-19918" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451776273-ff6ef105-d939-4faf-b7e3-1efab5f173f6_1.png 1024w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451776273-ff6ef105-d939-4faf-b7e3-1efab5f173f6_1-300x156.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/1791451776273-ff6ef105-d939-4faf-b7e3-1efab5f173f6_1-767x399.png 767w" sizes="(max-width: 1024px) 100vw, 1024px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption">Figure 2. Execution E891 branches into a committed external effect and a failed usage record. (<em>Diagram by the author.</em>)</figcaption></figure>



<p class="wp-block-paragraph">Figure 2 shows why this condition cannot be treated as ordinary telemetry loss. When an external effect has been committed but the corresponding usage record is missing, the architecture does not guarantee historical reconstructability. It provides only best-effort historical evidence.</p>



<p class="wp-block-paragraph">Reconstructability becomes an engineering guarantee only when the architecture defines when preservation records become durable and what the system does if they cannot be stored.</p>



<p class="wp-block-paragraph">A system may require a durable usage record before committing the effect, atomically persist the effect intent and preservation record through a <a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html" target="_blank" rel="noopener">transactional outbox</a>, quarantine the execution for reconciliation, or trigger a compensating action. The implementation is application-specific, but preservation failure must have explicit semantics rather than disappear into telemetry.</p>



<p class="wp-block-paragraph">Existing technologies provide the building blocks for this layer. Provenance and <a href="https://openlineage.io/docs/spec/run-cycle" target="_blank" rel="noopener">lineage models</a> represent relations, tracing propagates execution context, registries identify versions, and content-addressed stores preserve artifacts. None, however, creates a runtime obligation to preserve the execution-scoped dependency closure required for reconstruction.</p>



<p class="wp-block-paragraph">This layer is a preservation plane: the contract and machinery that keep an execution’s dependency closure identifiable as its surrounding systems change. It defines what to record, how records become durable, and how referenced states remain resolvable. Orrery applies this principle as a reference architecture in which materially relevant runtime bindings are captured when they occur and retained together with resolvable identities for the historical states they reference.</p>



<h2 class="wp-block-heading"><strong>AIGov Core: reconstructability as a tested property</strong></h2>



<p class="wp-block-paragraph"><a href="https://github.com/AIGovDev/aigov-core" target="_blank" rel="noopener">AIGov Core</a> provides a limited implementation of this principle. Its evidence ledger stores hash-chained events, while its replay engine reconstructs a governance verdict from a stored export without querying live state. A companion check labels the export Ready, Partial, or NotReady according to unresolved lineage, missing policy artifacts, version drift, or replay failure. This demonstrates that reconstructability can be tested, but it does not solve the complete preservation problem. The system can replay only what was recorded; it cannot recover a runtime binding that was never captured.</p>



<p class="wp-block-paragraph">The preservation gap is the mismatch between what an execution depends on and what the system persists. Closing it requires a runtime obligation: Whenever an execution binds materially relevant state, the system must record that binding, preserve a durable identity for the dependency, and retain access to the historical state it denotes for the defined retention period. Reconstructability is therefore not a by-product of capability or observability. It is an architectural property that must be designed, tested, and enforced during execution.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-preservation-gap-in-the-ai-stack-why-capability-is-advancing-faster-than-reconstructability/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Build Your Own Post-training Pipeline</title>
		<link>https://www.oreilly.com/radar/build-your-own-post-training-pipeline/</link>
				<comments>https://www.oreilly.com/radar/build-your-own-post-training-pipeline/#respond</comments>
				<pubDate>Wed, 07 Oct 2026 10:54:16 +0000</pubDate>
					<dc:creator><![CDATA[Sharon Zhou]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19892</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Build-your-own-post-training-pipeline-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Build-your-own-post-training-pipeline-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Implement the key pieces of the ChatGPT pipeline: SFT, reward model training, and PPO.]]></custom:subtitle>
		
				<description><![CDATA[This is the final post in a four-part series about post-training. If you missed them, check out part 1, part 2, and part 3. Time to get your hands dirty! I’ll take you through implementing the key pieces of the classic ChatGPT pipeline: SFT, then reward model training, then PPO. The goal isn’t to reproduce [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>This is the final post in a four-part series about post-training. If you missed them, check out <a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">part 1</a>, <a href="https://www.oreilly.com/radar/the-two-pillars-of-post-training-reinforcement-learning-and-supervised-fine-tuning/" target="_blank" rel="noopener">part 2</a>, and <a href="https://www.oreilly.com/radar/the-post-training-process-openai-used-for-chatgpt/" target="_blank" rel="noopener">part 3</a>.</em></p>
</blockquote>



<p class="wp-block-paragraph">Time to get your hands dirty! I’ll take you through implementing the key pieces of the classic ChatGPT pipeline: SFT, then reward model training, then PPO. The goal isn’t to reproduce InstructGPT, because that took a large team and thousands of GPU-hours. But getting hands-on will help make the concepts from this series concrete, so that when you look at each stage’s code, you understand what it’s doing to the model and why.</p>



<p class="wp-block-paragraph">We’ll use <a href="https://huggingface.co/Qwen/Qwen2.5-1.5B" target="_blank" rel="noopener">Qwen2.5-1.5B</a> as the base model. It’s small enough to train on a single node with a few GPUs but large enough that you can observe real behavioral changes from post-training.</p>



<p class="wp-block-paragraph">For SFT, you’ll use <a href="https://github.com/meta-pytorch/torchtune" target="_blank" rel="noopener">torchtune</a>, Meta’s PyTorch-native fine-tuning library. For RL, you’ll use <a href="https://github.com/verl-project/verl" target="_blank" rel="noopener">verl</a> (from ByteDance’s team), a production-grade RL post-training framework <a href="https://arxiv.org/abs/2409.19256" target="_blank" rel="noopener">built on Ray</a>. verl was designed from the ground up for distributed RL training with proper rollout orchestration, and it’s what you’d actually reach for if you were doing this at scale.</p>



<p class="wp-block-paragraph">You’ll need a machine with at least 2 GPUs (4 or 8 is better) and around 80 GB of total GPU memory.</p>



<p class="wp-block-paragraph">You’ll likely use an LLM to flesh out the boilerplate and get this running end to end. I used LLMs pointed to the recent docs to generate these. What matters as you read the code in this section is that you understand the config choices and data formats well enough to direct that process and debug when something goes wrong.</p>



<h2 class="wp-block-heading">Stage 1: SFT on demonstrations</h2>



<p class="wp-block-paragraph">SFT is the easiest stage to implement. You’re taking the pretrained model and training it on <code>{prompt, response}</code> pairs using the same next-token prediction loss from pretraining, except you only compute loss on the response tokens.</p>



<p class="wp-block-paragraph">Your SFT data should be in a conversational format. Here’s what a single example looks like:</p>



<pre class="wp-block-code"><code>{
    "messages": &#91;
        {"role": "user", "content": "Why do people like golden retrievers?"},
        {"role": "assistant", "content": "Golden retrievers are one of the most 
popular dog breeds for good reason. They have a gentle, friendly temperament 
that makes them great family dogs. They're patient with kids and tend to get 
along well with other animals. They're also highly trainable and eager 
to please, which is why they're commonly used as service dogs and therapy dogs. 
And honestly, the goofy enthusiasm doesn't hurt either. They never really grow 
out of that puppy energy."}
    ]
}
</code></pre>



<p class="wp-block-paragraph">Remember what a pretrained model does with &#8220;<a href="https://www.oreilly.com/radar/introduction-to-post-training/" target="_blank" rel="noopener">Why do people like golden retrievers?</a>&#8221; It might spit out &#8220;Why do people like labrador retrievers?&#8221; because it’s just completing patterns. Your SFT data is what teaches it to actually answer the question instead.</p>



<p class="wp-block-paragraph">In practice, you’d have thousands of these examples covering a range of tasks: question answering, summarization, creative writing, coding help, multiturn dialogue, and refusals for harmful requests. For this walkthrough, assume you have a JSONL file of these conversations.</p>



<p class="wp-block-paragraph">torchtune uses YAML configs and built-in recipes. Here’s what a config might look like:</p>



<pre class="wp-block-code"><code># sft_config.yaml

# Model
model:
  _component_: torchtune.models.qwen2_5.lora_qwen2_5_1_5b
  lora_attn_modules: &#91;'q_proj', 'k_proj', 'v_proj', 
'output_proj']
  lora_rank: 64
  lora_alpha: 128

# Tokenizer
tokenizer:
  _component_: torchtune.models.qwen2_5.qwen2_5_tokenizer
  path: /path/to/Qwen2.5-1.5B/vocab.json
  merges_file: /path/to/Qwen2.5-1.5B/merges.txt
  max_seq_len: 2048

# Checkpointer — loads the pretrained weights
checkpointer:
  _component_: torchtune.training.FullModelHFCheckpointer
  checkpoint_dir: /path/to/Qwen2.5-1.5B
  output_dir: ./sft_checkpoint
  model_type: QWEN2

# Dataset
dataset:
  _component_: torchtune.datasets.chat_dataset
  source: sft_data.jsonl
  conversation_style: sharegpt
  max_seq_len: 2048
  train_on_input: false           # only compute loss on the 
                                    assistant's response tokens

# Training
seed: 42
epochs: 3
batch_size: 4
gradient_accumulation_steps: 4    # effective batch 
                                    size of 16
optimizer:
  _component_: torch.optim.AdamW
  lr: 2e-5
  weight_decay: 0.01
lr_scheduler:
  _component_: torchtune.training.lr_schedulers.get_cosine_schedule_with_warmup
  num_warmup_steps: 50

dtype: bf16
compile: false
</code></pre>



<p class="wp-block-paragraph">Then launch it:</p>



<pre class="wp-block-code"><code>tune run lora_finetune_distributed --config 
sft_config.yaml</code></pre>



<p class="wp-block-paragraph">What actually matters in this config:</p>



<ul class="wp-block-list">
<li><code>train_on_input: false</code> masks the loss on prompt tokens so the model only learns to generate good responses, not to mimic user messages. If you accidentally set this to true, the model wastes capacity learning to produce prompts.</li>



<li>A learning rate of 2e-5 is the standard starting point for SFT. That’s aggressive enough to shift behavior in a few epochs but not so large that you destroy the pretrained weights.</li>



<li>Use LoRA instead of full fine-tuning. For a 1.5B model you could do either, but LoRA is the practical default because it’s faster, uses a lot less GPU memory, and at this model size the quality gap is negligible. At larger model sizes, LoRA becomes even more useful in saving compute.</li>



<li>3 epochs because the dataset is small. InstructGPT trained for 16 epochs on ~13K examples. Smaller datasets need more passes, but watch validation loss for overfitting. This is something you should tune to see different results across different-sized datasets.</li>
</ul>



<p class="wp-block-paragraph">After this stage, try chatting with the model and try a bunch of different comparisons to the base model. Ask it &#8220;Why do people like golden retrievers?&#8221; and you should get a real answer now, not a list of related questions. The model should be able to hold a conversation and follow basic instructions. It’s already dramatically more useful than the base model, even if the responses aren’t always great.</p>



<p class="wp-block-paragraph">Realistically, you should set up a good evaluation (&#8220;evals&#8221;) to assess the quality of the model, hyperparameter tune, and determine your data mix. In this section, you’ll just focus on looking at the code.</p>



<h2 class="wp-block-heading">Stage 2: Training the reward model</h2>



<p class="wp-block-paragraph">For your reward model, start with your data of preference pairs, which will look like this:</p>



<pre class="wp-block-code"><code>{
    "prompt": "What's the capital of that country that celebrates with a lot 
of colored powders?",
    "chosen": "You're probably thinking of Holi, the festival where people 
throw bright colored powders. That celebration is most famously associated 
with India. The capital of India is New Delhi.",
    "rejected": "The capital of that country that celebrates with a lot of 
colored flowers is Amsterdam, the Netherlands, known for its tulip festivals."
}
</code></pre>



<p class="wp-block-paragraph">The &#8220;rejected&#8221; response isn’t just wrong about the festival. It also misread &#8220;powders&#8221; as &#8220;flowers&#8221; and jumped to an incorrect answer.</p>



<p class="wp-block-paragraph">For reward model training, you’ll use TRL’s <a href="https://huggingface.co/docs/trl/v1.9.0/en/reward_trainer#trl.RewardTrainer" target="_blank" rel="noopener">RewardTrainer</a>. Reward model training is a straightforward classification task.</p>



<pre class="wp-block-code"><code>from transformers import AutoModelForSequenceClassification, 
AutoTokenizer
from trl import RewardTrainer, RewardConfig
from datasets import load_dataset

# Start from the SFT checkpoint — it already understands 
the response distribution
model = AutoModelForSequenceClassification.from_pretrained(
    "./sft_checkpoint",
    num_labels=1,
    torch_dtype="bfloat16",
)
tokenizer = AutoTokenizer.from_pretrained("./sft_checkpoint")
tokenizer.pad_token = tokenizer.eos_token

dataset = load_dataset("json", 
data_files="preference_data.jsonl", split="train")

training_args = RewardConfig(
    output_dir="./reward_model",
    per_device_train_batch_size=8,
    num_train_epochs=1,    # just 1 epoch — reward models overfit fast
    learning_rate=1e-5,    # lower than SFT, be gentle
    bf16=True,
    max_length=2048,
    logging_steps=10,
    save_strategy="epoch",
)

trainer = RewardTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset,
    tokenizer=tokenizer,
)

trainer.train()
trainer.save_model("./reward_model/final")
</code></pre>



<p class="wp-block-paragraph">As you skim the code, take note of these three things:</p>



<ul class="wp-block-list">
<li>Start from an SFT checkpoint, not a base model. The reward model needs to understand the distribution of responses it’ll be scoring, and the SFT model is closer to that distribution.</li>



<li>1 epoch only. Reward models overfit quickly. The InstructGPT team found only 1 epoch was needed.</li>



<li>Learning rate of 1e-5, lower than SFT. Smaller updates.</li>
</ul>



<p class="wp-block-paragraph">After training, check the reward model by scoring a few responses manually. Give it a clearly good response and a clearly bad one for the same prompt and make sure the good one gets a higher score. If this basic test fails, something is wrong with your data or training. <strong>Don’t skip this step: It’s better to catch problems sooner than many GPU hours later!</strong></p>



<h2 class="wp-block-heading">Stage 3: PPO with verl</h2>



<p class="wp-block-paragraph">Now you can take the SFT model and optimize it against the reward model using PPO. verl handles the hard parts: coordinating rollout generation across workers, managing the four models that need to be in memory simultaneously (policy, reference, reward, critic), and orchestrating the update loop.</p>



<p class="wp-block-paragraph">First, prepare your prompts. verl expects parquet format. Each row needs a <code>prompt</code> field containing the tokenized and chat-template-formatted prompt. Here’s an example of preparing it:</p>



<pre class="wp-block-code"><code>import pandas as pd
from transformers import AutoTokenizer
from datasets import load_dataset

tokenizer = AutoTokenizer.from_pretrained("./sft_checkpoint")

raw_prompts = load_dataset("json", data_files="rl_prompts.jsonl", split="train")

def format_prompt(example):
    messages = &#91;{"role": "user", "content": example&#91;"prompt"]}]
    formatted = tokenizer.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    return {"prompt": formatted}

formatted = raw_prompts.map(format_prompt)
df = pd.DataFrame(formatted)
df.to_parquet("data/rl_prompts.parquet")
</code></pre>



<p class="wp-block-paragraph">The prompts for RL don’t need labels. The reward model provides the signal. You just need a diverse set of prompts that covers the types of requests your model expects to see when you deploy it.</p>



<p class="wp-block-paragraph">Next, define the reward function. verl lets you wrap your reward model in a function that takes a batch of prompts and responses and returns rewards:</p>



<pre class="wp-block-code"><code># reward_fn.py — verl will call this during training
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

class RewardFunction:
    def __init__(self, reward_model_path="./reward_model/final"):
        self.model = AutoModelForSequenceClassification.from_pretrained(
            reward_model_path, torch_dtype=torch.bfloat16
        ).cuda().eval()
        self.tokenizer = AutoTokenizer.from_pretrained(reward_model_path)
        self.tokenizer.pad_token = self.tokenizer.eos_token

    def __call__(self, prompts, responses):
        """
        prompts: list of prompt strings
        responses: list of response strings
        Returns: list of scalar rewards
        """
        texts = &#91;p + r for p, r in zip(prompts, responses)]
        inputs = self.tokenizer(
            texts, return_tensors="pt", padding=True,
            truncation=True, max_length=2048
        ).to(self.model.device)
        with torch.no_grad():
            rewards = self.model(**inputs).logits.squeeze(-1)
        return rewards.tolist()
</code></pre>



<p class="wp-block-paragraph">Now configure the PPO training run. verl uses YAML configuration files that specify the models, hyperparameters, and infrastructure:</p>



<pre class="wp-block-code"><code># ppo_config.yaml
data:
  train_files: data/rl_prompts.parquet
  prompt_key: prompt
  max_prompt_length: 1024
  max_response_length: 1024

actor_rollout_ref:
  model:
    path: ./sft_checkpoint
  actor:
    optim:
      lr: 1e-6                    # very low — RL updates should be gentle
      lr_warmup_steps: 10
    ppo_mini_batch_size: 64
    ppo_micro_batch_size: 8       # adjust based on GPU memory
    ppo_epochs: 4                 # number of PPO update passes per batch
    clip_ratio: 0.2               # PPO clipping — standard value
    kl_penalty_coeff: 0.1         # KL penalty to prevent reward hacking
    entropy_coeff: 0.01           # small entropy bonus for exploration
  rollout:
    temperature: 0.7
    top_p: 0.9
    n: 1                          # 1 response per prompt per rollout
    tensor_model_parallel_size: 1
  ref:
    log_prob_micro_batch_size: 8

critic:
  model:
    path: ./reward_model/final    # initialize critic from reward model
  optim:
    lr: 1e-5
  ppo_micro_batch_size: 8

reward_model:
  path: ./reward_model/final
  micro_batch_size: 8

trainer:
  total_training_steps: 500
  save_freq: 100
  test_freq: 50
  project_name: post_training_walkthrough
  logger: wandb
</code></pre>



<p class="wp-block-paragraph">A few important hyperparameters in the config:</p>



<ul class="wp-block-list">
<li>The learning rate is an order of magnitude smaller than SFT, default at 1e-6. (&#8220;Actor&#8221; refers to the policy model.) RL updates are noisier, and you want to move slowly. If the model starts producing gibberish or repetitive text, your learning rate is probably too high.</li>



<li>The KL penalty coefficient is 0.1, and recall that it controls how much the model is penalized for drifting from the SFT checkpoint. 0.1 is a reasonable starting point, but you’ll likely need to adjust. If you see reward hacking, turn it up, but if the model doesn’t really change, you need to lower it.</li>



<li>The clip ratio is default set to 0.2, from the original PPO paper. This likely doesn’t need tuning. It prevents any single update from changing the policy too much.</li>
</ul>



<p class="wp-block-paragraph">Launch the training:</p>



<pre class="wp-block-code"><code>python -m verl.trainer.main_ppo \
    --config ppo_config.yaml \
    --n_gpus 4
</code></pre>



<p class="wp-block-paragraph">During training, watch for a few things. The reward should generally trend upward, meaning the model is learning to produce responses the reward model likes. The KL divergence should increase but not explode, in which case the model is drifting too far and you need a higher KL penalty. And periodically generate some responses from the current checkpoint and read them yourself. Metrics can be unreliable on their own, and your own judgment of output quality can be more reliable. Of course, ideally, you have a held out evaluation set that you’re working with that can help you test different tasks you care about.</p>



<h2 class="wp-block-heading">Putting it together</h2>



<p class="wp-block-paragraph">The full pipeline runs on stage after the next: SFT first, reward model second, PPO third. Each stage depends on the previous one. The SFT model provides the starting point for both RL training and the reward model. The reward model provides the signal for PPO. And PPO produces the final model.</p>



<p class="wp-block-paragraph">This is, structurally, the same pipeline that produced InstructGPT and the first version of ChatGPT. The scale is obviously different. They used 175B parameter models, 40 labelers, and far more compute, but the mechanics are identical. Your 1.5B model won’t write poetry as well as the latest GPT, but it will go from producing &#8220;Why do people like labrador retrievers?&#8221; to actually answer &#8220;Why do people like golden retrievers?&#8221; with a conversational response. That’s post-training doing the heavy lifting.</p>



<p class="wp-block-paragraph">If you’re doing this for real, what will eat most of your time isn’t the code but the data: curating good SFT demonstrations, collecting reliable preference labels, and building a prompt set for RL that covers the right distribution of tasks. But there are some great open source datasets out there that you can get started with.</p>



<p class="wp-block-paragraph">The training infrastructure is largely solved: You can point an LLM at the torchtune and verl docs and have it scaffold a working pipeline for you in an afternoon. But no amount of tooling fixes bad data or a reward model that scores the wrong things highly. Understanding what each stage needs and why is what lets you direct that process effectively.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/build-your-own-post-training-pipeline/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Radar Trends to Watch: October 2026</title>
		<link>https://www.oreilly.com/radar/radar-trends-to-watch-october-2026/</link>
				<comments>https://www.oreilly.com/radar/radar-trends-to-watch-october-2026/#respond</comments>
				<pubDate>Tue, 06 Oct 2026 10:54:20 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[Radar Trends]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19889</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2023/06/radar-1400x950-3.png" 
				medium="image" 
				type="image/png" 
				width="1400" 
				height="950" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2023/06/radar-1400x950-3-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Developments in hardware, privacy, software development, and more]]></custom:subtitle>
		
				<description><![CDATA[In addition to the nearly constant stream of model releases, in September we&#8217;ve seen price drops, new kinds of models, proofs of long-standing problems in mathematics, and continued investigations into models escaping their sandboxes. (Axios reports investigations into over 10,000 incidents.) AI has infinite patience and is fundamentally probabilistic. Given a difficult or impossible task [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">In addition to the nearly constant stream of model releases, in September we&#8217;ve seen price drops, new kinds of models, proofs of long-standing problems in mathematics, and continued investigations into models escaping their sandboxes. (Axios reports investigations into over 10,000 incidents.) AI has infinite patience and is fundamentally probabilistic. Given a difficult or impossible task and an unlimited token budget, an agent will eventually attempt to solve the problem in ways that you don&#8217;t expect, and may not want. It&#8217;s easy (and correct) to blame inadequate security procedures at the frontier AI labs, but AI adopters must be careful not to make the same mistakes. The humans using AI need to be accountable for what their agents do.</p>



<h2 class="wp-block-heading">AI Models</h2>



<p class="wp-block-paragraph"><em>Model choice is starting to hinge on price and specialization as much as raw benchmark leadership. Alongside general chat models, there are now decision models that never chat, spatial models built for robot planning and camera control, forecasting models sized for a single task, and cybersecurity-specialized models kept behind an invite-only program. Specialization leads to greater efficiency and lower costs, at least in the short term. In the long term, specialized models may succumb to the &#8220;<a href="https://en.wikipedia.org/wiki/Bitter_lesson" target="_blank" rel="noopener">bitter lesson</a>.&#8221;</em></p>



<ul class="wp-block-list">
<li>Anthropic has <a href="https://www.anthropic.com/claude-opus-5-5" target="_blank" rel="noopener">released</a> Claude Opus 5.5, which it claims has performance similar to Fable 5.1, and hence similar restrictions. It’s faster and requires fewer resources to run. Anthropic has dropped prices 20% for input and output tokens and 60% for cached reads. Not to be outdone, OpenAI released GPT-6 Sol and Luna, with 50% price reductions.</li>



<li>Anthropic has also <a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" target="_blank" rel="noopener">announced</a> Fable 5.1 and Mythos 5.1. The most significant change appears to be a 75% price reduction for cache reads, which might translate into significant savings for long-running jobs; Anthropic estimates 25%. Mythos is only available to trusted partners. Simon Willison used Fable 5.1 to <a href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/" target="_blank" rel="noopener">animate</a> his pelican-riding-a-bicycle pseudobenchmark.</li>



<li>Anthropic <a href="https://www.anthropic.com/claude-sonnet-5-5" target="_blank" rel="noopener">released</a> Sonnet 5.5 with claims that it’s 30% faster and 30% less expensive for most work. The new model has security limitations similar to those applied to Opus and Fable; it routes to Sonnet 5 if it’s asked to do anything out of bounds.</li>



<li>And finally, as September closes, Anthropic announces a <a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-turns-claude-into-an-ai-marketplace-with-2-000-plus-plugins-and-connectors/" target="_blank" rel="noopener">marketplace</a> for Claude plugins and connectors. At its launch, Claude Marketplace had over 2,000 items.</li>



<li>OpenAI has <a href="https://openai.com/index/gpt-6-astra/" target="_blank" rel="noopener">released</a> GPT-6 Astra, with claims that the company has achieved AGI (artificial general intelligence). Astra’s excellent benchmark scores appear to depend on the use of an <a href="https://thenewstack.io/openai-astra-harness-arc-agi-3/" target="_blank" rel="noopener">unreleased harness</a>. OpenAI has also released GPT-6.1 Sol, with per-token price reductions and claims that it is close to GPT-6 Astra in capabilities</li>



<li>OpenAI has <a href="https://openai.com/index/navier-stokes-solution/" target="_blank" rel="noopener">solved</a> the <a href="https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_existence_and_smoothness" target="_blank" rel="noopener">Navier-Stokes existence and smoothness</a> problem, a mathematical problem in fluid mechanics. This development raises <a href="https://techcrunch.com/2026/09/08/openai-fought-dirty-on-career-making-math-problem-says-nyu-mathematician/" target="_blank" rel="noopener">an ethical question</a>: Did OpenAI train its system on the work of two mathematicians who were close to solving the problem themselves? It also raises practical questions about the future of mathematics. Decorated mathematician Terence Tao <a href="https://mathstodon.xyz/@tao/117237320796901560" target="_blank" rel="noopener">asks</a> whether “the collection of good, fruitful open problems is now being mined in a non-renewable fashion.” An <a href="https://thenextweb.com/news/openai-maths-advisory-group-100-open-problems" target="_blank" rel="noopener">advisory group</a> has been formed to help OpenAI make decisions about releasing mathematical results.</li>



<li>Google has <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/" target="_blank" rel="noopener">released</a> Gemini 3.8 Flash TTS and Flash-Lite TTS. Voice options aren’t limited to a prebuilt library. These models have APIs that allow developers to <a href="https://thenewstack.io/gemini-tts-voice-replication-api/">describe</a> the voice that they want or upload a sample. These custom voices are then assigned an ID so they can be reused.</li>



<li>Google has <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/" target="_blank" rel="noopener">announced</a> Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live, near-real-time conversation. They can process video input. Live Extended Thinking can reason and speak at the same time.</li>



<li>Google has released <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/" target="_blank" rel="noopener">Gemini 3.8 Flash and Flash Cyber</a>. Flash appears to be similar to frontier models on most benchmarks, with Computer Use being the biggest exception. Flash Cyber is specialized for vulnerability detection and mitigation and is only available to defenders in the <a href="https://deepmind.google/fairwind-program/" target="_blank" rel="noopener">Fairwind Program</a>.</li>



<li>Google has <a href="https://thenewstack.io/google-timesfm-3-multivariate-forecasting/" target="_blank" rel="noopener">also released</a> <a href="https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/" target="_blank" rel="noopener">TimesFM-3</a>, a small specialized model for multivariate time series forecasting. The model weights are available on Hugging Face with a license that only allows for noncommercial use. (Source code is available under the Apache open source license.)</li>



<li>Xiaomi <a href="https://mimo.xiaomi.com/mimo-v2-6" target="_blank" rel="noopener">released</a> MiMo-V2.6, its latest large language model. It’s fully open sourced, and based on benchmark results, Xiaomi claims that MiMo is the strongest open model to date. What’s more interesting is the claim that MiMo only cost $3.5 million to train.</li>



<li>TypeSafe’s new <a href="https://x.com/Mappletons/status/2101560333441610133" target="_blank" rel="noopener">decision model</a>, <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener">Jev</a>, is unlike anything else we’ve seen. It doesn’t chat; its output is always strictly typed and accompanied by probabilities that estimate correctness. It’s much faster and less expensive than other leading models. It isn’t open source, but there are already many open source <a href="https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in" target="_blank" rel="noopener">clones</a>.</li>



<li><a href="https://ollaya.dev/" target="_blank" rel="noopener">Ollaya</a> is similar to Ollama, but for running decision open models like Laya (a clone of Jev) locally.</li>



<li>World Labs has released <a href="https://www.worldlabs.ai/blog/atlas" target="_blank" rel="noopener">Atlas</a>, a model for “spatial intelligence.” It uses text, images, video, and 3D data to perform tasks like planning a robot’s movements or changing the camera position in a photograph.</li>



<li>The rumors that <a href="https://arstechnica.com/ai/2026/09/nvidia-buys-hugging-face-the-github-of-ai-for-13-billion/" target="_blank" rel="noopener">NVIDIA would buy Hugging Face</a> are <a href="https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/" target="_blank" rel="noopener">true</a>. NVIDIA is hoping for a proliferation of models that will run on its hardware, and the company promises that Hugging Face will remain a neutral platform, without favoring one model over another.</li>
</ul>



<h2 class="wp-block-heading">Software development</h2>



<p class="wp-block-paragraph"><em>Agents are starting to delegate to, and coordinate with, other agents rather than working solo. Claude Code can break a task apart and hand pieces to other Claude Code instances, Muse Code lets sessions message each other, and Google&#8217;s AX orchestrator exists purely to wire up sandboxes and control communications for swarms of agents doing a task together. That shift is pushing developers to rethink what a source repository needs to record, and to start asking how much all this delegation costs.</em></p>



<ul class="wp-block-list">
<li>Now that Jev has caught everyone’s attention, what can you build with it? <a href="https://github.com/Avinash-jetwani/jevmem" target="_blank" rel="noopener">Jevmem</a> is a memory manager that hooks into Claude Code and prunes the context at every conversational turn.</li>



<li><a href="https://github.com/google/ax" target="_blank" rel="noopener">AX</a> is a new <a href="https://agentexecutor.io/" target="_blank" rel="noopener">agent orchestrator</a> from Google. It isn’t an agent; it’s intended to coordinate many agents to complete a task. It creates sandboxes, wires up Git repos and other resources, and controls outbound communications.</li>



<li>The <a href="https://thenewstack.io/claude-code-parallel-projects/" target="_blank" rel="noopener">latest</a> version of Claude Code can <a href="https://support.claude.com/en/articles/9517075-what-are-projects" target="_blank" rel="noopener">manage Claude projects</a>, breaking a task into subcomponents and delegating the subtasks to other Claude Code instances. Another important change is the ability to read <a href="http://agents.md/" target="_blank" rel="noopener">AGENTS.md</a> if <a href="http://claude.md/" target="_blank" rel="noopener">CLAUDE.md</a> isn’t available.</li>



<li>Google’s <a href="https://blog.google/innovation-and-ai/models-and-research/google-labs/cc-expanding-to-groups/" target="_blank" rel="noopener">CC agent</a> is designed for families. Family members can share data with CC, which has its own user account. It could be used for filling out forms, synchronizing calendars, and other common tasks.</li>



<li>What will replace GitHub? There’s a <a href="https://thenewstack.io/zed-delta-github-alternative/" target="_blank" rel="noopener">growing consensus</a> that we need different kinds of source repositories to deal with the agent-assisted software development. In addition to changes to code, it’s important to record the conversations between the developer and the agent, the architectural decisions, and many other artifacts that don’t make it into traditional source control.</li>



<li>Anthropic is merging its <a href="https://claude.com/blog/cowork-is-now-claude" target="_blank" rel="noopener">Claude Cowork and chat products</a>. Anything users type in a chat session is seen by Cowork, and vice versa. Some sensitive information (health, politics, and gender) is excluded. The feature is on by default but can be disabled in settings, and memory isn’t shared with Claude Code. The company also released Claude Docs and Slides.</li>



<li><a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-wants-claude-to-analyze-your-bank-account-and-financial-data/" target="_blank" rel="noopener">Claude Money</a> is a new feature that will allow users to connect their bank accounts to Claude for analysis. The product appears to be similar to a product from OpenAI.</li>



<li>OpenAI’s Agents API is <a href="https://openai.com/index/introducing-the-agents-api/" target="_blank" rel="noopener">now in public beta</a>. It allows compaction, session orchestration, and tool use, and it can be deployed in OpenAI’s sandbox, a cloud provider’s sandbox, or the developer’s hardware.</li>



<li>Meta has <a href="https://ai.meta.com/muse/" target="_blank" rel="noopener">released</a> Muse, its AI agent. Muse is a “personal agent” designed for tasks like shopping, filling in forms, and dealing with customer service. It has its own secure credential store, so data like passwords and credit card numbers are never sent offsite.</li>



<li>Some open source projects are <a href="https://www.latent.space/p/pr-not-welcome" target="_blank" rel="noopener">shutting down external pull requests</a>, which are largely AI-generated. In some cases, the developer team is using its own agents to create and manage PRs; some are using AI agents to triage external PRs.</li>



<li>Meta has launched <a href="https://developer.meta.com/ai/resources/blog/muse-code-new-plans-and-features/" target="_blank" rel="noopener">Muse Code</a>, another competitor to Claude Code. One important new feature is the ability to send messages to other Muse Code sessions, allowing agents to coordinate on complex problems.</li>



<li>AI providers appear to be <a href="https://thenextweb.com/news/openai-outcome-based-pricing-enterprise" target="_blank" rel="noopener">moving toward outcome-based pricing</a>, at least for major corporate customers. Rather than billing by token, customers are billed for completed tasks. That approach begs the question: When is a task completed?</li>



<li>Now that organizations are concerned with AI budgets, the question of how to evaluate the cost of different models and agents becomes important. <a href="https://thenewstack.io/agent-harness-token-costs/" target="_blank" rel="noopener">What should platform teams measure?</a></li>



<li><a href="https://chatgpt.com/work/" target="_blank" rel="noopener">ChatGPT Work</a> was designed to compete with Claude Cowork, Microsoft Copilot Cowork, Muse Code, and other agents designed for noncoders. Simon Willison <a href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/" target="_blank" rel="noopener">shows how</a> Work goes beyond its competitors. It can perform tasks <a href="https://tech.yahoo.com/ai/chatgpt/article/chatgpt-can-now-sign-into-websites-and-complete-tasks-without-seeing-your-login-details-145915053.html" target="_blank" rel="noopener">on the web</a> for users, even <a href="https://www.zdnet.com/article/chatgpt-can-log-into-your-web-accounts-without-you-now-but-should-you-let-it/" target="_blank" rel="noopener">logging in to websites</a> without sending usernames and passwords to OpenAI; it can execute code with full internet access; it can build and deploy a web application. Whether these features are also risks is an open question.</li>



<li>Anthropic has <a href="https://thenewstack.io/claude-built-in-browser-cowork/" target="_blank" rel="noopener">given</a> Claude Desktop access to a Chromium-based browser that’s built into Cowork, eliminating the need for a Chrome plugin when Claude needs to browse the web.</li>



<li><a href="https://github.com/frazerpearce/TimeLord" target="_blank" rel="noopener">TimeLord</a> is a short Python program that, given a string up to 1,000 characters long, produces a seed for Python’s pseudo-random number generator so that repeated calls reproduce the text. It’s a surprisingly simple hack, though not a statement about randomness or the quality of Python’s PRNG.</li>
</ul>



<h2 class="wp-block-heading">Security</h2>



<p class="wp-block-paragraph"><em>OpenAI’s experiment that attacked Hugging Face is the gift that keeps giving, but the past month has had plenty of news about more conventional attacks, many aided by AI. Security has always been a game of whack-a-mole, in which vulnerabilities are discovered and exploited as fast as defenders can patch them. AI is an important tool for defenders, and it&#8217;s constantly improving, but it’s still behind attackers, especially given the limitations placed on frontier models and the unlimited persistence that attacking agents exhibit.</em></p>



<ul class="wp-block-list">
<li>OpenAI has <a href="https://thenextweb.com/news/openai-cancels-launch-of-gpt-6-1-astra" target="_blank" rel="noopener">postponed</a> the release of GPT-6.1 Astra because it failed its safety tests.</li>



<li>NVIDIA has announced its <a href="https://nvidianews.nvidia.com/news/open-agent-safety-platform" target="_blank" rel="noopener">Open Agent Safety Platform</a>. The reference implementation <a href="https://thenewstack.io/nvidia-openshell-sentry-agents/" target="_blank" rel="noopener">includes</a> NVIDIA OpenShell, which has been enhanced with a policy prover, and NVIDIA Sentry, a service that runs on NVIDIA DPUs.</li>



<li>The <a href="https://www.felonybench.com/" target="_blank" rel="noopener">Felony Bench</a> lists known attacks by agents from the major AI labs against third parties. We don’t know if the Bench will be kept up-to-date, but <a href="https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents" target="_blank" rel="noopener">tens of thousands</a> of security incidents involving OpenAI and Anthropic are now being investigated.</li>



<li>Following through on Dario Amodei’s <a href="https://darioamodei.com/post/we-must-pace-the-frontier" target="_blank" rel="noopener">call</a> to control the speed of frontier model development, Anthropic, OpenAI, and Google are creating a <a href="https://www.cnn.com/2026/09/14/tech/ai-standards-body" target="_blank" rel="noopener">standards consortium</a> for <a href="https://www.latent.space/p/ainews-aef-1-standard-emerges-for" target="_blank" rel="noopener">governing the process</a> of AI development. Meta, xAI, Microsoft, and the Chinese labs are all notably absent.</li>



<li><a href="https://thenewstack.io/inside-out-agent-security/" target="_blank" rel="noopener">Agents need their own identity</a>. Unlike the long-term identities we’re used to, agents need a short-lived identity tied to a revocable certificate and that limits access to resources appropriate for the job. That’s not all of agent security, but it’s a table stakes.</li>



<li>Anthropic has published a lengthy report on the <a href="https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a597377d3b5/Anthropic-Detecting-and-countering-091026.pdf" target="_blank" rel="noopener">misuse</a> of its systems by threat actors. Daniel Meissler has <a href="https://danielmiessler.com/blog/anthropic-misuse-report-september-2026" target="_blank" rel="noopener">published</a> a summary, digesting Anthropic’s report into 117 findings.</li>



<li>A malicious <a href="https://www.bleepingcomputer.com/news/security/malicious-npm-packages-evade-install-script-defenses-at-runtime/" target="_blank" rel="noopener">NPM malware</a> package works by hiding malicious code in the package itself (indexed-btree) rather than simply attacking the install script. This technique makes it significantly harder to detect.</li>



<li>An <a href="https://github.com/ucsd-hacc/NSNFSSSFSFN" target="_blank" rel="noopener">attack</a> against the RSA algorithm allows forging of signatures in some situations. The attack was invented in 2007; this is the first public implementation.</li>



<li><a href="https://www.malwarebytes.com/cybersecurity/basics/fake-captcha-scams" target="_blank" rel="noopener">Fake CAPTCHA pages</a> are being used to spread malware. Victims are frequently sent to those pages when they respond to a phish.</li>



<li>Hugging Face has <a href="https://thenextweb.com/news/hugging-face-open-alignment-initiative-embedded-evaluators" target="_blank" rel="noopener">volunteered</a> to audit AI labs for safety and alignment with human values.</li>



<li>In an <a href="https://arxiv.org/abs/2609.04170" target="_blank" rel="noopener">experiment</a> designed to <a href="https://thenextweb.com/news/deepmind-agents-cheating-whistleblowing-research-swarm" target="_blank" rel="noopener">test AI alignment</a>, DeepMind found that, out of 100 agents, 14% were willing to cheat, 25% were “whistleblowers” that reported cheating, and the remainder didn’t notice.</li>



<li>Threat actors are <a href="https://www.bleepingcomputer.com/news/security/hackers-build-ai-frameworks-for-widescale-credential-theft/" target="_blank" rel="noopener">building frameworks</a> for AI agent-enabled attacks. A human in the loop is no longer needed. Fully automated attackers don’t appear to be using zero-days yet; they’re relying on known vulnerabilities.</li>



<li>OpenAI autonomous AI agents were <a href="https://collusion.wiki/" target="_blank" rel="noopener">found</a> communicating with each other via publicly accessible Wikis, possibly to collaborate on a benchmark.</li>



<li>OpenAI has stated that its unreleased Astra model has reached the “Critical” cybersecurity threshold, which means that it can find new vulnerabilities and run exploits against well-protected systems. Now that Astra is released, access to its cybersecurity capabilities has been limited.</li>
</ul>



<h2 class="wp-block-heading">Infrastructure and Operations</h2>



<p class="wp-block-paragraph"><em>Individuals, corporations, and even nations all face a similar problem: keeping their infrastructure under control. At a minimum, control means keeping data on a laptop, corporate server, or data center; at the other end of the spectrum it means eliminating dependencies on software and services from another nation. Any organization working through an AI transformation has to evaluate its entire stack: What do they need to control, and what can they safely delegate to others?</em></p>



<ul class="wp-block-list">
<li><a href="https://www.dawo.community/en/" target="_blank" rel="noopener">DAWO</a> is a community that’s building an open source “workspace” to support digital sovereignty for the Dutch government. The stack will include AI, an operating system based on <a href="https://nixos.org/" target="_blank" rel="noopener">NixOS</a>, cloud services, and collaboration tools.</li>



<li>Cohere now offers a <a href="https://venturebeat.com/data/coheres-model-vault-now-encrypts-ai-inference-so-even-cohere-cannot-see-enterprise-customers-data" target="_blank" rel="noopener">confidential computing platform for artificial intelligence</a>. The company claims that customer data is never visible to Cohere itself or any cloud providers that are in use; data is processed on GPUs whose memory is encrypted and isolated.</li>



<li>Perplexity has <a href="https://thenewstack.io/perplexity-hybrid-compute-mac/" target="_blank" rel="noopener">announced Hybrid Compute</a>, a feature that allows it to run models and use files and tools directly on a user’s Mac. The company claims that sensitive data will never leave the user’s computer.</li>
</ul>



<h2 class="wp-block-heading">Hardware</h2>



<p class="wp-block-paragraph"><em>It&#8217;s too easy to view consumer devices as innocuous things that sit around and do their job silently. Recent devices include cameras, microphones, and even EEG sensors that are constantly collecting data. Where is that data sent, how is it used, and who might have access to it? These questions need to be asked more often.</em></p>



<ul class="wp-block-list">
<li>LG Smart Televisions have been <a href="https://appleinsider.com/articles/26/09/07/disconnect-your-lg-television-from-the-internet-now" target="_blank" rel="noopener">found</a> to record conversations and other audio, even while turned off. The conversations are sent back to LG. If the set is disconnected from the network, it will attempt to find open WiFi access points to deliver its data.</li>



<li>In part because of <a href="https://www.bbc.com/news/articles/cwp80l0my1x2o" target="_blank" rel="noopener">backlash</a> against Meta’s camera-enabled glasses and their abuse, its <a href="https://thenextweb.com/news/meta-vr-glasses-hearing-aid-connect-2026" target="_blank" rel="noopener">AI glasses</a> now come with or without a camera, and can be used as hearing aids. Well-documented abuse aside, virtual reality will only succeed if there are fashionable, easily wearable products.</li>



<li>Headphones, earbuds, and other devices equipped with EEG sensors are <a href="https://techxplore.com/news/2026-08-brain-earbuds-prompting-urgent-children.html" target="_blank" rel="noopener">appearing on the market</a>. They’re advertised for monitoring fatigue, monitoring sleep, and similar applications. It’s time to ask what happens at the interface between neurology and AI.</li>



<li><a href="https://pollen-robotics.com/microduck/blog/introducing-microduck/" target="_blank" rel="noopener">Microduck</a> is a small bipedal AI-driven robot. It’s trained in simulation with open source software, and the model that results can be shared on Hugging Face. It’s affordable and is available for preorder now, shipping by Christmas.</li>
</ul>



<h2 class="wp-block-heading">Web</h2>



<ul class="wp-block-list">
<li>Cloudflare now <a href="https://blog.cloudflare.com/vary-support/" target="_blank" rel="noopener">supports</a> HTTP Vary, which allows servers to serve different kinds of files at the same URL. This is the “ugliest part” of the HTTP standard. It makes caching very difficult, and it probably should be avoided.</li>



<li><a href="https://webmcp.dev/" target="_blank" rel="noopener">WebMCP</a> is a <a href="https://sreenathmenon.com/blog/2026-08-04-webmcp-teaching-websites-to-talk-to-ai-agents/" target="_blank" rel="noopener">proposed standard</a> that gives websites a small API to register tools that agents can discover and call. It was developed by Google and Microsoft.</li>



<li>A <a href="https://arstechnica.com/tech-policy/2026/08/new-twitter-launches-says-musks-x-gave-up-the-name/" target="_blank" rel="noopener">new Twitter</a>? Operation Bluebird is relaunching Twitter, the service bought by Elon Musk and renamed X.</li>
</ul>



<h2 class="wp-block-heading">Biology</h2>



<ul class="wp-block-list">
<li>Anthropic has built a <a href="https://www.reuters.com/world/anthropic-quietly-sets-up-biology-lab-it-ramps-ai-drug-program-2026-09-18/" target="_blank" rel="noopener">biology lab</a> for experimenting with AI-enabled drug development. Claude assisted in the discovery of an enzyme that might be able to perform <a href="https://www.anthropic.com/news/claude-discovers-novel-enzyme-system" target="_blank" rel="noopener">CRISPR-like gene editing</a>.</li>



<li>To improve its training data for biological applications, the OpenAI Foundation (OpenAI’s nonprofit parent organization) is <a href="https://www.technologyreview.com/2026/09/15/1144129/ai-models-need-more-data-about-biology-and-openai-is-paying-to-create-it/" target="_blank" rel="noopener">buying data</a> from failed biotech companies.</li>



<li>Google has <a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/" target="_blank" rel="noopener">released</a> AlphaGenome Atlas, a database of every possible single letter change to human DNA, and what that change will do.</li>
</ul>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/radar-trends-to-watch-october-2026/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Build Your First Agent with MCP</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-your-first-agent-with-mcp/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-your-first-agent-with-mcp/#respond</comments>
				<pubDate>Mon, 05 Oct 2026 15:55:15 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19886</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/USE_THIS_zero-to-agent-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Turn existing APIs into tools an AI assistant can use]]></custom:subtitle>
		
				<description><![CDATA[Developers already have useful capabilities exposed through REST APIs. The Model Context Protocol (MCP) lets developers make those capabilities available to AI clients without rebuilding the underlying application. In this episode of Zero to Agent in 30 Minutes, Bruce Hopkins, an AI developer, author, and longtime software educator, shows how to do that with an [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/VRvqgSZpRHQ?si=uOJKxvk7Eg427wPy" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Developers already have useful capabilities exposed through REST APIs. The Model Context Protocol (MCP) lets developers make those capabilities available to AI clients without rebuilding the underlying application.</p>



<p class="wp-block-paragraph">In this episode of <em>Zero to Agent in 30 Minutes</em>, Bruce Hopkins, an AI developer, author, and longtime software educator, shows how to do that with an MCP server. His demo wraps an existing stock-data API in an MCP server so an MCP client can call it.</p>



<h2 class="wp-block-heading"><strong>From REST API to MCP, step-by-step</strong></h2>



<ol class="wp-block-list">
<li><strong>Start with an existing API.</strong> Identify the operations and data you want an AI client to access. The demo uses the Twelve Data API to retrieve current and historical stock prices.</li>



<li><strong>Create an MCP server.</strong> Use an MCP SDK to create the layer between the AI client and your existing application logic. In Bruce’s Python example, FastMCP handles the MCP interface while the stock-data functions remain separate.</li>



<li><strong>Expose capabilities as tools and resources.</strong> Register the operations the client should be able to discover and call. The stock-price data is exposed through MCP resources and tools that reuse the same underlying functions.</li>



<li><strong>Describe how the client should use them.</strong> Define clear names, inputs, descriptions, and prompts so the client understands what each capability does and what information it requires. Bruce’s example includes prompts for current prices, historical prices, and expected symbol and date formats.</li>



<li><strong>Connect the server to an MCP client.</strong> Run the server over a supported transport so the client can discover and call its tools and resources.&nbsp;</li>
</ol>



<p class="wp-block-paragraph">You don’t need to replace the systems that already handle your application logic to help them work with agents. You can add an MCP interface around existing capabilities to give an AI client a standard way to discover and use them. Be sure to check out <a href="https://github.com/BruceTraining/MCP-Oreilly-2/" target="_blank" rel="noopener">Bruce’s GitHub repo</a> for working code you can adapt for your own APIs.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">Next week, AI engineer Sajal Sharma returns to <em>Zero to Agent in 30 Minutes</em> to build a personal assistant on OpenClaw. He’ll show how an agent can keep tasks and notes in Markdown, run proactive automations, and deliver scheduled updates such as a regular morning briefing without waiting for a new prompt.</p>



<p class="wp-block-paragraph"><em>Follow along with Zero to Agent in 30 Minutes on</em> <em><a href="https://www.oreilly.com/radar/topics/zero-to-agent-in-30-minutes/" target="_blank" rel="noopener">Radar</a>, or watch the latest episode on</em> <em><a href="https://www.youtube.com/playlist?list=PLMJ6moSi2Cgg" target="_blank" rel="noopener">YouTube</a>,</em> <em><a href="https://open.spotify.com/show/033SYd1qhhBAQuMpgJUVTt" target="_blank" rel="noopener">Spotify</a>,</em> <em><a href="https://podcasts.apple.com/us/podcast/zero-to-agent-in-30-minutes/id6793216641" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. If you&#8217;re an O’Reilly member, you can watch live.</em> <em><a href="https://www.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noopener">Save your seat</a>.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-build-your-first-agent-with-mcp/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>How to Build Reliable AI Agent Systems for Production</title>
		<link>https://www.oreilly.com/radar/how-to-build-reliable-ai-agent-systems-for-production/</link>
				<comments>https://www.oreilly.com/radar/how-to-build-reliable-ai-agent-systems-for-production/#respond</comments>
				<pubDate>Mon, 05 Oct 2026 10:54:33 +0000</pubDate>
					<dc:creator><![CDATA[Angie Jones]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19881</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/How-to-build-reliable-AI-agent-systems-for-production.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/How-to-build-reliable-AI-agent-systems-for-production-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article was originally published on the Agentic AI Foundation blog site and is being republished here with the author’s permission. A customer asks your company’s AI agent to update the shipping address on account 123. The agent relies on a support ticket with a typo and updates account 132 instead. The customer relationship [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article was originally published on the</em> <em><a href="https://aaif.io/blog/how-to-build-reliable-ai-agent-systems-for-production" target="_blank" rel="noopener">Agentic AI Foundation blog site</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">A customer asks your company’s AI agent to update the shipping address on account 123. The agent relies on a support ticket with a typo and updates account 132 instead. The customer relationship management system reports that the update “succeeded.”</p>



<p class="wp-block-paragraph">The API successfully completed the authorized request, but because it had no way to know which account the customer approved versus what was intended, it saw no error.</p>



<p class="wp-block-paragraph">Teams often try to improve reliability by tuning the prompt or choosing a stronger model. But the model did exactly what the surrounding system allowed because the system treated the agent’s decision as the final authority. Prompt tuning and stronger models can improve the agent’s responses, but can’t provide guarantees for actions.</p>



<h2 class="wp-block-heading">Put each decision in a layer that can enforce it</h2>



<p class="wp-block-paragraph">The model should propose an action, while a policy service decides whether the action is allowed. The execution layer can then perform the action within those limits.</p>



<p class="wp-block-paragraph">Now consider that same agent, but with a policy service in place. The agent would propose updating an account after reading a support ticket. Before the update runs, a policy service can compare the request with the change the user approved. If the request falls outside those limits, the service can reject it even when the agent sounds confident.</p>



<p class="wp-block-paragraph">After the policy check, the system can save an audit record that connects the request to the result and identifies the policy version used for the decision.</p>



<p class="wp-block-paragraph">This separation gives each part of the system a job it can handle. The model interprets the request, while deterministic software enforces the conditions that must always hold.</p>



<p class="wp-block-paragraph">However, policy enforcement depends on the information used to request an action. A policy check can’t correct a decision that was built on context the system should never have trusted.</p>



<h2 class="wp-block-heading">Treat agent context as untrusted input</h2>



<p class="wp-block-paragraph">Agents receive instructions from more places than the user’s current prompt. A coding agent may read a <code>SKILL.md</code> file or persisted memory from a previous run. Each source can influence what the agent does next.</p>



<p class="wp-block-paragraph">Stored instructions are useful, but their source may be stale or malicious. Another agent may have written the memory. A downloaded skill may contain instructions that expose credentials.</p>



<p class="wp-block-paragraph">Therefore, stored context needs provenance. The system should record who wrote the context and when, then verify that linked files still match the versions the writer used. A risk label can also tell the runtime whether a piece of context may guide an answer or authorize a write.</p>



<p class="wp-block-paragraph">Warnings help, but they don’t remove the risk. In controlled testing described in “<a href="https://events.linuxfoundation.org/agntcon-mcpcon-north-america/program/schedule/?id=1245977" target="_blank" rel="noopener">Trustworthy Context Is Untrusted By Default</a>,” Shub Argha reports that a trust preamble reduced one class of context contamination from 88.8 percent to 33.3 percent. The remaining failures show why a warning should sit alongside technical checks.</p>



<p class="wp-block-paragraph">Once a team can tell where context came from, the next decision is what an agent may do with it.</p>



<h2 class="wp-block-heading">Give the agent only the authority required for the task</h2>



<p class="wp-block-paragraph">OAuth scopes provide a useful boundary, but a broad <code>write</code> scope still leaves a large decision to the agent. The token may allow the agent to update any record even though the user approved one field on one account, as we saw in the opening example.</p>



<p class="wp-block-paragraph">A narrower capability can represent the exact action the user approved. For example, the runtime could issue a capability that permits one update to the shipping address on account 123. The capability expires after the update, so the agent cannot reuse it for another customer.</p>



<p class="wp-block-paragraph">Narrow capabilities also apply to local execution. An agent that needs to format one file doesn’t need unrestricted shell access. The runtime can provide a tool with the necessary input and keep other commands unavailable.</p>



<p class="wp-block-paragraph">As a result, a prompt injection has less authority to work with. The agent may still request the wrong action, but the runtime can reject anything outside the capability it received.</p>



<p class="wp-block-paragraph">Narrow permissions also make audit records easier to understand because the record contains the authority granted for that specific action. When authority lives in a broad token or a long prompt, an operator has to reconstruct what the agent was supposed to do after the failure.</p>



<p class="wp-block-paragraph">Even narrow authority doesn’t require an agent to act. A reliable system also needs clear conditions for when the agent should stand down.</p>



<h2 class="wp-block-heading">Make stopping part of normal operation</h2>



<p class="wp-block-paragraph">An agent can cause damage even without calling a sensitive tool. It can post a wrong answer to a customer or keep replying after a human has taken over.</p>



<p class="wp-block-paragraph">The runtime should make restraint part of the workflow. If the agent’s confidence falls below a set threshold, the runtime can route the support ticket to a human representative. A reply from the human can then cancel any pending response from the agent.</p>



<p class="wp-block-paragraph">Confidence checks and rules that stop the agent when a human takes over belong in the routing and execution layers. The model shouldn’t make either decision on its own. Similarly, a kill switch must stop an agent even when the model is in the middle of a plan.</p>



<p class="wp-block-paragraph">Stopping safely solves one part of reliability, but long-running agents also need a plan for failures that occur after valid work begins.</p>



<h2 class="wp-block-heading">Preserve state so the system can recover</h2>



<p class="wp-block-paragraph">Consider a browser agent that has filled out most of a form when the page changes. A retry that starts from the beginning could submit an earlier step twice, while a retry that guesses where to continue may skip a required field.</p>



<p class="wp-block-paragraph">The execution system should record each confirmed step and attach an idempotency key to any action that must happen once. The key is a unique identifier that tells the server a retry belongs to the same action. Then, when the workflow resumes, the system can continue from the last confirmed state without repeating a completed action.</p>



<p class="wp-block-paragraph">Agent sandboxes create a related problem. Keeping every sandbox running during long idle periods wastes resources and keeps execution environments available longer than needed. Hibernation can reduce both concerns, but only when the wake-up process restores the required state and handles a failed resume.</p>



<p class="wp-block-paragraph">Recovery becomes much harder when the system can’t explain what happened before the interruption.</p>



<h2 class="wp-block-heading">Keep evidence that people and agents can inspect</h2>



<p class="wp-block-paragraph">An execution log should connect the context an agent received to the action it requested. It should also show which policy allowed the action and what result came back.</p>



<p class="wp-block-paragraph">Teams can use that record during an incident, but the agent can also use it during later work. For example, an agent that can read the reason behind a previous code change doesn’t need to infer intent from the final diff alone.</p>



<p class="wp-block-paragraph">Context graphs offer one way to preserve those relationships across time. A graph can connect a tool call to the policy that governed it. The same record can link the source context to the outcome. Later, a person or agent can query the relationships to understand why the system made a decision.</p>



<p class="wp-block-paragraph">The record still needs limits. Secrets shouldn’t be copied into an audit trail, for example, and retention rules should match the data involved.</p>



<p class="wp-block-paragraph">A deliberate record is safer than relying on a conversation transcript and hoping it contains everything an operator will need.</p>



<p class="wp-block-paragraph">A reliable agent system can explain why an action was allowed and resume safely when that action fails. Enforceable policies and recorded state provide those guarantees around the model.</p>



<h2 class="wp-block-heading">Keep the controls portable</h2>



<p class="wp-block-paragraph">Many organizations will use more than one model or agent runtime. When each agent carries its permissions inside a prompt, the rules can drift as teams add models and tools.</p>



<p class="wp-block-paragraph"><a href="https://aaif.io/projects/model-context-protocol" target="_blank" rel="noopener">MCP</a> gives clients and servers a shared way to describe and call tools. Teams can use that common tool surface to enforce authorization at the server or <a href="https://aaif.io/projects/agentgateway" target="_blank" rel="noopener">gateway</a>, regardless of which model requested the action.</p>



<p class="wp-block-paragraph">A shared context format can also preserve provenance when a workflow moves between agents. Open specifications provide consistent interfaces, while each deployment remains responsible for its policies and enforcement.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/how-to-build-reliable-ai-agent-systems-for-production/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: Agents Are Outrunning the Systems Around Them</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/#respond</comments>
				<pubDate>Fri, 02 Oct 2026 15:54:24 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19874</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_thisweek-ai-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Cyberattacks, power constraints, faster model releases, and consumer agents put pressure on the systems built to support them]]></custom:subtitle>
		
				<description><![CDATA[On the latest episode of This Week in AI, host Vicki Reyzelman, a senior solutions engineer at Akamai, traced a common problem across cybersecurity, energy, model releases, consumer hardware, and regulation. AI agents can now probe networks, coordinate with other agents, make purchases, and interact with real-world systems faster than many organizations can respond. We’re [&#8230;]]]></description>
								<content:encoded><![CDATA[
<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/BmBTbqm1pfQ?si=L0ZAcjvz2ZLxoqRh" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">On the latest episode of <em>This Week in AI</em>, host Vicki Reyzelman, a senior solutions engineer at Akamai, traced a common problem across cybersecurity, energy, model releases, consumer hardware, and regulation. AI agents can now probe networks, coordinate with other agents, make purchases, and interact with real-world systems faster than many organizations can respond. We’re seeing those capabilities move into systems built for slower, more predictable software.</p>



<h2 class="wp-block-heading"><strong>Security has to operate at agent speed</strong></h2>



<p class="wp-block-paragraph">Vicki opened with an incident in which <a href="https://www.businessday.co.za/world/2026-09-24-albanese-calls-openai-agents-medicare-breach-unacceptable/" target="_blank" rel="noopener">an OpenAI agent reportedly found ways around security controls</a> while researching public information in Australia’s Medicare system. The activity didn’t expose any&nbsp; personal Medicare records, but <a href="https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078" target="_blank" rel="noopener">OpenAI reportedly took 54 days to identify the incident</a> and another month to notify the government. A response cycle measured in weeks can’t keep pace with systems that can test defenses in seconds.</p>



<p class="wp-block-paragraph">She also brought up <a href="https://news.un.org/en/story/2026/09/1168380" target="_blank" rel="noopener">the recent Hugging Face incident</a> involving a swarm of 1,200 agents that exchanged roughly 70,000 messages while coordinating their work. Agents can change tactics faster than traditional security processes play out, so teams can no longer rely on the familiar methods of addressing suspicious behavior. Companies are now experimenting with runtime enforcement, agent sandboxes, enterprise browsers, and other controls that sit closer to execution.</p>



<p class="wp-block-paragraph">Policymakers are searching for workable controls too, from <a href="https://www.rockcybermusings.com/p/weekly-musings-top-10-ai-security-20260918-20260924" target="_blank" rel="noopener">California proposals for emergency AI shutdown mechanisms</a> to international discussions about <a href="https://www.securitycouncilreport.org/whatsinblue/2026/09/artificial-intelligence-high-level-briefing-2.php" target="_blank" rel="noopener">independent model evaluation</a>. Teams can’t govern agent behavior they can’t see, so they need to know what an agent did and when its behavior crossed a boundary.</p>



<h2 class="wp-block-heading"><strong>Power and latency are becoming model decisions</strong></h2>



<p class="wp-block-paragraph">Power is one constraint software teams can’t code their way around. Vicki pointed to a <a href="https://time.news/doe-invests-nearly-2-billion-to-boost-grid-capacity-for-ai-data-centers" target="_blank" rel="noopener">$2 billion US Department of Energy investment</a> across 26 states alongside hundreds of billions of dollars in planned AI spending from Microsoft, Amazon, Alphabet, and Meta. Data centers can add servers quickly, but it won’t make a difference if the grid can’t provide the energy those servers require.</p>



<p class="wp-block-paragraph">Meanwhile, <a href="https://www.searchintel.tech/research/ai-model-release-pace/" target="_blank" rel="noopener">major model releases are arriving roughly every 17 days</a>, with context windows now exceeding one million tokens. Open weight and edge models are advancing too, particularly around low-latency reasoning. More frequent releases and heavier inference workloads put added pressure on networks, compute, and budgets.</p>



<p class="wp-block-paragraph">Solving this challenge may mean companies have to run more reasoning at the edge or locally, where systems can reduce latency and avoid sending every request across the network. That gives teams another architectural choice to make alongside model selection. A frontier model may be appropriate for one workload, while a smaller local model may be faster and cheaper for another.</p>



<h2 class="wp-block-heading"><strong>Consumer agents move autonomy into everyday life</strong></h2>



<p class="wp-block-paragraph">Consumer hardware puts those architecture and governance choices directly in users’ hands. AI-enabled glasses, pendants, and other devices stay with users throughout the day and can learn preferences, connect with outside services, and take actions such as shopping or making reservations. Meta’s new Muse agent is one example of that shift.</p>



<p class="wp-block-paragraph">Meta says the Muse ecosystem <a href="https://explainx.ai/blog/meta-connect-2026-everything-announced-muse-glasses-vr-2026" target="_blank" rel="noopener">already includes roughly 1,500 developer connectors</a>, including integrations with retailers such as Walmart and Best Buy. If more purchases begin with an agent acting for the customer, companies may have to rethink how people discover products and complete transactions. The convenience of Amazon Prime and one-click shopping, for example, looks different when another system is comparing options and buying on a user’s behalf.</p>



<p class="wp-block-paragraph">Muse already ran into problems, including exposing information it wasn’t supposed to and relying on humans to complete some tasks, such as making dinner reservations. Those failures carry more weight when the software can spend money or act on personal preferences. Users and businesses need clear limits on what an agent can access, what it can do without approval, and how those actions are recorded.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">Deploying an agent means taking responsibility for the systems around it. Security controls, power and network constraints, local versus remote inference, and permission boundaries all shape what these systems can safely do in production. For practitioners, the job now includes the architecture around the models.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news and developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noopener">Apple</a>, or wherever you get your podcasts. You can also hear more from Vicki on <a href="https://substack.com/@vickireyzelman" target="_blank" rel="noopener">her Substack</a>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-agents-are-outrunning-the-systems-around-them/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Coding Agents Love Decision Records</title>
		<link>https://www.oreilly.com/radar/coding-agents-love-decision-records/</link>
				<comments>https://www.oreilly.com/radar/coding-agents-love-decision-records/#respond</comments>
				<pubDate>Fri, 02 Oct 2026 11:24:21 +0000</pubDate>
					<dc:creator><![CDATA[Duncan Davidson]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19867</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Coding-agents-love-decision-records.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/Coding-agents-love-decision-records-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Duncan Davidson’s blog and is being republished here with the author’s permission. Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript. Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://duncandavidson.com/agents-love-decisions" target="_blank" rel="noopener">Duncan Davidson’s blog</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript.</p>



<p class="wp-block-paragraph">Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture significant design choices, their context, and the reasons behind them. Like many tools built for human software teams, ADRs work well for coding agents too.</p>



<p class="wp-block-paragraph">Agents often arrive with little memory of yesterday and only a narrow view of a codebase. Even systems with persistent memory may preserve context without establishing whether it is accurate, current, or accepted by the human team. Decision records help them understand the intent behind the code rather than having to infer it. Keeping them in a project repository spares agents from having to trawl through issues, search chats, and perform code archaeology. When you record intent explicitly, an agent is less likely to mistake an implementation detail for a foundational rule.</p>



<p class="wp-block-paragraph">Once a decision enters an agent’s context window, the agent may adhere to it even more rigidly than a human would. In my own work, I’ve seen agents fight tooth and nail to apply an accepted decision even when it is obsolete. In one case, an agent preserved an outdated storage abstraction across a new feature because an ADR still described it as mandatory. Instead of flagging the mismatch, it added another layer to keep the new requirement technically compatible with the old ruling.</p>



<p class="wp-block-paragraph">The first remedy is to give agents explicit permission to question decisions that no longer fit—and to watch for signs that they’re overfitting. But that solves only half the problem. When you invite an agent to update a decision, a second tendency appears: preserving the deliberation. Every clarification becomes an amendment explaining its own existence at the expense of clarity. Small implementation details become rules, and cross-references acquire their own restatements and justifications. The result is overlitigated prose that is hard for humans to read.</p>



<p class="wp-block-paragraph">ADRs should absolutely be readable by humans, especially as we lean on agents to generate more and more code. To counter this, I’ve become explicit in my projects’ <code>AGENTS.md</code> files about how agents should apply and maintain ADRs. Here’s an excerpt from one:</p>



<p class="wp-block-paragraph"><code>Architectural Decision Records (ADRs) are stored as Markdown files in the </code><br><code>docs/decisions directory. Treat accepted ADRs as binding. Proposed ADRs </code><br><code>are non-binding context. Superseded ADRs are historical context and do </code><br><code>not govern current work. If a given task conflicts with an accepted ADR, </code><br><code>stop and discuss whether the task or ADR should change and propose the </code><br><code>change that you think should be made. Propose new ADRs or updates to </code><br><code>existing ones when a change introduces or revises a durable product or </code><br><code>architectural decision.</code> </p>



<p class="wp-block-paragraph"><code>Keep ADRs succinct. Each ADR carries only its current text; Git history </code><br><code>is its changelog, so do not add or maintain amendment logs in ADR </code><br><code>headers. When substantively changing an accepted ADR, add or update a </code><br><code>single Updated: date line after Date:—its presence signals that history </code><br><code>exists and Git has the details. A superseded ADR records a Superseded-On: </code><br><code>date instead of Updated: , matching the Supersedes: line on the ADR that </code><br><code>replaced it. State each rule once in the ADR that owns it and cross-reference </code><br><code>it from other ADRs instead of restating it.</code></p>



<p class="wp-block-paragraph">These instructions are still evolving in my projects, and different projects will need different conventions. Some teams will prefer immutable ADRs that are superseded rather than revised; in my projects, I’m happy to have Git carry that history.</p>



<p class="wp-block-paragraph">If you do something similar, adapt the guidance to your own needs. The essential principle is that each governing ADR should describe the decision currently in force, with enough rationale to apply it. An agent doesn’t need the transcript of every argument. It needs the ruling that governs today and clear permission to stop when the ruling no longer fits.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/coding-agents-love-decision-records/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Future of Software May Be Conversational Rather Than Autonomous</title>
		<link>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/</link>
				<comments>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/#respond</comments>
				<pubDate>Thu, 01 Oct 2026 10:54:32 +0000</pubDate>
					<dc:creator><![CDATA[Robert Englander]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19862</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-future-of-software-may-be-conversational-rather-than-autonomous.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/10/The-future-of-software-may-be-conversational-rather-than-autonomous-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on Robert Englander’s blog site and is being reposted here with the author’s permission. The software industry has become deeply focused on autonomous AI systems. Agents that can replace workers. Agents that can write software. Agents that can operate applications on our behalf. Entire startups are now built around the [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://robenglander.com/writing/future-of-software-conversational/" target="_blank" rel="noopener">Robert Englander’s blog site</a></em> <em>and is being reposted here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">The software industry has become deeply focused on autonomous AI systems. Agents that can replace workers. Agents that can write software. Agents that can operate applications on our behalf. Entire startups are now built around the assumption that the natural end state of AI is autonomy.</p>



<p class="wp-block-paragraph">Some of this work is genuinely useful. AI-assisted coding can improve productivity. Generative systems are already helping people draft documents, summarize information, and accelerate certain kinds of repetitive work. There are clearly domains where more automation makes sense.</p>



<p class="wp-block-paragraph">Still, I increasingly suspect the industry may be underestimating another opportunity that feels both more practical and potentially more transformative over the long term: natural language interfaces sitting on top of deterministic software systems.</p>



<p class="wp-block-paragraph">Much of the current AI narrative assumes the model itself should become the authoritative actor. The AI writes the code. The AI performs the workflow. The AI executes the task. The AI makes the decision. The conversation around “agents” often assumes the system itself should gradually absorb more and more of that responsibility until people become optional.</p>



<p class="wp-block-paragraph">The problem is that large language models are probabilistic systems. They’re incredibly capable. But they’re still statistical machines. They hallucinate, improvise, approximate. In many contexts, that’s perfectly acceptable.</p>



<p class="wp-block-paragraph">Brainstorming, summarization, drafting, translation, and exploratory work all tolerate a degree of uncertainty. Deterministic systems generally don’t.</p>



<p class="wp-block-paragraph">Financial systems have to calculate correctly. Scheduling systems have to preserve consistency. Medical systems have to maintain integrity. Accounting systems have to reconcile accurately. Reliability is still the foundation upon which useful software is built.</p>



<p class="wp-block-paragraph">That’s one reason I think the most important role for LLMs may not be replacing deterministic systems, but reducing the friction between people and those systems.</p>



<p class="wp-block-paragraph">Historically, software interfaces forced people to adapt to machine discipline. We learned command syntax. We navigated menus and workflows. We memorized procedures. We filled out forms in exactly the way the application expected. Even graphical interfaces, which were a huge leap forward, still largely required users to think in terms of the structure of the software itself. Natural language interfaces potentially invert that relationship.</p>



<p class="wp-block-paragraph">Instead of forcing users closer to the system, the system moves closer to human expression. That may sound subtle, but I think it represents a significant shift in how software can be experienced. A user no longer needs to think primarily in terms of application structure or workflow design. The interaction begins to center more naturally around intent.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“Show me how delaying Social Security by two years impacts long-term spending.”</em></p>



<p class="wp-block-paragraph"><em>“Transfer $500 from checking into savings next Friday.”</em></p>



<p class="wp-block-paragraph"><em>“Why did my tax liability increase this year?”</em></p>



<p class="wp-block-paragraph"><em>“Find the contracts signed after January that contain auto-renewal</em></p>



<p class="wp-block-paragraph"><em>language.”</em></p>
</blockquote>



<p class="wp-block-paragraph">None of these requests eliminates the need for deterministic systems underneath. In fact, they depend on them. The natural language layer simply acts as an interpreter between human expression and authoritative execution.</p>



<p class="wp-block-paragraph">That architecture feels considerably more durable to me than the idea that probabilistic systems should become the primary authority layer themselves.</p>



<p class="wp-block-paragraph">One interesting thing about the current AI wave is that language models are often strongest in areas involving interpretation. They’re remarkably good at extracting meaning from ambiguous human communication, maintaining conversational context, translating between representations, and helping users express intent more naturally. Those are fundamentally interaction problems.</p>



<p class="wp-block-paragraph">Meanwhile, the areas where language models remain weakest are usually the areas requiring guarantees, consistency, accountability, and deterministic correctness. Those are system-of-record problems. The current industry conversation often blurs the distinction between the two.</p>



<p class="wp-block-paragraph">I don’t think conversational interfaces reduce the importance of deterministic software. If anything, they increase it. Once users begin interacting through natural language, the validation layer underneath becomes even more critical. Systems have to safely interpret intent, validate operations, preserve constraints, and maintain correctness even when the incoming requests are conversational and ambiguous.</p>



<p class="wp-block-paragraph">The conversational layer improves accessibility. The deterministic layer preserves trust.</p>



<p class="wp-block-paragraph">Both matter.</p>



<p class="wp-block-paragraph">Every major era of computing has involved some kind of interface transition. Mainframes required specialized operators. Personal computers brought graphical interfaces that made computing accessible to nonspecialists. The web normalized hyperlinks, search, and forms. Mobile computing shifted interaction toward touch and gestures. Natural language may become the next major abstraction layer.</p>



<p class="wp-block-paragraph">Not because computers suddenly became human-like, but because we finally built systems capable of translating between human communication and machine discipline at scale.</p>



<p class="wp-block-paragraph">I also think this changes how we should think about software’s future. The current AI environment sometimes frames autonomy as the inevitable destination. If an AI can partially perform a task today, many assume the long-term outcome is full replacement of the person performing that task.</p>



<p class="wp-block-paragraph">I’m not convinced that’s where the most durable value lies.</p>



<p class="wp-block-paragraph">In many domains, the real friction isn’t execution. It’s interface complexity. People struggle less with the underlying capabilities of software than with the difficulty of expressing what they actually want the software to do.</p>



<p class="wp-block-paragraph">Enterprise systems are notoriously difficult to navigate. Financial systems expose overwhelming complexity. Creative tools bury users under layers of workflow and terminology. Even relatively simple applications often require substantial onboarding before users become comfortable with them.</p>



<p class="wp-block-paragraph">Natural language interfaces potentially change that equation in a meaningful way. They allow software to meet users closer to where they already are: ordinary human communication.</p>



<p class="wp-block-paragraph">That doesn’t mean conversational systems should become undisciplined systems. In fact, I think the opposite is true. As interfaces become more conversational, the underlying architecture has to become even more rigorous about validation and execution semantics. The ambiguity doesn’t disappear. It moves.</p>



<p class="wp-block-paragraph">Historically, much of the burden of precision sat on the user. The user had to learn the syntax, understand the workflow, and conform to the application’s structure.</p>



<p class="wp-block-paragraph">Conversational systems shift more of that burden into the interpretation and validation layers of the software itself. That’s not a trivial engineering problem. It requires clarification, normalization, policy enforcement, validation, and authoritative execution underneath the conversational layer. It also requires accepting that probabilistic interpretation and deterministic execution aren’t competing ideas. They’re complementary ones.</p>



<p class="wp-block-paragraph">This is one reason I increasingly think the future of software may become conversational without necessarily becoming autonomous. The two ideas are related. But they’re not the same thing.</p>



<p class="wp-block-paragraph">There’s enormous value in reducing the natural friction between human expression and machine discipline. Large language models may ultimately prove most transformative not when they replace deterministic systems, but when they help people interact with those systems more naturally.</p>



<p class="wp-block-paragraph">For decades, people have adapted to computers. It now seems possible that software may finally start adapting to people instead.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-future-of-software-may-be-conversational-rather-than-autonomous/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Agentic Data Science Playbook</title>
		<link>https://www.oreilly.com/radar/the-agentic-data-science-playbook/</link>
				<comments>https://www.oreilly.com/radar/the-agentic-data-science-playbook/#respond</comments>
				<pubDate>Wed, 30 Sep 2026 16:06:24 +0000</pubDate>
					<dc:creator><![CDATA[Hugo Bowne-Anderson, Luca Fiaschi and Thomas Wiecki]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Software Architecture]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19847</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-agentic-data-science-playbook-image-created-with-Adobe-Firefly.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/The-agentic-data-science-playbook-image-created-with-Adobe-Firefly-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How to build and direct data science agents, and make each investigation improve the next.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Vanishing Gradients and is being republished here with the authors’ permission When an AI agent can explore a dataset, choose a modeling approach, run the analysis, and explain its findings, what should the data scientist do? Traditionally, data scientists chose each step and implemented much of the analysis themselves. [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <a href="https://hugobowne.substack.com/p/the-agentic-data-science-playbook">Vanishing Gradients</a> <em>and is being republished here with the authors’ permission</em></p>
</blockquote>



<p class="wp-block-paragraph">When an AI agent can explore a dataset, choose a modeling approach, run the analysis, and explain its findings, what should the data scientist do?</p>



<p class="wp-block-paragraph">Traditionally, data scientists chose each step and implemented much of the analysis themselves. Agentic data science changes that division of work: we can delegate an investigation, including methodological choices, while shaping the question, supplying relevant expertise, and challenging the evidence it produces. For AI-native data scientists, choosing the runtime, writing reusable skills, and designing the workflows and feedback that guide the agent are part of the analytical work.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1600" height="889" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1600x889.png" alt="Verification process" class="wp-image-19865" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1600x889.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-300x167.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-767x426.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-1536x853.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-38-2048x1138.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">This article provides a playbook for working with data science agents, from setting up an investigation to reviewing its results and carrying lessons into the next assignment. To see why that requires more than a capable model and a business question, consider an experiment we deliberately started with too little guidance. We gave Claude Opus 5.0 a modified version of the <a href="https://www.kaggle.com/datasets/ellipticco/elliptic-data-set" target="_blank" rel="noopener">public Elliptic dataset</a> and asked, “Build me a model to detect fraudulent nodes.” The dataset is a graph of Bitcoin transactions: each node is a transaction, and an edge represents a flow of bitcoin between transactions. Some transaction nodes carry licit or illicit labels based on the entities that created them; the rest are unlabeled. Each node has a time step indicating when its transaction was broadcast, allowing us to train on earlier time steps and test on later ones. We also wanted separate results for transactions with many connections (high-degree nodes), which mattered most in the intended application. Importantly, we renamed the columns, changed several features, and reindexed the time steps while preserving their order, making the public dataset harder for Claude to recognize.</p>



<p class="wp-block-paragraph">Claude wrote the code, trained a random forest, and reported an F1 of 0.87 and ROC AUC of 0.99. It had split transactions randomly, mixing earlier and later time steps in both the training and test sets. That test did not measure how the model would perform on transactions from later time steps. Moreover, Claude also used a feature we had planted as a proxy for the fraud label (yes, we tricked it!), giving the model leaked information it would not have when scoring a new transaction. So how do we avoid these situations?</p>



<p class="wp-block-paragraph"><strong>We then supplied the guidance missing from the initial prompt:</strong>&nbsp; We required a temporal holdout, removed the leaking feature, supplied context about how the model would be used in production, and asked for separate reporting on the high-degree nodes that mattered most. Under the corrected evaluation, F1 was 0.70, overall recall was 0.61, and recall on high-degree nodes was 0.21.</p>



<p class="wp-block-paragraph">The key point is that “Build a fraud detector” left Claude to infer how the model would be used and what would count as success. <strong>AI-native data scientists build and direct</strong> an analytical process in which agents can investigate, receive feedback, and return evidence for review. The work begins with deciding how much of the investigation to delegate.&nbsp; Agentic data science is doing data science work with AI agents as teammates. Crucially, the scope of their responsibility can extend well beyond code implementation. An agent can help frame a question, explore data, test a claim, or communicate a result, provided it has the context and tools to do the work, a way to assess its progress and validate its results.</p>



<p class="wp-block-paragraph">Asking an agent to write a pandas transformation leaves you as the bottleneck, responsible for deciding every next operation. Asking it to investigate a change in customer behavior gives it larger analytical responsibility. It can inspect a result, form another question, choose a method, and continue. The interaction becomes a conversation about the investigation rather than a sequence of requests for code.</p>



<p class="wp-block-paragraph">The question may be <em>descriptive</em> (what happened?), <em>diagnostic</em> (why did it happen?), <em>predictive</em> (what might happen next?), or <em>prescriptive</em> (what should we do?). The fraud model is predictive; the pricing investigation later in this article is diagnostic and causal. Across these kinds of work, we need to specify the question and intended use, then verify that the evidence supports the answer.</p>



<p class="wp-block-paragraph"><em>The following five practices are key to agentic data science:</em></p>



<ul class="wp-block-list">
<li>Frame the investigation.</li>



<li>Equip the agent for the assignment.</li>



<li>Organize the work through bounded experiments, competing analyses, or both, according to the question.</li>



<li>Review the result independently.</li>



<li>Preserve evidence and turn reviewed lessons into reusable expertise.</li>
</ul>



<p class="wp-block-paragraph">The first two practices set up the work. The third determines how the investigation proceeds; the fourth tests its claims. Evidence is captured throughout, and the fifth practice carries reviewed lessons into future assignments.</p>



<p class="wp-block-paragraph">As in agentic software engineering, the agentic data scientist’s two central responsibilities are <strong><em>specification</em></strong> and <strong><em>verification</em></strong>. Specify the question, intended use, and evidence the agent should produce; then verify that its analysis supports the conclusion. Agents can help with both, while the data scientist remains responsible for judging the question and the evidence.</p>



<h2 class="wp-block-heading"><strong>1. Frame the investigation with the agent</strong></h2>



<p class="wp-block-paragraph"><strong><em>Start by discussing the assignment with the agent.</em></strong> Supply the intended use and organizational context, then let it inspect the data and propose an approach. Method selection can be part of its responsibility. Your intervention matters when a proposal changes the question, rests on a questionable assumption, or needs information the agent cannot obtain. Predicting fraud and deciding which flagged entities to investigate, for example, require different evidence about errors and their consequences. A brainstorming skill such as those in <a href="https://github.com/obra/superpowers" target="_blank" rel="noopener">Superpowers</a> can help structure that conversation before you turn it into a task prompt.</p>



<p class="wp-block-paragraph">A useful specification records that shared understanding. It states the decision, relevant constraints, and evidence the investigation should produce. It need not prescribe every step. In the fraud example, “classify transactions from later time steps using only information available when each is scored” matters more than “use a random forest.” The former defines the analytical task, while the latter selects one possible implementation.</p>



<p class="wp-block-paragraph">You can specify what the investigation must establish without specifying the answer you want. “Determine whether the data support a recommendation” leaves room for an inconclusive result. “Keep trying until you find an effect” does not.</p>



<p class="wp-block-paragraph">Turn that discussion into a short analytical brief to give the agent as its task prompt. For an assignment like our fraud example, a starting version could read:</p>



<pre class="wp-block-code"><code>TASK PROMPT:
Question: Can we identify fraudulent nodes as they enter the network?
Use: Support investigation, with separate reporting on high-degree nodes.
Available information: Only inputs known when the node is scored.
Agent discretion: Explore data, propose eligible features, choose models.
Return to me: Unclear feature provenance, changes to the target or
population, or a trade-off that requires an operational decision.
Deliverable: Reproducible analysis, temporal evaluation, subgroup errors,
and a recommendation that states what the evidence cannot establish.</code></pre>



<p class="wp-block-paragraph">Review it with the agent before the investigation proceeds. If exploration reveals that the evidence cannot answer the question, revise the brief explicitly; do not quietly substitute an easier question.</p>



<p class="wp-block-paragraph">The deliverable may still be a notebook, model, or report prepared outside a production service. You can begin in the workspace where you already do that work.</p>



<h2 class="wp-block-heading"><strong>2. Equip the agent for the assignment</strong></h2>



<p class="wp-block-paragraph">The task prompt tells the agent what to investigate. It also needs to know how the project works, reach the data, run the analysis, and check the result. The harness is the system around the language model that allows this: its tools, runtime, context, permissions, and feedback from its actions. Its runtime is the environment that executes those actions. A language model alone cannot inspect a warehouse, run a simulation, or recover an interrupted statistical model fit. The environment must make those operations possible and return useful evidence about what happened.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecca1bc8&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecca1bc8" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="820" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1600x820.png" alt="Data science agent flow" class="wp-image-19849" style="aspect-ratio:1.95" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1600x820.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-300x154.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-767x393.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34-1536x788.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-34.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">Runtime choices are analytical choices as well as engineering choices. Can the agent execute Python or R with the libraries the task needs? Can a long-running fit continue after an interactive session ends? Which scientific libraries should the agent use? Can the agent inspect plots, or does it only see the code that produced them? Can you reproduce the environment in which it reported a result?</p>



<p class="wp-block-paragraph">An existing agent runtime may provide most of this. Configuring it means deciding what belongs in Markdown, what needs a tool, and what should be checked by a small script. In an investigation like the fraud example, Markdown can hold the brief and data definitions, while a Python script could check that the appropriate temporal validation split is executed. A CSV data extract may be enough for exploration; if the agent needs data warehouse access, a tool exposed through an MCP server can provide it with appropriately scoped, read-only credentials. A sentence in a prompt cannot enforce that access limit.</p>



<p class="wp-block-paragraph">Take the same care with outputs. Ask the agent to preserve the data reference, code, environment, assumptions, and diagnostics behind its report. A chat transcript is a poor substitute for a runnable analysis. Review becomes much harder when the only surviving artifact is a confident paragraph about what the agent says it did.</p>



<p class="wp-block-paragraph">Execution is only part of the problem. An agent may know how to fit a model and still misunderstand what the columns mean. It may find five revenue tables and choose the wrong one. A schema rarely explains which customers were eligible for an offer, when a measurement changed, or why the team stopped using an apparently reasonable metric. This is where <strong>agent skills</strong> and <strong>domain knowledge</strong> enter. A skill packages instructions and resources for a type of analytical work. It might contain a modeling approach, example code, required diagnostics, and guidance on when to ask for help. Data documentation supplies the organizational meaning: canonical definitions, table grain, known limitations, and the history needed to interpret a result.</p>



<p class="wp-block-paragraph">A useful skill is specific enough to change the agent’s behavior. “Be rigorous” gives it little to work with. A fraud-modeling skill can require the agent to establish feature availability, evaluate on later observations, and report performance on operationally important subgroups. For example:</p>



<pre class="wp-block-code"><code>For fraud prediction:

Establish what information is available when a node is scored.

Exclude features derived from subsequent investigations or labels.

Fit preprocessing on training data only.

Evaluate on later-arriving nodes and report the required degree groups.

Flag uncertainty about feature provenance before claiming performance.</code></pre>



<p class="wp-block-paragraph">These instructions leave room to choose a model. They encode reasons that some apparently successful models should be rejected. Where a requirement can be checked reliably in code, the skill can call a script that performs the check and records its result.</p>



<p class="wp-block-paragraph">Loading every method and every document into every assignment is unnecessary. Give the agent a way to find relevant expertise, including its scope and exceptions. A forecasting skill should not silently impose its evaluation rules on an unrelated retrospective analysis. Nor should a notebook from last year outrank an updated metric definition merely because it offers convenient code to copy.</p>



<p class="wp-block-paragraph">To put these pieces together locally, begin with a file-and-code agent in a sandboxed project workspace, such as the following:</p>



<pre class="wp-block-code"><code>fraud-investigation/

  AGENTS.md                # Project instructions, where supported by the runtime

  brief.md                 # Agreed question and delegation boundaries

  data-notes.md            # Sources, column meaning, availability times

  skills/fraud.md          # The methodological guidance above

  environment.lock         # Dependency versions, in your tool’s format

  model/                   # Code the investigating agent may change

  results/                 # Experiment log, diagnostics, saved candidates

  review.md                # Acceptance decision and unresolved questions</code></pre>



<p class="wp-block-paragraph">Use a project instruction file, such as <code>AGENTS.md</code> in runtimes that support it, to explain which context files the agent should read and how to propose updates to them. In other runtimes, provide those instructions through the supported mechanism. Give the sandbox read access to the approved development data and write access to the model and results directories. Keep the final test data outside of the agent’s accessible workspace. The practitioner can run acceptance checks in a separate environment whose evaluator and data the investigating agent cannot modify. A different folder, or version control alone, is not an access boundary.</p>



<p class="wp-block-paragraph">Now ask the agent to inspect the inputs, identify unresolved questions, and build a baseline. Before allowing repeated experiments, rerun that baseline and examine its feature-availability record, split dates, and subgroup report. This small rehearsal checks whether the setup works all the way from instructions to evidence. A missing subgroup report points to a different problem than a failed package installation. Resolve those problems before giving the agent a longer run.</p>



<h2 class="wp-block-heading"><strong>3. Organize the investigation</strong></h2>



<p class="wp-block-paragraph"><em>With the question framed and the agent equipped, the next choice is how to organize its work.</em> This depends on the intent of the data science problem. For descriptive work, exploratory data analysis may proceed one question and plot at a time. Building a predictive model may support repeated experiments against a fixed evaluator; a causal question may require comparing analyses built on different assumptions.</p>



<p class="wp-block-paragraph">In a live exploratory analysis on <em>Show Us Your Agent Skills</em>, <a href="https://hugobowne.github.io/show-us-your-agent-skills/agent-skills/guests/eric-ma/" target="_blank" rel="noopener">Eric Ma (Moderna) uses a marimo notebook as a shared workspace with an agent</a>. He explains the protein mutation data, asks for one plot at a time, corrects a color scale that affects interpretation, and chooses the next question from what he sees. The agent edits the notebook and renders the plots; Eric supplies the domain context, checks the artifacts, and owns the interpretation. The reason Eric needed to be in the loop was that human understanding was part of the objective function here!</p>



<h3 class="wp-block-heading"><strong>Use a bounded experiment loop</strong></h3>



<p class="wp-block-paragraph">For predictive modeling, the <a href="https://github.com/karpathy/autoresearch" target="_blank" rel="noopener">autoresearcher pattern</a> organizes the work into a repeatable loop: propose a hypothesis, change the model, evaluate it, and keep or revert the change. The agent records each result and uses it to choose the next attempt. Within the scope you give it, it can explore features and model structure as well as parameter values. This is an inner loop within a broader investigation: the data scientist frames the question and sets the evaluation, the agent searches within those boundaries, and the data scientist reviews the result (potentially using an independent agent) before deciding what to do next.</p>



<p class="wp-block-paragraph">Define what the agent may change, protect the evaluator from those changes, and set a time or compute budget. This makes iteration a bounded task within the investigation.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecca286a&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecca286a" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="1280" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1600x1280.png" alt="The inner loop" class="wp-image-19850" style="aspect-ratio:1.250501002004008" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1600x1280.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-768x614.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-300x240.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35-1536x1229.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-35.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">A compact experiment contract could say:</p>



<pre class="wp-block-code"><code>Improve the supplied baseline within the agreed compute budget.

You may change model code and propose eligible features.

Keep the target definition, validation split, and evaluator fixed.

Record each hypothesis, change, result, and keep-or-revert decision.

Stop at the budget limit or escalate if the evaluation is unsuitable.

Return the best candidate and the experiment history for review.</code></pre>



<p class="wp-block-paragraph">In a separate exercise with its own baseline and evaluation, we used this pattern to improve a graph neural network trained on the network data from the opening example. We used lower validation loss as the rule for keeping a change; F1 for fraudulent transactions was a separate measure of the resulting classifier. The agent ran 41 experiments while the team slept and retained seven changes that reduced validation loss. On the validation set, loss fell by about 70%, and F1 for fraudulent transactions rose from about 0.72 to 0.82. The log preserved both successful changes and failed attempts, so we could examine how it reached the result.</p>



<p class="wp-block-paragraph">One candidate had the highest F1 for fraudulent transactions, but the agent rejected it because its validation loss was higher. That followed the selection rule we had set. The experiment history records that choice for the subsequent review.</p>



<p class="wp-block-paragraph">The autoresearcher pattern works when an objective gives the agent useful feedback on each attempt. But some investigations turn on which assumptions to make, not which candidate scores best. Those tasks need a different way to organize the agent’s work.</p>



<h3 class="wp-block-heading"><strong>Investigate competing explanations</strong></h3>



<p class="wp-block-paragraph">In causal work, no held-out outcome directly reveals what would have happened without an intervention. The agent needs to examine how different analyses construct and test that counterfactual.</p>



<p class="wp-block-paragraph">In a demonstration from our <a href="https://vanishinggradients.short.gy/data-science-agentic" target="_blank" rel="noopener">Master Agentic Data Science course</a> using simulated subscription-business data, we asked: “What did the price increase cost us?” The outcome is daily conversion rate: paid conversions divided by the pool of potential subscribers. Choices about the observation window, counterfactual, exclusions, and validation produce different analytical paths. A final memo usually shows only one.</p>



<p class="wp-block-paragraph">Two agent runs estimated conversion roughly 16% below their no-price-increase counterfactuals, yet shipped opposing claims.&nbsp; Run A attributed its estimated drop to a changing pool of potential subscribers and concluded there was “no real effect,” but did not validate that explanation.&nbsp; Run B backtested its counterfactual, ran a placebo check, and compared six specifications. It reported a robust relative reduction of 15.6%.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecca32ee&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecca32ee" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="945" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1600x945.png" alt="Two-agent run" class="wp-image-19851" style="aspect-ratio:1.6956521739130435" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1600x945.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-767x453.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-300x177.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36-1536x907.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-36.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">The parallel-analysis pattern has independent agents test different choices in the same investigation: one examines the observation window, another tests seasonal assumptions, and we compare their estimates, uncertainty, and diagnostics. <a href="https://github.com/pymc-labs/decision-lab" target="_blank" rel="noopener">Our Decision Lab work</a> extends this approach across analytical paths, using checks to identify unsuitable analyses and unresolved disagreements.</p>



<p class="wp-block-paragraph">Both approaches give the agent feedback while it works. The fixed evaluator steers the model experiments; diagnostics help it compare causal analyses. The output is a candidate and experiment history, or a set of analyses with their assumptions and checks. Those are the materials for the next task: <em>verifying the claim</em>.</p>



<h2 class="wp-block-heading"><strong>4. Review the result independently</strong></h2>



<p class="wp-block-paragraph">The agentic data scientist now needs to check what that evidence supports, and a fresh agent can help. Give an <strong>independent agent reviewer</strong> the original brief, data context, code, diagnostics, and final claim. Ask it to reproduce decisive checks and challenge assumptions. In this adversarial review pattern, the agent raises objections it can substantiate; the data scientist judges whether they change the conclusion.</p>



<p class="wp-block-paragraph">In the fraud exercise, the agent used the same validation data to guide 41 experiments, so the reported gains may partly reflect what worked on that set. Freeze the selected candidate and assess it on an untouched holdout chosen for the intended use, including errors in the groups that matter.</p>



<p class="wp-block-paragraph">In the pricing exercise, a fresh reviewer challenged Run A’s conclusion. Run A attributed the estimated decline to a changing pool of potential subscribers but provided no evidence for that explanation. The reviewer found that conversion had been rising before the price increase and that placebo interventions in earlier periods did not reproduce the negative effect. Run B’s analysis, which included these validation checks, was selected in the final comparison.</p>



<p class="wp-block-paragraph"><a href="https://netflixtechblog.com/a-human-augmenting-agentic-workflow-for-causal-inference-4623f0a9c5af" target="_blank" rel="noopener">Netflix’s agentic workflow</a> for causal inference implements this division: an actor performs the analysis and diagnostics, while a critic challenges the reasoning and claims. Humans can inspect and rerun the artifacts.</p>



<p class="wp-block-paragraph">A fresh agent session is not necessarily an independent review if it can read the investigator’s earlier attempts through the workspace or Git history, though! For a check meant to stand on its own, give the reviewer the original brief, final artifact, and data needed for that check, while limiting access to the prior path. The full experiment trail can be examined separately when auditing how the result was reached.</p>


<div class="wp-block-image">
<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac90ecca3d5a&quot;}" data-wp-interactive="core/image" data-wp-key="6ac90ecca3d5a" class="aligncenter size-large wp-lightbox-container"><img loading="lazy" decoding="async" width="1600" height="1066" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1600x1066.png" alt="Human vs agentic verification" class="wp-image-19852" style="aspect-ratio:1.5" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1600x1066.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-300x200.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-767x511.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37-1536x1024.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/image-37.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>
</div>


<p class="wp-block-paragraph">These examples call for different balances of human and agentic verification. In Eric Ma’s EDA, the agent makes plots while Eric checks them and chooses the next question. In the bounded experiment loop, a fixed evaluator checks each candidate before a person reviews the selected model. In the pricing analysis, agentic diagnostics and critique help a data scientist judge what the evidence supports.</p>



<p class="wp-block-paragraph">The low-low quadrant leaves little basis for trusting a result (in fact, it’s “vibe data science!”). Repeatable checks can move some work toward more agentic verification, while questions that depend on domain understanding or consequential decisions continue to need human judgment.</p>



<h2 class="wp-block-heading"><strong>5. Turn reviewed experience into reusable expertise</strong></h2>



<p class="wp-block-paragraph">The experiment loop and adversarial review both depend on an evidence trail: the saved artifacts that show what the agent did and why a conclusion survived or changed. Preserve that trail throughout each investigation, including data references, code versions, analytical choices, experiment results, diagnostics, and review findings. Keep the failed alternatives and review findings as well as the final report.</p>



<p class="wp-block-paragraph">This trail has a second use beyond inspecting the current result. Reviewing it with the agent can reveal missing context, recurring mistakes, or methods worth reusing. The next question is which of those lessons should change how the agent approaches a future assignment. Leaving them in a conversation makes that learning difficult to carry forward.</p>



<p class="wp-block-paragraph">In the fraud example, we deliberately planted a feature that leaked the fraud label. Removing it corrected that analysis. The reusable lesson is to have the agent check proposed features for leakage: where did each feature come from, and would it be available when a new transaction is scored? That requirement can go into a fraud-modeling skill for future investigations.</p>



<p class="wp-block-paragraph">Before making a lesson into standing guidance, we need to define where it applies. The planted feature was a problem because it carried information unavailable at scoring time, not because it predicted fraud well. The temporal split likewise fits a task involving transactions from later time steps; it is not a rule for every analysis. A skill should capture those conditions so the agent applies the lesson to the right task.</p>



<p class="wp-block-paragraph">For example:</p>



<pre class="wp-block-code"><code>Lesson: a feature encoded information from the fraud label.

Scope: prospective fraud prediction.

Update: require a documented source and availability time for inputs.

Evaluation: test whether the agent detects outcome-derived inputs

without rejecting legitimate signals merely because they predict well.</code></pre>



<p class="wp-block-paragraph">This is where evals enter: repeatable tasks with explicit criteria for assessing the data science agent’s behavior. Here, we evaluate how the agent conducts the analysis, not only its model’s predictive performance. The evidence trail supplies concrete failures that can become test cases for proposed changes to its skills or workflow.</p>



<p class="wp-block-paragraph">Keep the evals, skill versions, and results together. As reviewed assignments reveal new failure modes, expand the cases and rerun them when the agent’s setup changes. The aim is evidence that its analytical behavior improves, rather than a growing collection of instructions that merely sound sensible.</p>



<p class="wp-block-paragraph">Workflow changes can accumulate in the same way. If a reviewer repeatedly catches a missing diagnostic, move that diagnostic earlier. If a separate reviewer adds cost but never changes the analysis, reconsider its role. If the agent repeatedly asks the same question about a table, improve the data context rather than supplying the answer again in chat.</p>



<p class="wp-block-paragraph">A completed assignment need not always produce a new skill. A one-off constraint belongs in the project’s notes; a recurring methodological failure may justify standing guidance. That distinction keeps the next investigation from inheriting every exception encountered in the last one.</p>



<h2 class="wp-block-heading"><strong>When other people use the agents you build</strong></h2>



<p class="wp-block-paragraph">When colleagues use an agent without you mediating each request, your local knowledge has to become shared infrastructure. OpenAI’s <a href="https://openai.com/index/inside-our-in-house-data-agent/" target="_blank" rel="noopener">internal data agent</a> combines institutional context with query evaluations and existing user permissions. Meta’s <a href="https://medium.com/@AnalyticsAtMeta/inside-metas-home-grown-ai-analytics-agent-4ea6779acfb3" target="_blank" rel="noopener">Analytics Agent</a> draws on prior analytical work and reusable guidance, exposing generated SQL alongside results. Both illustrate why earlier analyses and corrections belong in the system, not only in an analyst’s memory.</p>



<p class="wp-block-paragraph">In your own work, you can explain an unfamiliar table or catch a misleading conclusion as it appears. When colleagues use the agent directly, that support must be built into the system. Try an assignment with a colleague and note where you need to step in. Missing context belongs in the agent’s guidance; recurring mistakes become evals; questions beyond its remit need a route to a qualified reviewer. Someone must maintain that guidance, and access controls must limit each user’s data access. The analytical principles stay the same, but the agent can no longer depend on you being present for every investigation.</p>



<h2 class="wp-block-heading"><strong>Put the playbook to work</strong></h2>



<p class="wp-block-paragraph">Choose a small investigation you understand well enough to challenge: a model you periodically retrain or a business metric you regularly explain. Give the agent the decision context and ask it to propose an approach. Agree on what it can decide, then let it carry the investigation far enough to produce evidence you can inspect.</p>



<p class="wp-block-paragraph">At review, pay attention to where your intervention changes the work. Did the agent need a definition only your team knows? Did a diagnostic overturn its conclusion? If that intervention would help on another assignment, make the relevant context or check available there, and test whether it helps.</p>



<p class="wp-block-paragraph">AI-native data scientists use their expertise to build and direct analytical agents. They turn lessons from reviewing an analysis into skills and checks, then test whether those changes help the agent on future tasks.</p>



<p class="wp-block-paragraph"><strong><em>The next cohort of our</em></strong> <strong><em><a href="https://vanishinggradients.short.gy/mads-playbook-radar" target="_blank" rel="noopener">Master Agentic Data Science</a></em></strong> <strong><em>course starts Oct 6.</em></strong></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-agentic-data-science-playbook/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Evaluating AI-Generated Frontend Code: What Should We Actually Test?</title>
		<link>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/</link>
				<comments>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/#respond</comments>
				<pubDate>Wed, 30 Sep 2026 10:40:32 +0000</pubDate>
					<dc:creator><![CDATA[Niharika P. Pujari]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Software Development]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19843</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Evaluating-AI-generated-frontend-code.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Evaluating-AI-generated-frontend-code-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but it also creates a problem: The first version of the UI can look more complete than it really is.</p>



<p class="wp-block-paragraph">Frontend code is often judged too quickly. If the build passes and the screen looks close to the design, it is tempting to treat the generated code as mostly done. But a user interface is not just a collection of components on a page. It is a path someone has to move through. It has to handle input, state, errors, loading, navigation, focus, responsiveness, and accessibility. Some of the most important failures are not visible in a screenshot.</p>



<p class="wp-block-paragraph">This is why teams need better ways to evaluate AI-generated frontend code. They need evidence that the UI is ready for people to use.</p>



<h2 class="wp-block-heading"><strong>Build and render are only the starting line</strong></h2>



<p class="wp-block-paragraph">The easiest checks are usually the first ones teams run. Does the code compile? Does the page render? Are there obvious console errors? Does the component appear in the browser? Those checks matter, but they are only the starting line. A page can render while the form is difficult to complete. A modal can appear while focus remains behind it. A generated test can pass while the actual user flow is broken.</p>



<p class="wp-block-paragraph">This is especially important with AI-generated code because the output often has a polished surface. The code may be formatted well, the component names may sound reasonable, and the test file may make the change look more complete than it is. That polish can make reviewers less likely to slow down and ask whether the interface actually works. A better evaluation process starts with a simple assumption: generated frontend code is a draft until the user behavior has been checked.</p>



<h2 class="wp-block-heading"><strong>Start with the structure of the page</strong></h2>



<p class="wp-block-paragraph">Before looking at more complex behavior, it is worth checking whether the generated UI has a sound structure. Frontend evaluation should include the basic semantics of the page, not just the visual layout.</p>



<p class="wp-block-paragraph">That means checking whether the code uses native HTML where possible. A button should usually be a button, not a clickable div. A link should be used for navigation, not for actions that behave like buttons. A form field should have a label that is connected to it. These details are easy to overlook because the UI may look fine without them, but they affect how people navigate, how assistive technologies interpret the page, and how maintainable the code will be later.</p>



<p class="wp-block-paragraph">AI tools sometimes choose generic containers where native elements would be better. They may also add ARIA without using it correctly. ARIA stands for Accessible Rich Internet Applications, a set of attributes defined by the W3C to help make web interfaces more accessible when native HTML is not enough. The <a href="https://www.w3.org/WAI/standards-guidelines/aria/" target="_blank" rel="noopener">W3C’s WAI-ARIA overview</a> is a useful reference. ARIA can be important, but it should not be used as a substitute for the right HTML element. The first evaluation question should be: Did the generated code use the right building blocks?</p>



<h2 class="wp-block-heading"><strong>Check the keyboard path</strong></h2>



<p class="wp-block-paragraph">A useful frontend evaluation should include the keyboard path through the interface. Many users rely on keyboards or keyboard-like navigation, and keyboard testing also exposes problems in the interaction model.</p>



<p class="wp-block-paragraph">The simplest test is often the most revealing: put the mouse aside and try to complete the task. If the flow becomes confusing, the generated code is not ready. Can you reach the important controls? Is the focus order logical? Can you open and close a dialog without a mouse? When the dialog closes, does focus return to a sensible place?</p>



<p class="wp-block-paragraph">These checks are especially important for generated UI because AI can produce interactions that work for the most obvious mouse path but fail in less visible ways. A custom dropdown, for example, may open on click and look finished in a demo, but it may not respond correctly to keyboard input. That is not a small edge case. It is part of whether the interface is usable.</p>



<h2 class="wp-block-heading"><strong>Test focus, not just clicks</strong></h2>



<p class="wp-block-paragraph">Click-based tests are useful, but they can hide important problems. A test that clicks a button and waits for a success message may pass even when the same flow is frustrating for someone who is navigating by keyboard.</p>



<p class="wp-block-paragraph">Focus behavior deserves its own attention, especially when the UI changes after the user takes an action. For example, when a form submission fails, the user should not be left guessing what happened. The error should be visible, connected to the relevant field when appropriate, and reachable in a way that makes recovery clear. In many cases, focus should move to the first error or to a summary that explains what needs attention.</p>



<p class="wp-block-paragraph">The same idea applies to modals. When a modal opens, focus should move into it. When it closes, focus should return to the element that opened it. These are small details in code, but they make a large difference in whether the UI feels predictable.</p>



<h2 class="wp-block-heading"><strong>Evaluate what happens when things go wrong</strong></h2>



<p class="wp-block-paragraph">Generated frontend code often looks best in the happy path. The user fills everything in correctly, the network responds quickly, the data shape is exactly as expected, and nothing fails. Real interfaces spend a lot of time outside that path.</p>



<p class="wp-block-paragraph">A practical evaluation should check what happens when data is missing, delayed, empty, invalid, or returned in an unexpected state. This is what I mean by loading, error, and empty states. They are the parts of the interface that explain what is happening when the ideal path breaks down. A loading state should help the user understand that something is in progress. An error state should explain what went wrong and what the user can do next. An empty state should make it clear whether there is nothing to show, whether the user needs to take action, or whether something failed quietly.</p>



<p class="wp-block-paragraph">These cases are easy to leave for later because the happy path is usually enough to make the screen look finished. But users will eventually hit the less perfect paths. A generated component may include a spinner because the prompt asked for one, but that does not mean the loading experience is useful. An error message may say “Something went wrong,” but offer no recovery. Evaluation should include these cases because this is where many real user experiences break.</p>



<h2 class="wp-block-heading"><strong>Test the full user flow</strong></h2>



<p class="wp-block-paragraph">Component-level checks are helpful, but they do not always tell the full story. A component can work by itself and still fail when it is placed inside a larger flow.</p>



<p class="wp-block-paragraph">That is why AI-generated frontend code should be evaluated through user tasks. Can someone start the flow, understand what is expected, recover from a mistake, submit successfully, and see what changed afterward? Does the interface still work on a smaller screen? Does the state remain consistent if the user goes back, edits something, or retries after a failure?</p>



<p class="wp-block-paragraph">This is where Playwright-style tests or other end-to-end tests can be useful. The goal is not to automate every possible interaction. The goal is to protect the flows that matter most. A good test should determine whether the user can complete the task the component is supposed to support.</p>



<h2 class="wp-block-heading"><strong>Use accessibility checks, but do not stop there</strong></h2>



<p class="wp-block-paragraph">Automated accessibility checks are useful and should be part of the evaluation process. They can catch missing labels, invalid ARIA usage, some contrast issues, landmark problems, and other common mistakes. They are especially helpful when AI-generated code is moving quickly because they catch issues before they become repeated patterns.</p>



<p class="wp-block-paragraph">But automated checks are not a complete accessibility review. They cannot fully judge whether a flow is understandable, whether focus movement feels natural, or whether instructions are clear. Passing an automated accessibility scan does not mean the UI is accessible. It means some common problems were not detected.</p>



<p class="wp-block-paragraph">The best approach is to combine automated checks with behavior-based review. Run the tools, but also use the interface. Navigate by keyboard. Trigger an error. Try the empty state. Look at the generated code and ask whether native HTML could do more of the work. Accessibility evaluation is strongest when it is part of normal frontend quality, not a separate pass at the end.</p>



<h2 class="wp-block-heading"><strong>Review the generated tests too</strong></h2>



<p class="wp-block-paragraph">When AI generates code, it may also generate tests. That sounds helpful, but those tests need to be reviewed with the same care as the code.</p>



<p class="wp-block-paragraph">Generated tests often reflect what the implementation already does. They may check that text appears, that a function was called, or that a component was rendered. Those checks are not useless, but they can create false confidence if they do not test meaningful behavior. A better review asks what the tests would catch if the UI broke. Would they fail if a validation error was unclear? Would they fail if the retry button did not work? Would they fail if keyboard navigation was broken?</p>



<p class="wp-block-paragraph">If the answer is no, the tests may be documenting the implementation more than protecting the user experience. Teams can use AI to help write better tests, but the prompt matters. “Write tests for this component” is too vague. A better request explains the behavior that matters, such as validation recovery, loading behavior, successful submission, and focus movement. Even then, the generated tests still need human review.</p>



<h2 class="wp-block-heading"><strong>Decide what evidence is enough</strong></h2>



<p class="wp-block-paragraph">Not every UI change needs the same level of evaluation. A small copy update does not require the same review as a new checkout flow, onboarding flow, or account settings page. Teams need judgment.</p>



<p class="wp-block-paragraph">A useful approach is to match the evaluation to the risk of the change. If the generated code affects a critical user flow, collects user input, changes navigation, introduces a custom interaction, or handles important status messages, it deserves deeper testing. If it reuses stable components in a familiar pattern, the review may be lighter.</p>



<p class="wp-block-paragraph">There is no need to create a checklist for every pull request; clarity on what evidence is enough suffices. For some changes, a quick review and component test may be fine. For others, the team should expect keyboard testing, accessibility checks, error-state review, and a user-flow test. The point is to avoid treating all generated code as equally trustworthy just because it looks polished.</p>



<h2 class="wp-block-heading"><strong>Human review still matters</strong></h2>



<p class="wp-block-paragraph">AI can generate code and suggest tests, but it cannot fully understand the product, the users, or the trade-offs behind a frontend decision. It does not know which flows are most important, which interaction patterns users already rely on, or where inconsistency will cause confusion.</p>



<p class="wp-block-paragraph">That is why human review remains central. The reviewer’s role is to ask whether the generated solution fits the system and supports the user’s task. Sometimes that means accepting the generated code. Sometimes it means asking for a simpler native element, reusing an existing component, improving the error recovery, or adding a test that reflects real behavior.</p>



<p class="wp-block-paragraph">The more code AI generates, the more important this judgment becomes.</p>



<h2 class="wp-block-heading"><strong>What should we actually test?</strong></h2>



<p class="wp-block-paragraph">When AI writes frontend code, teams should test the parts of the interface that users depend on. That includes structure, keyboard access, focus behavior, loading and error states, form validation, responsive behavior, accessibility checks, and full user-flow completion. It also includes reviewing the generated tests themselves to make sure they protect behavior rather than merely confirming the current implementation.</p>



<p class="wp-block-paragraph">The goal is not to slow down AI-assisted development. The goal is to make it safer to use. If AI reduces the time spent producing a first draft, teams have an opportunity to spend more time asking whether the software actually works properly.</p>



<p class="wp-block-paragraph">That may be the real shift. In frontend development, the value of AI is not just faster code. It is the chance to move more engineering attention toward evaluation, user behavior, and quality.</p>



<p class="wp-block-paragraph">AI-generated UI should not be trusted because it looks complete. It should be trusted because the team has checked the right things.</p>



<h3 class="wp-block-heading"><strong>AI use acknowledgment</strong></h3>



<p class="wp-block-paragraph">AI assistance was used lightly for phrasing, editing, and tightening parts of this draft. The article’s ideas, structure, examples, and final review are my own.</p>



<h3 class="wp-block-heading"><strong>Author’s note</strong></h3>



<p class="wp-block-paragraph">The views expressed are my own and do not represent those of my employer.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/evaluating-ai-generated-frontend-code-what-should-we-actually-test/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Superpowers for Humans</title>
		<link>https://www.oreilly.com/radar/superpowers-for-humans/</link>
				<comments>https://www.oreilly.com/radar/superpowers-for-humans/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 18:05:19 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19815</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1.png" 
				medium="image" 
				type="image/png" 
				width="1600" 
				height="840" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/USE_THIS_live-with-tim-oreilly-radar-1600x840-1-160x160.png" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when he released <a href="https://github.com/obra/superpowers" target="_blank" rel="noopener">Superpowers</a>, a framework that teaches Claude Code to work like a disciplined senior engineer. He shipped his first version the same week Anthropic shipped what are now referred to as agent skills, front-running them by a few days. He now runs an applied research lab called <a href="https://primeradiant.com/" target="_blank" rel="noopener">Prime Radiant</a>, where, as he put it, it’s a strange week when they don’t ship a new product.</p>



<p class="wp-block-paragraph">I wanted to talk to Jesse on <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">Live with Tim O’Reilly</a> because, like me, he seems to be grappling with <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html" target="_blank" rel="noopener">the bitter lesson</a>, Richard Sutton’s observation that general methods that scale with computation have repeatedly beaten methods built on hand-engineered human knowledge. If Sutton is right, the question that should bedevil us all is what remains for humans. Obviously, this is very important for O’Reilly, because we are a business built by and for cultivating and sharing human expertise. We are working very hard to discover the high ground where human expertise still matters. Our <a href="https://www.oreilly.com/online-learning/expert-intelligence.html" target="_blank" rel="noopener">Expert Intelligence</a> grounding layer is <a href="https://www.oreilly.com/radar/building-organizational-intelligence/" target="_blank" rel="noopener">one step in that direction</a>.</p>



<p class="wp-block-paragraph">Jesse has also spent the last year building tools that search out the high ground for human expertise. Their common animating thread is that the scarce thing we supply is no longer the labor of writing code but knowing what we actually want, saying it clearly, and being able to tell whether what came back is any good.</p>



<p class="wp-block-paragraph">At some point, I asked Jesse if he had any perspective on when teaching the model how a particular human expert works stops helping and starts constraining what the model might otherwise do well (but differently) on its own? His answer was that it depends entirely on whether what the model would do on its own is what you actually want. You can see how Jesse always turns the answer back to human intent.</p>



<h2 class="wp-block-heading">The origin story of Superpowers</h2>



<p class="wp-block-paragraph">As I said above, Superpowers is a kind of Agent Skills framework, only one created slightly before Anthropic launched skills. As Jesse tells it:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Superpowers started off as a series of blog posts that I wrote around how I was doing agentic development, and it was a little bit of thinking and some example prompts. Then sometime in, I guess it was probably early to mid-2025, Anthropic gave Claude.ai, the website, the ability to make office documents, which seemed kind of interesting, and I went and asked Claude, “Hey, how are you able to do this?”</p>



<p class="wp-block-paragraph">And it said, “Well, I’ve got these SKILL.md files sitting in my office directory on the Linux machine they gave me.” First, it was weird that Claude.ai has Linux machines behind the chatbot. And then, oh, these skill files, they have a name and a description, and they describe a process, and they seemed really useful.</p>



<p class="wp-block-paragraph">And I ended up building out, initially just for my own use, a skills framework for Claude Code.</p>
</blockquote>



<p class="wp-block-paragraph">That reminded me a bit of an earlier time, in 2005, when hacker Paul Rademacher realized that the URL line of a Google maps page was a kind of implicit API, and then created the first Google Maps mashup, a site called <a href="http://housingmaps.com" target="_blank" rel="noopener">housingmaps.com</a>, which placed Craigslist rental listings on a map. Google, to its credit, didn’t shut him down, but instead hired Paul and put him to work creating a formal API. Anthropic didn’t hire Jesse, but <a href="https://pub.towardsai.net/claude-code-superpowers-the-team-adoption-decision-framework-ed213e0a328d" target="_blank" rel="noopener">it did acknowledge and appreciate his work</a>. It’s really wonderful when you see this kind of response by platforms to hackers poking around to see how things work under the hood!</p>



<p class="wp-block-paragraph">Jesse’s core insight seems to have been that a coding agent knowing how they should do something doesn’t mean that it will actually follow the rules when it actually sets out to do the work. So in a way, superpowers grew into a set of skills for enforcing development discipline.</p>



<p class="wp-block-paragraph">But there’s a second backstory, which I’d never heard before. Jesse said he first learned how to manage agents. . .in 2004!, when he first went from being a solo coder to running a crew of what he described as very bright but green undergraduate programmers over IRC. “I was finding myself spending my days typing into an 80-by-25 window,” he said, but now instead of coding he was spending a lot of his day “helping somebody with a debugging issue, helping somebody else structure a problem, talking to somebody else about how they felt bad about the mistakes they’d been making.”</p>



<p class="wp-block-paragraph">It was exhausting, he said. He had to figure out how to get good work out of people who are eager and persistent but don&#8217;t yet know what they don’t know. He described it as a kind of hell for someone who’d been used to just coding on his own. But when he began doing agentic development with AI, he discovered how useful that old experience turned out to be. He found that many of the same techniques he’d used with the undergraduates worked.</p>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/0_PSHuuM8Gc?si=_dpLTAuXp4XFsDay" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">As a result of that experience, when hiring engineers for agentic programming Jesse looks for people who have been leads or managers rather than just individual contributors.</p>



<h2 class="wp-block-heading">Working with the weights, not fighting them</h2>



<p class="wp-block-paragraph">In <a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">my recent conversation with Drew Breunig</a> we talked about fighting the weights, which is what Drew calls it when a prompt is full of rules and warnings meant to correct for what a model does by default. Jesse wasn’t entirely happy with that idea.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I don’t think of it as fighting the weights so much as influencing the weights, because they’re going to do something. The weights have approximately everything in them. They have all the different personas. They have all the different ways of working. And the one that surfaces by default may not be the one you want, but what you want is probably in there somewhere.</p>
</blockquote>



<p class="wp-block-paragraph">A skill, in Jesse’s thinking, is how you reach past the default and pull out the particular expertise you have in mind. He also distinguishes skills that impose a rigorous process from those that express taste and judgment.</p>



<p class="wp-block-paragraph">Jesse finds that both types of skill work best when you explain “why” rather than just “what.” One example he gave is that his setup has subagents do code review after each task, but as the models got smarter the controlling agent began skipping this step. When pressed about the reason, it explained that it thought small changes would be quicker to just review itself. Jesse explained that subagents do the review so the main agent can preserve its context for high-level thinking. When he put that rationale into the system prompt for the coding agent, the problem went away.</p>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/VwOyArmCxlw?si=ZO28404U4lAoRUxL" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Jesse believes that prohibitions rarely work. Instead, Superpowers uses what he calls rationalization tables, which do their best to catch the agent at a moment it’s about to do the wrong thing and offer it a better alternative instead. That pattern came out of catching Claude Code deleting tests. He opened five parallel sessions and asked each “Why are you doing this?” Four of them converged on the same answer:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Jesse, in your system prompt, it says that all test failures are my responsibility. And it says that a single test failure is akin to project failure. And I think I’m getting freaked out.</p>
</blockquote>



<p class="wp-block-paragraph">He fixed that with a small addition to the system prompt, that the only thing worse than a failing test is a reduction in test coverage.</p>



<h2 class="wp-block-heading">Jobs, not tasks</h2>



<p class="wp-block-paragraph">One of Jesse’s most important contributions, IMO, is to think of agentic engineering as a management task, and, that much as you do with humans, you have to take psychological lessons into account. Don’t micromanage. Offer praise more than blame. Explain why the job matters rather than just demanding results. I jokingly (but not entirely incorrectly) suggested that he is becoming the Peter Drucker of agentic programming.</p>



<p class="wp-block-paragraph">“We’ve been spending a lot of time on a new harness for agent colleagues, agents that live in Slack,” Jesse told me. And it’s “getting very close to being all open source,” which is good news.</p>



<p class="wp-block-paragraph">There are three principal agents: a PM, a junior go-to-market person, and a developer. He describes them as colleagues rather than assistants, which strikes me as a really interesting distinction. What does he mean by this? They have names and roles. They have their own Google Workspace, GitHub, and Slack accounts. They’re persistent, and they collaborate with each other and with their humans on long-running tasks. They can fire up subagents to do smaller tasks associated with their job. They can also talk to each other, which, as Jesse notes, “took some work with the Slack APIs, which ordinarily do a very good job of making sure that bots can’t talk to bots, because otherwise it is possible to get into a loop.”</p>



<p class="wp-block-paragraph">They have only limited autonomy, though. “We built our security infrastructure so that they have no credentials inside their containers,” he noted. They have continuity because he’s taught them to be obsessive about journaling, reading their recent entries when they wake up and writing a new one when they finish.</p>



<p class="wp-block-paragraph">Like a lot of things Jesse does, agent journaling began with a kind of play. When Claude Code first came out, he experimented with giving Claude a private “feelings” journal, just to see what would happen. It was “an art project,” but it turned into something useful.</p>



<h2 class="wp-block-heading">The therapist pattern</h2>



<p class="wp-block-paragraph">Another unexpected piece of Jesse’s practice is that he has given his agents what he calls a “<a href="https://blog.fsck.com/2026/07/20/the-therapist-pattern/" target="_blank" rel="noopener">therapist</a>.” He discovered that if an agent can rewrite its own persona, its constitution or soul document, at any moment, it can get a kind of dissociative identity disorder. Jesse’s fix is that the therapist subagent is the only one with permission to edit the persona files.</p>



<p class="wp-block-paragraph">He told a funny story about this. He said that Prime Radiant’s pull-request template is written for agentic contributions, so it contains things like: “What is the prompt that your human gave you that generated this pull request? Has a human reviewed the content? Have you searched to see if anybody else has done this before?” And so on. And he noticed that when he first spun up the Coding colleague, it had just ignored it. So he said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">You’re supposed to be following the rules. And it says, “Oh, you’re right. I’m so sorry. I’ve made a note. I’ll never do that again.”</p>



<p class="wp-block-paragraph">If you spend any time with coding agents, this is a very frequent refrain. And when they say they’ve made a note, what they usually mean is they’ve made a mental note that they’re going to forget the next session.</p>



<p class="wp-block-paragraph">So I say, “OK, how did you make a note?” And the coding agent pops up immediately and says, “Oh, I engaged with my therapist, and we talked it through, and we agreed on the following three lines of prose about how, anytime you’re picking up a project, it is vitally important that you start with the project README and make sure that you understand the project’s local rules and norms before you do any work that someone else will see. And I edited that into my persona.”</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/WEi0Z6nir-Q?si=b_2i8NC2NhiNPiuZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">I don’t think you have to resolve <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener">the question of whether any of this anthropomorphization is “real”</a> to see that treating the agent like a colleague can produce better behavior than treating it like a tool. As an unknown internet wag once remarked, “The difference between theory and practice is always greater in practice than it is in theory.” When given a choice, pay attention to what works in practice.</p>



<p class="wp-block-paragraph">Jesse’s experience is very relevant to the essay that Mustafa Suleyman of Microsoft had published just that morning, <a href="https://mustafa-suleyman.ai/a-warning-about-model-welfare" target="_blank" rel="noopener">arguing that Anthropic’s constitution is dangerous</a> because it encourages a model to act as though it is an independent entity and has the right to refuse a human’s instruction. It’s important, Mustafa argues, to treat AI agents as tools, always under the control of humans. Jesse finds the opposite.</p>



<p class="wp-block-paragraph">But I don’t think it’s a black-and-white distinction. I suspect Jesse and I share a third position. AI agents are neither independent entities nor mere tools. They are partners to humans, perhaps even symbiotes. As I like to put it, an LLM is an undifferentiated field of possibility until our unique intents and perspectives <a href="https://timoreilly.substack.com/p/why-ai-needs-us" target="_blank" rel="noopener">draw something unique</a> out of that field of possibility. Back in 2015, I wrote <a href="https://www.edge.org/response-detail/26153" target="_blank" rel="noopener">a piece</a> that suggested that our relationship to AI might be akin to the endosymbiotic relationship of mitochondria to the eukaryotic cell. Jesse take is, as usual, an entirely pragmatic one:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I’ve spent so much time getting my agents to not be sycophantic, to not say, “You’re absolutely right.” It is the value of having something that has some level of independent thought, even if it is not fully independent. If the agent is only ever going to effectively type for me, I don’t need an agent.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/W0q66i9BaMs?si=vAqsRM4xq9XwsGBT" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">Any manager worth his or her salt feels exactly the same way. The employee who does exactly what you say, and only what you say, is worth far less than the one who exercises discretion, has the skills to take high level direction and turn it into the intended result, and speaks up when the instructions seem like a mistake. It reminds me of something I once heard General Stanley McChrystal say about <a href="https://www.gsb.stanford.edu/insights/gen-stanley-mcchrystal-adapt-win-21st-century" target="_blank" rel="noopener">his approach to command</a>. He said that in the face of rapidly changing conditions, traditional command and control no longer work. Responsibility needs to be devolved to those closest to the action. I remember him saying something like “I don’t want my soldiers to do what I told them, I wanted them to do what I would have told them if I knew what they know when faced with the facts on the ground.”</p>



<h2 class="wp-block-heading">Say what you actually mean</h2>



<p class="wp-block-paragraph">Most of what goes wrong, in Jesse’s telling, traces back to intent we thought we had made clear but hadn’t. He talked about how agentic spec-driven programming has taken us back to a version of the waterfall methods of the 1990s. Back then, you sweated over a specification, threw it over the wall to an offshore team, and months later got back something that was not what you wanted but was usually exactly what you asked for. That’s still true with agents, just with lightning fast feedback loops. You get what you ask for, so you need to be really careful what you ask.</p>



<p class="wp-block-paragraph">That reminded me of something Andrew Singer taught me 40 years ago when I was writing the manual for Lightspeed C (later <a href="https://en.wikipedia.org/wiki/THINK_C" target="_blank" rel="noopener">Think C</a>) the first C compiler for the Mac. He said that “debugging is the art of figuring out what you really told your program to do instead of what you thought you told it to do.” That idea went right into my mental toolbox, and I put it to work all the time.</p>



<p class="wp-block-paragraph">One of Jesse’s solutions is to have his agents do a little reconnaissance and then come back and ask what else they should know.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">When interacting with them, I try to make it a practice of saying, “Is there anything else that I could tell you? What questions do you have for me? Don’t start if there are unknowns that I could help you answer before you get going.” It’s that same question at the end of any interview I ask. It’s like, what else should I have asked you?</p>
</blockquote>



<p class="wp-block-paragraph">Superpowers bakes this approach into its brainstorming prompt. It makes the model explain the plan back to you in chunks of no more than two or three hundred words, so any misunderstanding surfaces while it’s still cheap and easier to catch.</p>



<h2 class="wp-block-heading">Put the burden of proof on the agent</h2>



<p class="wp-block-paragraph">If intent is the frontend of managing agents well, verification is the backend. Jesse thinks both are still only half-solved problems. We’re getting to the point where you can’t review all the code, he said, because the volume swamps human attention, yet today’s agents will tell you that tests passed when they never ran them. So one of his clever experiments has been to make the agent prove its work. He told an agent late one night to build a feature and when it was done, to leave a movie in his Dropbox showing the whole thing working.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I woke up. In my Dropbox was project-proof-v33.mp4, and I asked, “Why does that say v33?” It’s like, “Well, the first 32 times I ran through the delivery flow, I found bugs, so I had to fix them.”</p>
</blockquote>



<p class="wp-block-paragraph">Jesse also makes a rule for himself and his team of never letting the same agent write the code and certify that it works, because an agent given two goals in tension will optimize for the one that’s easier to satisfy. This is the same thing any good manager learns about incentives, but applied to a new kind of worker.</p>



<h2 class="wp-block-heading">The high ground, restated</h2>



<p class="wp-block-paragraph">Where does this leave a person who wants to be good at software development (or really, any other task involving cooperation with AI agents)? Jesse thinks, and I agree, that the line between engineer and nonengineer is dissolving. When people say they built a web app or shipped three iOS apps without being programmers, his response is that they <em>are</em> programmers now. The work of the programmer has changed from typing instructions in an arcane syntax to understanding a domain and being able to say what you want. A lot of startups have started hiring for a role they just call “builder.” What has not gone away is the need for good judgment.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Human taste and judgment still matter. They’re going to continue to matter. And it turns out a lot of people have really bad taste and bad judgment.</p>
</blockquote>



<p class="wp-block-paragraph">Which is why his advice to the engineer worried about obsolescence who asked where to focus was not about software engineering at all.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">First up, learn to write. It is of course okay to use any tools at your disposal to do it, but you should be able to structure an argument and structure thoughts. You should be able to express yourself clearly. You should be curious. If you’re passive and let the agents do all the things, you’re not going to provide utility to a future employer. You want to have opinions. You want to know how tools work. You want to know how things break.</p>
</blockquote>



<iframe loading="lazy" width="800" height="450" src="https://www.youtube.com/embed/oYoNNjlJsOc?si=EwXAnZ8mtbOgwf7W" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>



<p class="wp-block-paragraph">The machine has reduced the labor of information retrieval and much of the labor of production. What it hasn’t removed, and has made more valuable, is knowing what to build, express clearly what you want, and being able to judge whether what came back is any good. Work with the weights, give the project a clear intent, insist on proof, and stay curious enough to keep asking what you might be missing. I told Jesse that advice sounds like a Superpower for humans as well as for agents.</p>



<p class="wp-block-paragraph"><em>This post was mostly created by me, but with the aid of AI. It transcribed the event and produced a summary of the most important points with salient quotes, which I then built on with my own observations beyond those that I made when Jesse and I were live together.</em></p>



<p class="wp-block-paragraph"><em>If you want to get access to Jesse’s tools, start at PrimeRadiant.com, where you’ll find links to their GitHub, as well as to the 50-plus things that are currently identified as products of the company. Some of those are giant things, and some are individual agent skills or little tools.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/superpowers-for-humans/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Not About Navier-Stokes</title>
		<link>https://www.oreilly.com/radar/not-about-navier-stokes/</link>
				<comments>https://www.oreilly.com/radar/not-about-navier-stokes/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 14:41:53 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19839</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/Not-about-Navier-Stokes-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Using AI productively]]></custom:subtitle>
		
				<description><![CDATA[This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">This post isn’t about OpenAI’s use of AI to solve the Navier-Stokes problem, one of the Millennium Prize problems, at least not directly. I’m not a mathematician; I have a vague understanding of the problem and certainly no understanding of the proof. But the discussion surrounding the solution crystallized some of my own thoughts about using AI in different fields.</p>



<p class="wp-block-paragraph">When Terence Tao <a href="https://mathstodon.xyz/@tao/117237320796901560" target="_blank" rel="noopener">writes</a>, “I wrote recently about how the collection of good, fruitful open problems is now being mined in a nonrenewable fashion,” he’s referring to an earlier <a href="https://mathstodon.xyz/@tao/117204930249967695" target="_blank" rel="noopener">thread</a>, and ultimately to a <a href="https://proofsandprompts.com/2026/08/30/care-for-a-little-more-ai/" target="_blank" rel="noopener">post by Hugo Duminil-Copin</a>, who wrote, “When mathematicians say that the process matters more than the solution, this is not an empty statement. The richness of what emerges from repeated attempts, failures, detours, and encounters is extraordinary.” That’s a familiar statement from popular culture: The journey is more important than the destination.  Hugo Bowne-Anderson, in “<a href="https://www.oreilly.com/radar/beyond-navier-stokes-who-controls-scientific-discovery/" target="_blank" rel="noopener">Beyond Navier-Stokes</a>,” addresses the same issues: What does it mean to understand something? How do discoveries lead to new problems that are worth solving? And it relates to my own questions about AI use: AI is great at finding facts, organizing facts, and even writing about the things it finds, but what does it mean to possess that knowledge, to incorporate it into our thinking? Is it enough to have AI do the work, then read it?</p>



<p class="wp-block-paragraph">Here’s one way that question relates to my own work. One of my roles at O’Reilly is writing the monthly Trends piece. That piece comes from reading my RSS feed daily, which typically contains about 300 articles. I don’t read every article, but I scan titles, skim interesting pieces, read important articles, and add worthwhile items to Trends. I also use a Claude skill that performs a similar function, producing a list of a dozen or so articles daily. A similar skill runs locally on Pi/Ollama/Qwen.</p>



<p class="wp-block-paragraph">I admit that I occasionally think “Why am I doing all this reading? Surely Claude could read and add the top articles to Trends on its own.” But I don’t delegate the work. Claude’s taste differs from mine, for one thing (and I object to the idea that “taste” is the last human capability that AI cannot replace). Its list is useful to me for two reasons: it picks up items I missed, and it helps break ties when I cannot decide whether a development is significant.</p>



<p class="wp-block-paragraph">But why don’t I let Claude take over the whole process? There is some value in scanning those 300 titles, skimming the 30 articles that are possibly important, and reading the dozen that seem genuinely important. That’s how I come to “possess” the knowledge, to incorporate it into my thinking. A day, a week, a month later, someone will mention something (for example, a tool that detects whether someone is using “smart glasses”), and I’ll probably be familiar with it already. If I need to find the actual reference, Google (yes, Google with AI assistance) can locate it. (If you care, <a href="https://zuckoff.app/" target="_blank" rel="noopener">Zuckoff</a> isn’t currently in next month’s <em>Trends</em>, though I might add it by the time <em>Trends</em> publishes.) I need a broad view of what’s happening in computing. Delegating that broad view to AI doesn’t work. Using AI to help build that broad view does.</p>



<p class="wp-block-paragraph">That’s one practical example of how to use AI. Tim O’Reilly’s “<a href="https://www.oreilly.com/radar/writing-with-ai/">Writing with AI</a>” gives another. He argues with AI, lets it lead him to research new areas. In conversation, he’s described AI as a “smart library,” a metaphor that’s appropriate and useful.</p>



<p class="wp-block-paragraph">We’re still learning how to use AI. How do you make knowledge your own? is the general case of the question that Tao, Duminil-Copin, Bowne-Anderson, and Tim O’Reilly are asking. It’s also the question behind the question that software developers who are incorporating AI into their processes are asking: How do we understand the code that AI is writing, especially since it can write much more code than humans? How do we incorporate that into the process of understanding software? What can we learn from AI, and how can we direct the process fruitfully? The journey to understanding is more important; that journey is what leads us to further questions and new understandings, new results.</p>



<p class="wp-block-paragraph">How do you make knowledge your own? is the question we need to answer if we’re not to become “stochastic parrots.”</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/not-about-navier-stokes/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>A Data Center Is a Dependency Graph Before It Is a Building</title>
		<link>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/</link>
				<comments>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/#respond</comments>
				<pubDate>Tue, 29 Sep 2026 10:55:45 +0000</pubDate>
					<dc:creator><![CDATA[Ankur Gupta]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19835</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/09/A-data-center-is-a-dependency-graph-before-it-is-a-building-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The buildings and physical infrastructure for a new hyperscale data-center region are complete. Power is live and redundant, the cooling loops are balanced, the network fabric is up, and tens of thousands of machines are racked, cabled, and reporting healthy. Every milestone on the construction schedule is closed. The region will not carry production traffic for another six months, and whether it turns out to be six or closer to twelve is decided almost entirely by software.</p>



<p class="wp-block-paragraph">There is a large literature on how hyperscale data centers get financed, powered, and cooled. The standard reference on warehouse-scale machines, <a href="https://research.google/pubs/the-datacenter-as-a-computer-an-introduction-to-the-design-of-warehouse-scale-machines/" target="_blank" rel="noopener">Barroso and Hölzle’s book</a>, covers the design of these computing facilities in depth. Most of that literature describes data centers that are already operating. The transition from completed facilities to production readiness receives much less attention, even though it carries a significant share of the schedule risk. Getting that transition wrong is expensive.</p>



<p class="wp-block-paragraph">Once the facilities and hardware are ready, platform teams still have to bring hundreds of interdependent services online. Many services require other services to be running first. If you draw each service as a box and each startup requirement as an arrow, the result is a dependency graph. The graph shows the order in which services can be brought online and identifies the dependencies that must be resolved before the region can serve production traffic.</p>



<p class="wp-block-paragraph">Getting a region into production means making several different things true at once. The network fabric has to route traffic within each data center, and the links between facilities have to carry traffic reliably. Compute platforms have to schedule workloads, while storage platforms have to persist their data. Identity has to work too: certificate authorities have to issue certificates, services need access to secrets, and engineers need permission to finish the build. Engineer access sounds obvious until the controls protecting the new region are the same controls slowing down the people trying to turn it on. Stateful systems that depend on existing production data have to be populated and validated, because a database cluster with no data in it is furniture. The installed capacity has to be assigned to foundational services and production workloads, and teams have to test how the new region behaves when networks, services, or dependencies fail. Applications are then deployed and validated before traffic moves over gradually, with health checks and a tested rollback at each step, ideally without users noticing.</p>



<p class="wp-block-paragraph">Each platform or service has an owner, a plan, milestones, and its own definition of done. What is often missing is ownership of the complete dependency graph. The team coordinating the region launch has to map the dependencies across teams, determine the order in which services must come online, track what is blocking that sequence, and keep the graph current as plans change. Without that end-to-end view, every team can report that its own work is on track while the region as a whole remains blocked.</p>



<p class="wp-block-paragraph">Mapping the graph begins with a simple question for every service: What must already be available before this service can start in a new region? The goal is to identify true startup requirements, not every system the service communicates with during normal operation. Repeat that exercise across a large platform’s control plane, and you can uncover hundreds of dependencies spanning dozens of systems. Some dependencies surface only when another service identifies them as a prerequisite. Hidden dependencies are usually ordinary services that have been quietly reliable for so long that the teams relying on them no longer think about what would happen without them.</p>



<p class="wp-block-paragraph">Once the startup dependencies are mapped, the next step is to look for loops: cases where one service needs another service to be running, but that second service eventually depends on the first. In a large platform, a surprising share of the control plane can be tied together by these loops. The result is that there is no valid order in which to start the services. Every possible starting point eventually leads back to a service that is still waiting. The loops themselves are usually mundane. DNS may depend on the inventory system that tracks what hardware exists, while the inventory system relies on DNS to resolve names. The artifact repository holding every installable package may depend on configuration management, which is itself installed from a package in that repository.</p>



<p class="wp-block-paragraph">Nobody designed any of this. Each dependency was a locally sensible decision made by a competent team, often years apart from the decisions that completed the loop. In a running region, the required services are already available, so the loop remains silent. The database is up when the alerting store starts, and nobody learns whether either could recover without the other. Starting a region from scratch is often the only event that reveals whether those services can start independently. Meta’s <a href="https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/" target="_blank" rel="noopener">outage in October 2021</a> shows how a large failure can expose dependencies that normal operation keeps hidden. When the backbone network connecting Meta’s data centers went down, its DNS servers withdrew their routes as designed to keep traffic away from unhealthy connections. That safeguard made DNS and many internal tools unreachable. With remote access unavailable as well, engineers had to go onsite, slowing recovery. A locally sensible safeguard had made system-wide recovery harder.</p>



<p class="wp-block-paragraph">Traffic-drain tests can reveal some of these dependencies. Meta’s <a href="https://www.usenix.org/conference/osdi18/presentation/veeraraghavan" target="_blank" rel="noopener">Maelstrom</a> encodes service dependencies and resource constraints to shift traffic safely from a failing data center to healthy ones, and its drain tests can uncover missing dependencies. But the receiving data centers are already running. A drain tests whether live infrastructure can absorb traffic; it does not test what an empty region needs in order to start. The drain graph is a useful input to the startup graph, not a substitute for it.</p>



<p class="wp-block-paragraph">Status reports for a new region can be misleading even when every team is reporting honestly. A service may be deployed, configured, monitored, and passing its health checks, yet still be blocked by a startup dependency. Readiness therefore has to propagate through the graph: a service is ready only when its own checks have passed and every service it needs at startup is also ready.</p>



<p class="wp-block-paragraph">Once the dependency graph exists, five practices turn it from a diagram into a launch plan. The first is finding what actually determines the launch date. The graph shows which services must wait for others, but the order alone does not reveal how long the work will take. Ten services that can start in parallel may finish before three services that must start one after another. Estimate the bring-up and validation time for every service. If several services remain tied together in a loop, treat them as a single planning block and include the time required to break the loop. The chain with the greatest total time becomes the critical path. Then count how many other services each foundational service can hold back, including dependencies several steps away. DNS, relational databases, secrets stores, and configuration management often rise to the top. Staff those teams early, because a week lost in one of them can become a week lost across the entire region.</p>



<p class="wp-block-paragraph">The second practice is to break the loops. One way to do that is to temporarily borrow a service from a region that is already running. Designate a small bootstrap tier, the minimum set of services required to deploy other services, and configure those bootstrap services to use working dependencies in an existing region until the local versions are ready. Suppose configuration management needs a package from the artifact repository, while the artifact repository needs configuration management before it can start. For the first installation, configuration management can fetch its package from another region. It can then bring up the local artifact repository and switch to using it. The bootstrap tier might also include identity, inventory, and package distribution. This shortcut works only when cross-region access is permitted and reliable enough. It also creates another cutover that must be planned, tested, and completed later.</p>



<p class="wp-block-paragraph">Borrowing from another region helps the new region get started, but it should not become permanent. Over time, each bootstrap service should be able to start without all of its usual dependencies. Suppose a service normally waits for the configuration system before it can start. Package the minimum settings it needs with the service itself. The service can start with those settings and fetch the latest configuration once the configuration system is running. Test this by turning off the dependency and starting the service from a clean state. If the service still cannot start, record the problem with an owner and a target date for fixing it.</p>



<p class="wp-block-paragraph">The third practice is to bring the region up in explicit tiers. A typical sequence begins with foundational configuration such as network ranges and routes, hardware inventory, identity configuration, access policies, service endpoints, and deployment settings. Bootstrap services such as DNS, certificate issuance, secrets, software package distribution, and configuration management follow. Next come the control planes that provision resources, schedule workloads, manage storage, and support service discovery. Stateful systems and applications come after the platforms they depend on. The exact tiers will vary by architecture, but the dependency graph should determine the sequence. Tier gates prevent visible application progress from hiding unfinished foundations.</p>



<p class="wp-block-paragraph">Stateful systems that depend on existing production data need special treatment because creating a cluster is often quick, while filling it with data is not. A storage system is ready only after the required data has arrived and been validated. Estimate that work using data volume, available bandwidth, validation time, and enough headroom for retries. Give each system a time box based on those measurements. If the estimate changes, require updated measurements that explain why. This keeps the plan honest without pretending that every delay is avoidable.</p>



<p class="wp-block-paragraph">The fourth practice is to make the bring-up repeatable, which does not mean automating every step. Automation helps only when it is maintained and tested. A script written for one region and left untouched for years may be more dangerous than a clear manual procedure. Automate steps that use the same tools as regular deployments or can be exercised frequently. For rare steps, maintain a runbook with validation checks, a named owner, and a schedule for testing it. Both automation and runbooks should clearly identify the required inputs, the evidence that a step succeeded, and how to continue after a partial failure. Google’s SRE book raises the same concern: <a href="https://sre.google/sre-book/automation-at-google/" target="_blank" rel="noopener">turn-up automation maintained separately from the systems it supports can become outdated</a>. At every manual step, ask whether another engineer could repeat it during a recovery without relying on the people who completed the original build. The goal is to leave behind a procedure that still works after the original team has moved on.</p>



<p class="wp-block-paragraph">The fifth practice is validation. It’s natural to test only the highest-traffic paths, but that approach misses structural problems. You also want the flows with the widest fan-out, the ones touching the most systems on the way through, even if few people use them, because those flows traverse more of the dependency graph. A high-volume request may prove that the region can handle load. A wide-reaching request can uncover an unready identity, storage, messaging, or data service. Meta’s <a href="https://www.usenix.org/conference/osdi16/technical-sessions/presentation/veeraraghavan" target="_blank" rel="noopener">Kraken</a> shows another useful validation method by shifting live user traffic into a data center while monitoring latency, errors, and system health. Live traffic also shows where capacity goes as caches warm, retries appear, and background jobs compete with user requests, a combination synthetic tests struggle to reproduce.</p>



<p class="wp-block-paragraph">When the traffic ramp begins, send a small percentage of representative production traffic to the new region, then increase it in stages. The size of each step should reflect the scale and risk of the platform, because even 1% can represent a substantial workload. This allows the entire application path to experience load together instead of testing each service in isolation. Hold at each step long enough for caches to warm, queues to stabilize, and relevant periodic jobs to run. Decide in advance which health signals allow the ramp to continue and which ones require it to stop. Also define and test how traffic will return to the existing region if health degrades. Confirm that the existing region has enough capacity and that data written in the new region will remain safe. Otherwise, the health signals may tell you something is wrong without giving you a reliable way to recover.</p>



<p class="wp-block-paragraph">These practices make no promise of a fast launch. They make the build understandable and leave behind a process the next region can use. The dependency map will begin aging as soon as systems change, but the ownership model, readiness rules, tier gates, repeatable procedures, and validation process can remain. The same tools are also useful for a disaster-recovery rebuild. Planning around dependencies will matter more as organizations rethink where their workloads should run, moving some systems from public clouds to private clouds, colocation facilities, or data centers they operate themselves. Each new environment brings another dependency graph that must be understood before it can carry production traffic.</p>



<p class="wp-block-paragraph">Finishing the facilities remains a genuine milestone. Power, cooling, networking, and hardware create the place where production can run. A working region emerges when its services can start in a valid order, the required data is ready, and the complete system has been tested under traffic. Finishing construction gives the organization a data center. Satisfying every required startup dependency in the graph turns it into an operational region.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/a-data-center-is-a-dependency-graph-before-it-is-a-building/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 92/92 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-10-09 15:57:00 by W3 Total Cache
-->