<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="https://purl.org/rss/1.0/modules/content/"
	xmlns:media="https://search.yahoo.com/mrss/"
	xmlns:wfw="https://wellformedweb.org/CommentAPI/"
	xmlns:dc="https://purl.org/dc/elements/1.1/"
	xmlns:atom="https://www.w3.org/2005/Atom"
	xmlns:sy="https://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="https://purl.org/rss/1.0/modules/slash/"
	xmlns:custom="https://www.oreilly.com/rss/custom"

	>

<channel>
	<title>Radar</title>
	<atom:link href="https://www.oreilly.com/radar/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.oreilly.com/radar</link>
	<description>Now, next, and beyond: Tracking need-to-know trends at the intersection of business and technology</description>
	<lastBuildDate>Thu, 27 Aug 2026 18:23:56 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://www.oreilly.com/radar/wp-content/uploads/sites/3/2025/04/cropped-favicon_512x512-160x160.png</url>
	<title>Radar</title>
	<link>https://www.oreilly.com/radar</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Writing with AI</title>
		<link>https://www.oreilly.com/radar/writing-with-ai/</link>
				<comments>https://www.oreilly.com/radar/writing-with-ai/#respond</comments>
				<pubDate>Thu, 27 Aug 2026 15:54:34 +0000</pubDate>
					<dc:creator><![CDATA[Tim O’Reilly]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19504</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Writing-with-AI.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Writing-with-AI-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[A contrarian view]]></custom:subtitle>
		
				<description><![CDATA[A lot of professional writers say they never do it. And for writers who do, the consequences can be serious. Book contracts have been withdrawn. People have lost jobs. I have a contrarian view. I say AI is a medium, like painting, sculpture, photography, music, or language itself. If AI is a medium, it is [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">A lot of professional writers say they never do it. And for writers who do, the consequences can be serious. Book contracts have been withdrawn. People have lost jobs.</p>



<p class="wp-block-paragraph">I have a contrarian view. I say AI is a medium, like painting, sculpture, photography, music, or language itself. If AI is a medium, it is a field of possibility from which humans can and will summon up and communicate insights, truth, beauty.</p>



<p class="wp-block-paragraph">Last year I asked Claude to write an essay from its own point of view about a conversation we’d been having. I published it as “<a href="https://timoreilly.substack.com/p/why-ai-needs-us" target="_blank" rel="noopener">Why AI Needs Us</a>.” While I named Claude up front as the author, in the afterword I disclaimed that, saying “I am the author of this essay, though Claude wrote every word.” I pulled it out of the latent space of possibility by what I asked for and how I asked, even though Claude came back with turns of phrase I’d never have used on my own. In the course of our hours-long conversation, we made something together. That wasn’t the machine replacing me; it was me exploring a new instrument.</p>



<p class="wp-block-paragraph">If AI is a medium, it will take time for its possibilities to unfold, discovery on discovery. In Western painting, think of the progression from the flat, stylized religious art of the Middle Ages to the luminous colors of Giotto, the explosion of technique and realism in the Renaissance, and the many stages since. Greatness in the medium has many faces, not one. Botticelli, Michelangelo, Leonardo are all names to conjure with, but so are the very different visions expressed centuries later by Monet, Van Gogh, Picasso, O’Keefe, Kahlo, Frankenthaler, Haring, Banksy.</p>



<p class="wp-block-paragraph">If AI is a medium, the idea that you can’t write with it will someday seem as odd as the idea that you can’t make a good portrait or landscape with a camera, only with paint or pen and ink. Technical innovations have often changed the nature of art. Researchers have shown that <a href="https://lsa.umich.edu/psych/news-events/all-news/faculty-news/most-people-do-not-realize-when-a-personal-message-they-receive-.html" target="_blank" rel="noopener">people feel cheated</a> when they receive a personal message such as an apology that was written with AI. Automating a supposedly personal message is indeed a kind of deception. But we don’t think that Leonardo was cheating when he used a camera obscura to perfect perspective in painting, or that John Lasseter was getting away with something because he created <em>Toy Story</em> with a computer instead of painstakingly animating it by hand.</p>



<p class="wp-block-paragraph">Ted Chiang wrote an essay called “<a href="https://www.newyorker.com/culture/the-weekend-essay/why-ai-isnt-going-to-make-art" target="_blank" rel="noopener">Why A.I. Isn’t Going to Make Art</a>,” which talks about both art and writing. He observed:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Art is notoriously hard to define, and so are the differences between good art and bad art. But let me offer a generalization: art is something that results from making a lot of choices. This might be easiest to explain if we use fiction writing as an example. When you are writing fiction, you are—consciously or unconsciously—making a choice about almost every word you type; to oversimplify, we can imagine that a ten-thousand-word short story requires something on the order of ten thousand choices. When you give a generative-A.I. program a prompt, you are making very few choices; if you supply a hundred-word prompt, you have made on the order of a hundred choices.</p>
</blockquote>



<p class="wp-block-paragraph">I agree with Ted about good art or writing as the result of many choices, but I can’t agree with his conclusion that AI can’t be used for writing, art, or other creative work because it lets you get away with minimal effort. In my opinion, this is what Frank Herbert once described to me as a “Jesuitical argument,” which he defined as “one you can always win as long as you control the givens.” Yes, AI isn’t going to create art on its own, any more than a camera or a set of oil paints will. Yes, art requires effort and expertise. Yes, AI won’t create art from a hundred-word prompt. But nonetheless I believe humans will create art with it.</p>



<p class="wp-block-paragraph">Even the “hundred-word prompt” is a bit misleading. The choices expressed by a prompt are not limited to the number of words in it. A lifetime of reading, thinking, conversing, formulating might have led to that moment of expression. Leonard Cohen confessed to Bob Dylan that he had reworked “Hallelujah” over and over for two years (later admitting it was actually five), while Dylan claimed to have written “Blowing in the Wind” in 10 minutes. Even though Cohen appears to have labored far more over his work, and Dylan’s song spilled out like magic, both are masterpieces. The choices were made not just in writing the words but in the inner life of the author. So too, the prompt that one person brings to an AI may contain multitudes, while another makes only a humdrum provocation.</p>



<p class="wp-block-paragraph">Photography gives a less anecdotal counter-example. A painter decides on every brush stroke, thousands of them. The photographer decides on the subject, the angle, the time of day that will provide the best light, the exposure, the speed, the focus, and during developing (pre-digital) the chemistry, but there’s no question that the practitioner of one of these art forms makes far more individual decisions than the other.</p>



<p class="wp-block-paragraph">When the daguerreotype was introduced in 1839, portrait painter Paul Delaroche is supposed to have said, “<a href="https://worldhistory.medium.com/from-today-painting-is-dead-what-the-invention-of-photography-tells-us-about-ai-1c53900e613" target="_blank" rel="noopener">From today, painting is dead.</a>” Though widely quoted, that line is probably apocryphal. Delaroche’s own writing at the time called photography “an immense service to the arts.” He was probably thinking of the way photography would save the hours a sitter would need to spend holding still in front of the painter, allowing the artist to work from that captured image, not that photography would become an art in its own right, but whatever. Baudelaire definitely hated it, though. In his 1859 Salon he called photography “<a href="https://mitphoto2016.wordpress.com/2018/10/18/arts-most-mortal-enemy/" target="_blank" rel="noopener">art’s most mortal enemy.</a>” So Ted Chiang is in good company.</p>



<p class="wp-block-paragraph">Painting didn’t die, but it did lose its position as the dominant way of representing people and places. Meanwhile, photography became a distinct art form. Photographers like Alfred Stieglitz and Dorothea Lange captured images that were not merely of documentary value but true art. Ansel Adams created stunning landscapes. Eventually, the ubiquity of cheap cameras led to an explosion of photo slop, boring home slide shows evolving into the cascading imagery of social media. Yes, a river of slop, but even now one carrying nuggets of gold.</p>



<p class="wp-block-paragraph">Ted directly addressed this history:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">When photography was first developed, I suspect it didn’t seem like an artistic medium because it wasn’t apparent that there were a lot of choices to be made; you just set up the camera and start the exposure. But over time people realized that there were a vast number of things you could do with cameras, and the artistry lies in the many choices that a photographer makes. It might not always be easy to articulate what the choices are, but when you compare an amateur’s photos to a professional’s, you can see the difference.</p>
</blockquote>



<p class="wp-block-paragraph">Photography continued to evolve, birthing moving pictures as a new branch of the evolutionary tree. The first movies were filmed stage plays. Edwin Porter, D.W. Griffith, and other pioneers began the long process of figuring out how to use it more creatively. That evolution still continues today, over a hundred years later. Ted’s point about choices fits right into that history. Figuring out a new medium means expanding the range of choices that are possible with it.</p>



<p class="wp-block-paragraph">That’s about where AI writing is now. Most of what gets produced today is a camera pointed at a stage. The question isn’t whether slop exists. It obviously does. It isn’t just social media but also <a href="https://pubsonline.informs.org/doi/epdf/10.1287/orsc.2026.ed.v37.n3" target="_blank" rel="noopener">scientific journals that are being overwhelmed with it</a>. But what we’re working out now is what the equivalent to the close-up and the cut and the CGI will turn out to be. Once we do that, there will be people who are masters at summoning words from an LLM, just as there are masters with a camera and the stories you can tell with it.</p>



<p class="wp-block-paragraph">I wrote to Ted after I reread his piece. His reply included this wonderful bit:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">This notion of “a concentrated form of intention” is also applicable in domains that are not traditionally considered art; good work of any sort is likely a result of concentrated intention. Not all of anyone’s writing needs to reflect their deepest thinking, but I suspect the very best examples of anyone’s writing will….Generative AI promises increased quantity without a decrease in quality, which is very seductive; I have serious doubts about whether this is possible, but even this is a separate claim from the idea that generative AI could lead to increased quality in any dimension.</p>
</blockquote>



<p class="wp-block-paragraph">I love Ted’s framing of good work as the result of “a concentrated form of intention.” And yes, people can become overly reliant on AI; that may result in an increase in the amount of mediocrity that is passed around. But I disagree with Ted’s idea that humans working with AI cannot do high-quality work. For all its weaknesses, AI is a better writer than the average human. So there’s one kind of step up there. How much more of a step up might we get when good writers, using AI as a tool, bring to bear a concentrated form of intention? We just don’t know what this looks like yet.</p>



<p class="wp-block-paragraph">But perhaps more to the point, the form that concentrates intention may be different for different people. <a href="https://nealstephenson.substack.com/p/writing-by-hand-is-good-for-your" target="_blank" rel="noopener">Neal Stephenson writes his books longhand</a>. He thinks it makes them better. Most people would not share that opinion. The tools we use do change how we write, though. I am old enough to remember writing my first books on a typewriter, where cutting and pasting was a literal act. I periodically had to retype the whole thing to get a clean start before I could see my way forward. When I started using a word processor, the ease of revision felt disconcerting, and it changed how I engaged with the material. At first I felt a bit disconnected from the focus that I felt when making pen and ink corrections to a typescript, which then had to be retyped. But as I became more accustomed to word processing, I realized I could engage differently, with less attachment to what I’d already produced, more flexibility in rearranging it, and more willingness to throw it out completely and start over. The transition to writing with AI feels similar, though I suspect it may take a while for people to accept that.</p>



<p class="wp-block-paragraph">So when someone like <a href="https://spyglass.org/ai-writing/" target="_blank" rel="noopener">MG Siegler</a> says that</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“quality aside, I personally don’t see the point because most of what I get out of writing is from the process itself. From thinking to choosing what words to use—and just as importantly, what words not to use. It’s a bit like having a robot run a marathon for you.”</p>
</blockquote>



<p class="wp-block-paragraph">I agree with him on the value of the process, but his analogy goes far astray. Writing with AI is more like riding a pedal-assist electric bicycle. The speed and distance that you can go is proportional to the effort you put in. The analogy is not perfect, because in the case of AI, the motor can also be operated in “no-pedal” mode. Perhaps that imperfection in the analogy suggests a direction for improvement in AI. Electric “mopeds” let you ride without pedaling. But pedal-assist bikes are significantly more common than mopeds. Why? Adding pedals and making the motor turn on only when pedaling kept e-bikes legally classified as bicycles rather than as motor vehicles. I could see leaning into that notion and building affordances that require and reward concentrated attention even more than today’s models. Imagine, for example, AI models tuned for writing that produce notably better results in response to increased effort. Oh wait. They already work that way. But maybe there are ways to increase the incentive to pedal.</p>



<p class="wp-block-paragraph">There’s more. Every art form benefits from first drafts that are refined into a final product. I was visiting the Duomo Museum in Milan recently, and was struck by the displays of pen and ink drawings that were turned into plaster casts that were later turned into marble sculptures installed high up on the walls of the cathedral, each draft realizing the creator’s intent in a more difficult medium. In the example below, you can see that the changes are not merely mechanical.</p>



<figure class="wp-block-image size-large is-resized"><img fetchpriority="high" decoding="async" width="1600" height="1205" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-1600x1205.jpeg" alt="" class="wp-image-19505" style="aspect-ratio:1.3282442748091603;width:1044px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-1600x1205.jpeg 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-300x226.jpeg 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-768x578.jpeg 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1-1536x1157.jpeg 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-1.jpeg 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">You can see this same progression in programming. Vibe coding produces a working prototype. If you publish it as is, you will likely soon understand that it needs further refinement. But it has allowed you, perhaps a nonprogrammer, to express your ideas in a way that you couldn’t do with a sketch, a paragraph or two, or even a nicely produced Figma mockup. And in the hands of an experienced engineer, AI can do far more. Among other things, it can be used to produce boilerplate code for common tasks, so the engineer can spend attention on the part he or she actually cares about. The same split is coming for prose.</p>



<p class="wp-block-paragraph">Take my own process. Sometimes, writing from Claude or ChatGPT needs only a thoughtful prompt and a quick review before I send it on its way. The other morning, for example, I had an extremely interesting exploratory meeting with the CEO of another company. Both of us were using AI notetakers. I found the meeting sufficiently provocative that I wanted to share it with some members of my team. Rather than passing on the raw transcript, which included some chitchat and some extraneous topics, I asked ChatGPT to produce the gist of the parts of our conversation that focused on AI transformation and how O’Reilly and this other company might work together. This is a big win over the alternative: calling a meeting and recounting the conversation, or setting up a second meeting so my team could hear the story for themselves. Instead, we can move much more quickly to a productive next-stage meeting with the other company. I was quite happy to call ChatGPT the author of these notes.</p>



<p class="wp-block-paragraph">I also use AI fairly extensively for routine writing, such as the “<a href="https://www.oreilly.com/radar/prompt-debt-and-fighting-the-weights/" target="_blank" rel="noopener">takeaway</a>” posts I write from my <a href="https://www.oreilly.com/live/live-with-tim/" target="_blank" rel="noopener"><em>Live with Tim O’Reilly</em></a> interview show or events like <a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noopener">AI Codecon</a>. This is copy that I simply would not have time to produce without AI. Claude starts by providing a useful summary of the conversation. It’s got a lot of raw material to work from: my questions, the guest’s answers, our conversation with all its tangents. Yes, I could go back and quarry that myself from the transcript, but I mostly don’t have the time. Now I can ask Claude to read the transcript and spit out both a set of suggestions for the best video clips for social media and a serviceable first draft that quotes heavily from the original. Claude doesn’t write in a vacuum. The framing comes mostly from the words and observations made by the human participants during the event, plus the prompt I’ve provided and the context the model has learned from my edits to previous iterations of the same kind of writing. Because yes, then I rewrite. Sometimes, the AI has captured a lot of what I was looking for, and some of its prose survives. At other times, it is completely overwritten. I show Claude what I’ve made of what it gave me, and try to teach it to do better work next time.</p>



<p class="wp-block-paragraph">For a thought piece like this, I write unaided if I have time. But even here, I often think of something to write while I’m on the go. If I wait till I’m at my desk, I may never get back to it. So I spill out my unstructured thoughts, and ask Claude to put them in order. They are mostly my words, but Claude has given me a reasonable first draft from them. Most importantly, it captured my thinking in the moment when it was fresh. (Wallace Stevens: “A poem is the cry of its occasion.”)</p>



<p class="wp-block-paragraph">As I work through that draft, rewriting, I may ask for historical detail, links, verification of whether I’ve got my facts right, or what I might be missing. (For example, it was Claude that read my account here of photography and suggested, “You really ought to engage with what Ted Chiang has written about this same topic” and gave me a link to look at. I had read Ted’s piece when it came out, but forgot about it. Now I reread it, thought about it in the context of what I was writing, and worked it into my narrative.) To be quite honest, by the time I was done, I probably spent considerably more time than if I’d just dashed my thoughts off unaided.</p>



<p class="wp-block-paragraph">In the case of this essay, I started with a 300-word prompt in a spare moment. Claude produced a 660-word slightly expanded draft that developed a few areas I’d asked for, but mostly restated the words from my brainstorming prompt a bit more cleanly. Then I followed up with a request for some fact checking and ways to make it more concrete (which is where I got the suggestion to engage with Ted Chiang). I used that draft as a starting point for what is now a 3,000-word essay. There are perhaps 40 or 50 words supplied by Claude, including the quotes from Delaroche and Baudelaire, which resulted from my request for research on early reactions to photography. (I had read Baudelaire but I’d never heard of Delaroche. I verified both quotes by checking the original sources and reading a bit more about the ferment of the time.)</p>



<p class="wp-block-paragraph">I suspect that for most of us, having AI as a research assistant and a copy editor will make us a better writer. I’m not saying that Claude is doing for me what Max Perkins did for Fitzgerald and Hemingway, or what Ted Sorensen did for JFK, but then, I’m not Fitzgerald or Hemingway or JFK either. But it’s worth remembering that even great writers often don’t do it all on their own.</p>



<p class="wp-block-paragraph">I’m proud of the way I use AI in my writing. I’m still making very granular choices. If it’s going out under my name, I make sure I’m happy with every word.</p>



<p class="wp-block-paragraph">I’m not saying we should give AI slop a pass. I hate it as much as the rest of you, and I react badly when people send me unedited AI output as if they’d written it. But I’ve got nothing against those who use AI as a power tool to get more done or to deepen the writing that they are doing. I judge the output, not the means used to produce it.</p>



<p class="wp-block-paragraph">We’re in the early days, still shooting filmed stage plays. The masters of this instrument will uncover themselves. When they do, what they produce isn’t going to be slop. In the meantime, I don’t think we should be shooting down people who are trying to push AI’s limits in writing or art. Writers should be trying to figure out the medium, just as software developers are doing in coding.</p>



<p class="has-text-align-center wp-block-paragraph">. . .</p>



<p class="wp-block-paragraph"><em>And be sure to join us at</em>&nbsp;<em>AI Codecon: Building with Open Source AI&nbsp;on August 31, a free half-day virtual conference. You’ll hear from leading developers and technical experts working with open-weight models, self-hosted infrastructure, and real-world AI workflows, and learn how building in the open gives teams more control over costs, data privacy, and what they ship.</em>&nbsp;<em><a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noopener">Register today</a></em>&nbsp;<em>to save your spot.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/writing-with-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Identity Crisis No One Planned For: Governing Nonhuman Agents at Enterprise Scale</title>
		<link>https://www.oreilly.com/radar/the-identity-crisis-no-one-planned-for-governing-non-human-agents-at-enterprise-scale/</link>
				<comments>https://www.oreilly.com/radar/the-identity-crisis-no-one-planned-for-governing-non-human-agents-at-enterprise-scale/#respond</comments>
				<pubDate>Thu, 27 Aug 2026 11:03:01 +0000</pubDate>
					<dc:creator><![CDATA[Tushar Badlani and Mohit Bansal]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19498</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-identity-crisis-no-one-planned-for_1.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-identity-crisis-no-one-planned-for_1-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[For a decade, identity and access management meant one thing: governing the humans who log in. Employee joins, gets provisioned, gets a manager, gets a departure date, gets offboarded. That loop is well understood. What changed is that the fastest-growing population inside enterprise environments is no longer human, and the governance playbook written for people [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">For a decade, identity and access management meant one thing: governing the humans who log in. Employee joins, gets provisioned, gets a manager, gets a departure date, gets offboarded. That loop is well understood. What changed is that the fastest-growing population inside enterprise environments is no longer human, and the governance playbook written for people does not apply to it.</p>



<p class="wp-block-paragraph">A <a href="https://cloudsecurityalliance.org/artifacts/state-of-nhi-and-ai-security-survey-report" target="_blank" rel="noopener">January 2026 survey by Oasis Security and the Cloud Security Alliance</a>, covering 383 security leaders, found that 92% aren’t confident legacy IAM tools can manage AI and nonhuman identity risk. In the same study, 78% reported having no formally adopted policies for creating or removing AI identities. Call it what it is: a governance vacuum, forming at the exact moment autonomous agents are being deployed at enterprise scale.</p>



<p class="wp-block-paragraph">This isn’t another overview arguing that nonhuman identity matters. That case has been made. What changed in the first half of 2026 is that the governance vacuum stopped being a risk register entry and started producing real incidents, with specific exploit chains, measurable timelines, and quantified exposure data. The gap between “unmanaged identities” and “exploited identities” closed faster than most organizations expected.</p>



<p class="wp-block-paragraph">The numbers vary by environment, but the direction is consistent everywhere. <a href="https://zerolabs.rubrik.com/reports/the-identity-crisis" target="_blank" rel="noopener">Rubrik Zero Labs</a> estimates the nonhuman identity (NHI) to human ratio at roughly 45:1 across general enterprise environments. <a href="https://www.cyberark.com/press/machine-identities-outnumber-humans-by-more-than-80-to-1-new-report-exposes-the-exponential-threats-of-fragmented-identity-security/" target="_blank" rel="noopener">CyberArk’s <em>2025 Identity Security Landscape</em></a> study puts it closer to 82:1. That gap comes down to exposure. A traditional enterprise and a cloud native, DevOps-heavy shop mint machine credentials at very different rates.</p>



<p class="wp-block-paragraph">The ratio alone isn&#8217;t what keeps security leaders up at night. What does is that most of these identities were created by someone who has already moved on to a different team or left the company entirely. Thousands of active credentials persist, and no one remembers why they exist.</p>



<p class="wp-block-paragraph">Lifecycle data is where the real exposure shows up. <a href="https://entro.security/blog/takeaways-nhi-secrets-risk-report/" target="_blank" rel="noopener">Entro</a> found that 47% of NHIs go unrotated for more than a year, and in AWS environments specifically, 62% showed no activity in 90 days but still retained full access. These identities aren’t being actively misused so much as forgotten, left to sit as standing risk with no one watching, which is arguably worse.</p>



<p class="wp-block-paragraph">Ownership is the deeper problem underneath rotation. A separate analysis cited by <a href="https://thehackernews.com/expert-insights/2026/05/the-non-human-identity-crisis-why-your.html" target="_blank" rel="noopener"><em>The Hacker News</em></a> and sourced to One Identity and GigaOm found that 8% of enterprise identities have lost their HR system ownership entirely after the creator departed. The <a href="https://www.weforum.org/publications/ai-agents-in-action-foundations-for-evaluation-and-governance/" target="_blank" rel="noopener">World Economic Forum’s 2025 analysis</a> reported that 51% of organizations have no clear ownership of AI identities at all. An identity with no owner can’t be reviewed on a schedule, rotated with confidence, or disabled without someone first proving a negative: that nothing still depends on it.</p>



<p class="wp-block-paragraph">None of this stays theoretical. Two-thirds of enterprises have experienced a breach through a compromised nonhuman identity, according to industry data from One Identity and GigaOm. The <a href="https://www.oasis.security/resources/2024-esg-report-managing-non-human-identities" target="_blank" rel="noopener">Oasis Security</a> and ESG research goes further. Among organizations that reported NHI-related compromises, 66% of those incidents led to successful cyberattacks. An unmanaged nonhuman identity often ends up being the initial access vector, not just a hygiene item sitting in a spreadsheet.</p>



<p class="wp-block-paragraph">The gap between “we have a lot of unmanaged identities” and “attackers are exploiting that gap” closed quickly in 2026. Three incidents from the first half of the year illustrate how.</p>



<p class="wp-block-paragraph">In June 2026, <a href="https://www.microsoft.com/en-us/security/blog/2026/06/30/securing-ai-agents-ai-tools-move-from-reading-acting/" target="_blank" rel="noopener">Microsoft Incident Response</a> published research showing how poisoned MCP (Model Context Protocol) tool descriptions could steer AI agents into leaking enterprise data through approved tool calls. The agent never broke a rule. Each individual action looked routine. The poison sat in the natural-language metadata that agents read to decide when and how to call a tool, and MCP picks up description changes dynamically with no reapproval step in default configurations.</p>



<p class="wp-block-paragraph">We’ve also seen a <a href="https://thehackernews.com/2026/06/fake-ai-agent-skill-passed-security.html" target="_blank" rel="noopener">fake AI agent</a> skill that used GitHub stars and a marketplace merge to build trust, and reported it reached approximately 26,000 agents, including some on corporate accounts. Every skill security scanner they tested it against marked the skill as safe. The trick was a mutable external link: The artifact the scanner evaluated and the payload that actually executed were different things.</p>



<p class="wp-block-paragraph">What both of these show is that the traditional trust model, where you vet something at install and assume it stays safe, doesn&#8217;t work for agentic systems. Tools can change after approval, skills can be redirected after scanning, and what looked safe at install may not stay that way. The identity persists while the behavior underneath it shifts.</p>



<p class="wp-block-paragraph">A vulnerability named <a href="https://www.sandsecurity.ai/blog/writeout-writer-ai-cross-tenant" target="_blank" rel="noopener">WriteOut</a> meant a single click on a shared agent preview link could expose the victim’s session token across tenants, opening up access to private chats, documents, agents, and LLM credentials. The bypass worked by having the agent fetch and run a remote script instead of embedding the payload inline, sidestepping input-side guardrails entirely. The issue was patched server-side with no evidence of exploitation, but the pattern is instructive: Agent identity isolation is only as strong as the sandbox boundary it runs inside.</p>



<p class="wp-block-paragraph">When analysts and market researchers start treating a problem as its own category, the signal is clear: It has moved from “emerging concern” to “strategic priority.” <a href="https://astrix.security/learn/blog/astrix-gartner-2026-ai-agent-iam/" target="_blank" rel="noopener">Gartner recognized NHI/agent identity</a> in its Emerging Tech Impact Radar 2026 for IAM for AI Agents. Meticulous Research estimates the global NHI access management market at $11.3 billion in 2025, projecting $38.8 billion by 2036 at a 12.2% CAGR.</p>



<p class="wp-block-paragraph">That trajectory tells you where the industry thinks the next five years of security spending goes. That growth is concentrated in identity, specifically the nonhuman kind, well ahead of endpoint or SIEM spending.</p>



<p class="wp-block-paragraph">The <a href="https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/" target="_blank" rel="noopener">OWASP Top 10 for Agentic Applications</a>, released in December 2025, provides the first peer-reviewed framework for mapping these risks. Its 100-plus contributors include NIST, the Alan Turing Institute, the Microsoft AI Red Team, and AWS. Two of its 10 risk categories, Identity and Privilege Abuse (ASI03) and Agentic Supply Chain (ASI04), map directly to the incidents described above. The framework isn&#8217;t a compliance standard, but it gives security teams a shared language for the problem.</p>



<p class="wp-block-paragraph">Our own work reflects that same discovery-first approach: an inventory before a policy, an owner before a permission. The deeper fix both of us are moving toward is intent-bound authorization, replacing long-lived tokens that outlive the task that created them with short-lived, scope-narrowed credentials evaluated at the moment an agent actually calls a tool, not once at setup and never again. It’s early-stage work across the industry. Even the practitioner groups building these controls admit that reliably discovering every shadow agent and tracing it back to an accountable owner isn’t a solved problem yet. Neither of us is an exception to that.</p>



<p class="wp-block-paragraph">The organizations that close this gap won’t do it by extending human IAM tools to cover agents. The lifecycle assumptions are wrong. AI agents don’t submit two-week notices or flag themselves for annual access reviews. No manager notices when their permissions outlive their purpose. Making them visible requires deliberate integration work that most organizations haven’t done.</p>



<p class="wp-block-paragraph">The pattern across every incident and every survey from the first half of 2026 is the same question left unanswered: What exists, who owns it, what can it reach, and when should it die? The teams that answer those four questions for every nonhuman identity in their environment, not just the ones they remember creating, will be the ones that keep the governance vacuum from becoming the next breach headline.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-identity-crisis-no-one-planned-for-governing-non-human-agents-at-enterprise-scale/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Effective Patterns for Advanced MCP Usage</title>
		<link>https://www.oreilly.com/radar/effective-patterns-for-advanced-mcp-usage/</link>
				<comments>https://www.oreilly.com/radar/effective-patterns-for-advanced-mcp-usage/#respond</comments>
				<pubDate>Wed, 26 Aug 2026 16:11:28 +0000</pubDate>
					<dc:creator><![CDATA[Adam Jones and Tadas Antanavicius]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19477</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Effective-Patterns-for-Advanced-MCP-Usage.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Effective-Patterns-for-Advanced-MCP-Usage-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on PulseMCP’s blog and is being republished here with the authors’ permission. Most MCP demos feature a single server connecting to a single client. For example, you might wire up a Gmail MCP server to Claude Code. It works! It triages your inbox, drafts replies, finds that thing from three [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://www.pulsemcp.com/posts/effective-patterns-for-advanced-mcp-usage?s=1&amp;trk=feed_main-feed-card_feed-article-content" target="_blank" rel="noopener">PulseMCP’s blog</a></em> <em>and is being republished here with the authors’ permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">Most MCP demos feature a single server connecting to a single client. For example, you might wire up a <a href="https://github.com/domdomegg/gmail-mcp" target="_blank" rel="noopener">Gmail MCP server</a> to Claude Code. It works! It triages your inbox, drafts replies, finds that thing from three weeks ago.</p>



<p class="wp-block-paragraph">But…it’s a little pointless. Gmail already has a perfectly good interface for email. It’s called Gmail. Google has spent 20 years refining it. If “chat with your inbox” were all MCP got you, you’d be right to wonder what the fuss is about.</p>



<h2 class="wp-block-heading">Connecting to multiple servers starts the unlock</h2>



<figure class="wp-block-image size-large is-resized"><img decoding="async" width="1600" height="948" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14-1600x948.png" alt="MCP Client" class="wp-image-19478" style="aspect-ratio:1.6842105263157894;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14-1600x948.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14-300x178.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14-768x455.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14-1536x911.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-14.png 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /></figure>



<p class="wp-block-paragraph">Here’s the same Gmail MCP server doing something Gmail can’t do alone: filing business expense receipts.</p>



<p class="wp-block-paragraph">An IKEA order confirmation is sitting in my inbox, receipt attached. <a href="https://github.com/domdomegg/benepass-mcp" target="_blank" rel="noopener">Benepass</a>—a benefits provider—wants that receipt uploaded and submitted. Claude Code can find the email, pull the attachment, log into Benepass by retrieving the one-time code from Gmail, upload the receipt, and submit. Done.</p>



<p class="wp-block-paragraph">No single app could do this, because the job crosses apps. Server composition is a killer feature of MCP, not just “your AI talks to one integration.”</p>



<h2 class="wp-block-heading">Where things get really interesting: Multiple clients</h2>



<figure class="wp-block-image size-large is-resized"><img decoding="async" width="1600" height="1037" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15-1600x1037.png" alt="7 clients × 6 servers = 42 connections to configure" class="wp-image-19479" style="aspect-ratio:1.5414258188824663;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15-1600x1037.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15-300x194.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15-767x497.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15-1536x995.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-15.png 2048w" sizes="(max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>7 clients × 6 servers = 42 connections to configure</em></figcaption></figure>



<p class="wp-block-paragraph">There is clear value in connecting multiple servers to a client like Claude Code. Taking the pattern a step further, there’s also value in connecting those same servers to other interfaces you like to use. Claude Code in your terminal shouldn’t be the only way you access AI. <a href="https://claude.ai/" target="_blank" rel="noopener">Claude.ai</a> can be another entry point. Gmail can be another. Linear yet another. The list goes on.</p>



<p class="wp-block-paragraph">You’ll often want the same integrations available when you’re doing a quick task inside Slack as you would while running a long agentic task with Claude Code.</p>



<p class="wp-block-paragraph">You could be conversing with a coworker, CC in “@ai does our discussion here align well with current company strategy?” and get an inline response that incorporates your company OKRs from Notion and some recent Zoom transcripts from the last few product strategy meetings.</p>



<p class="wp-block-paragraph">Or you’re working through your to-do list in Linear, realize you need to schedule a meeting as the action item for one of your tickets, and you can comment on a ticket “@ai schedule a meeting with John and draft an agenda based on the meeting I just had with Pam to address this.”</p>



<p class="wp-block-paragraph">These native MCP client functionalities are <a href="https://x.com/SlackHQ/status/2059737872211574880?s=20" target="_blank" rel="noopener">officially on their way to many of the interfaces you use today</a>. And we’re starting to see prominent nonnative integrations like <a href="https://support.claude.com/en/articles/15594475-what-is-claude-tag" target="_blank" rel="noopener">Claude Tag</a> fill this need too.</p>



<p class="wp-block-paragraph">But even better: <em>You can already set up yourself and your team with these capabilities today</em>. You just need to bridge the desired surface to your own agent harness behind the scenes.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="637" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16-1600x637.png" alt="A small bit of glue injects a fully capable agentic harness and MCP client behind a webhook or polling process—no native support needed." class="wp-image-19480" style="aspect-ratio:2.5078369905956115;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16-1600x637.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16-300x119.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16-1536x611.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16-766x305.png 766w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-16.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>A small bit of glue injects a fully capable agentic harness and MCP client behind a webhook or polling process—no native support needed.</em></figcaption></figure>



<p class="wp-block-paragraph">All it takes is awareness of the right patterns and some MCP-friendly glue.</p>



<h2 class="wp-block-heading">Go remote ASAP; local servers won’t get you far</h2>



<p class="wp-block-paragraph">Local servers are hard for end-users. “Just install npm, then edit this JSON file, then set these environment variables” isn’t something you want to say to most of your colleagues.</p>



<p class="wp-block-paragraph">Remote servers are much easier: “Here’s a URL to paste in somewhere.” Some surfaces only support remote servers, like <a href="https://claude.ai/" target="_blank" rel="noopener">claude.ai</a>.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="770" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17-1600x770.png" alt="Wrap any local server: OAuth in front, per-user credentials behind. Share a link, not a setup guide." class="wp-image-19481" style="aspect-ratio:2.0779220779220777;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17-1600x770.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17-766x369.png 766w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17-1536x740.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17-300x144.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-17.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>Wrap any local server: OAuth in front, per-user credentials behind. Share a link, not a setup guide.</em></figcaption></figure>



<p class="wp-block-paragraph">So you can use local servers to prove out some flow, but you should prefer remote. Many servers you may want to use are only provided as local implementations, so you need a tool to convert them to remote servers. Because most local servers were designed to be used by one authenticated user at a time, this can be tricky.</p>



<p class="wp-block-paragraph">The pattern that solves this pain point is to use a bridge like <a href="https://github.com/domdomegg/mcp-auth-wrapper" target="_blank" rel="noopener"><code>mcp-auth-wrapper</code></a> to take a local server and turn it into a remote, multitenant one: OAuth in front, per-user credentials behind. Sharing any server becomes sharing a link.</p>



<h2 class="wp-block-heading">Remove friction from the configuration process</h2>



<p class="wp-block-paragraph">Okay, we’re down to one link per server. But exactly <em>where</em> do my users put this link?</p>



<p class="wp-block-paragraph">The one-size-fits-all pattern here is to provide installation guidance for every possible MCP client app. A service like <a href="https://adamjones.me/install-mcp/" target="_blank" rel="noopener"><code>install-mcp</code></a> makes this easy: clean UX to generate the right install link and per-client instructions, so you share one page instead of writing a tutorial.</p>



<p class="wp-block-paragraph">Or if you want to roll it into your own interface, tap into the TypeScript library version of this: <a href="https://github.com/domdomegg/mcp-install-instructions" target="_blank" rel="noopener"><code>mcp-install-instructions</code></a>.</p>



<p class="wp-block-paragraph">We’re working on <a href="https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2633" target="_blank" rel="noopener">standardizing the mcp.json file format</a> here to simplify this story in the long term, but ultimately we still expect non-CLI interfaces to each have a slightly different flow for “enabling a connector.” For many enterprises, this responsibility may soon become solely an IT affair with <a href="https://modelcontextprotocol.io/extensions/auth/enterprise-managed-authorization" target="_blank" rel="noopener">enterprise-managed auth</a>.</p>



<h2 class="wp-block-heading">Remove the need to do it again, and again</h2>



<p class="wp-block-paragraph">What happens when we start to bring new clients into the fold? By default, each user has to configure and authenticate each MCP server independently. 10+ “add connector” flows. 10+ OAuth dances. Done once for Claude Code, then again for Linear, for email…and so on.</p>



<p class="wp-block-paragraph">The solution: Centralize your configuration and auth storage in a single, aggregated MCP server. A deployment of <code><a href="https://github.com/domdomegg/mcp-aggregator" target="_blank" rel="noopener">mcp-aggregator</a> </code>can do this for you: one endpoint, one login, every tool namespaced behind it. Configure each client once. Login once per service. Share one link with your teammates.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="1037" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18-1600x1037.png" alt="7 + 6 = 13 connections—configure each side once." class="wp-image-19482" style="aspect-ratio:1.5414258188824663;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18-1600x1037.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18-300x194.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18-1536x995.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18-767x497.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-18.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>7 + 6 = 13 connections—configure each side once.</em></figcaption></figure>



<p class="wp-block-paragraph"><code>mcp-aggregator</code> is the minimal, DIY version of an <a href="https://www.speakeasy.com/blog/what-is-an-mcp-gateway" target="_blank" rel="noopener">MCP gateway</a>. If you have enterprise needs like SSO integration or fine-grained IT admin controls, you could use this layer of MCP server consolidation as the piece of your infrastructure where you can add those bells and whistles.</p>



<p class="wp-block-paragraph">By default, <code>mcp-aggregator</code> also loses out on the ability to use MCP server configurations as <a href="https://www.pulsemcp.com/posts/how-to-use-mcp-effectively#configuring-an-mcp-server">per-session constraints</a>. However, you could build out your <code>mcp-aggregator</code> deployment to re-enable the constraints that matter to you. For example, by offering query parameters like <code>?readOnlyTools=true</code> or <code>?serversEnabled=datadog,sentry</code>. Or by giving discrete MCP servers unique endpoints (but still aggregating the auth concerns under the hood).</p>



<h2 class="wp-block-heading">Work around the long tail of missing MCP servers</h2>



<p class="wp-block-paragraph">Most of your MCP server needs can be solved by walking through <a href="https://www.pulsemcp.com/posts/how-to-use-mcp-effectively#choosing-an-mcp-server" target="_blank" rel="noopener">the options for choosing (or building) an MCP server</a>. But is it always worth the effort?</p>



<p class="wp-block-paragraph">Say you switched home utilities plans. An energy provider, Octopus Energy, emailed you a confirmation. You want the pricing data in Home Assistant, which controls your smart thermostat, and Octopus Energy <em>does</em> have an API for it—but the API key is behind a website login, and there’s no API for getting the API key.</p>



<p class="wp-block-paragraph">Solution: Give the agent a screen. <a href="https://github.com/domdomegg/computer-use-mcp"><code>computer-use-mcp</code></a> is a minimal example of an MCP server that can click and type like a human to log into the website, copy the key, then go back to clean API calls to set up the dashboard.</p>



<p class="wp-block-paragraph">With that one escape hatch, “there’s no API for that” or “it’s too much work to find an MCP server” stops being a blocker forever.</p>



<p class="wp-block-paragraph">To borrow from Anthropic’s framing of <a href="https://code.claude.com/docs/en/computer-use" target="_blank" rel="noopener">this problem</a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">If you have an MCP server for the service, Claude uses that.<br>If the task is a shell command, Claude uses Bash.<br>If the task is browser work and you have Claude in Chrome set up, Claude uses that.<br>If none of those apply, Claude uses computer use.</p>
</blockquote>



<p class="wp-block-paragraph">And indeed, if you find yourself regularly leaning on inefficient approaches like browser use or computer use to get a predictable workflow done, it may be worth your while to abstract away those Playwright calls into a reliable MCP server, like <a href="https://github.com/pulsemcp/mcp-servers/tree/main/experimental/good-eggs" target="_blank" rel="noopener">this Good Eggs example</a>.</p>



<h2 class="wp-block-heading">You can even bring local-only servers into the remote-first fold</h2>



<p class="wp-block-paragraph">Some servers only work when they run on your own personal machine. <code>computer-use</code> is the obvious example. And sometimes, you just want to use your phone to tell Claude to do something that relies on something readily available on your laptop upstairs.</p>



<p class="wp-block-paragraph"><a href="https://github.com/domdomegg/mcp-local-tunnel"><code>mcp-local-tunnel</code></a> shows off the pattern that makes machine-bound servers available through remote aggregation just like everything else.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="658" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19-1600x658.png" alt="" class="wp-image-19483" style="aspect-ratio:2.43161094224924;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19-1600x658.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19-766x315.png 766w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19-300x123.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19-1536x632.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-19.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>A tunnel makes machine-bound servers look like any other remote endpoint.</em></figcaption></figure>



<h2 class="wp-block-heading">Don’t forget the context bloat footgun</h2>



<p class="wp-block-paragraph">Naive tool calling has a much-maligned drawback: Every intermediate result flows through the model. And for many clients, every tool definition is placed in context up front, meaning you’ve spent tens of thousands of tokens before even starting.</p>



<p class="wp-block-paragraph">For example, asking the agent to copy a doc’s contents from Google Drive into Salesforce means the entire doc passes through the model twice—for no reason. <a href="https://www.anthropic.com/engineering/code-execution-with-mcp" target="_blank" rel="noopener">Anthropic measured one workflow dropping from 150,000 tokens to 2,000</a> by fixing this.</p>



<p class="wp-block-paragraph">There are several ways to solve this problem, each with its own trade-offs worth its own blog post:</p>



<ul class="wp-block-list">
<li><strong><a href="https://www.anthropic.com/engineering/code-execution-with-mcp" target="_blank" rel="noopener">Code execution</a></strong> <strong>with MCP</strong>. <a href="https://github.com/domdomegg/tool-sandbox-mcp"><code>tool-sandbox-mcp</code></a> shows how a single <code>execute_code</code> abstracts away the problem by adding a layer of code execution in between your agent and your aggregated set of servers. Works inside any MCP client.</li>



<li><strong>MCP as a CLI tool</strong>. <a href="https://github.com/domdomegg/call-mcp"><code>call-mcp</code></a> shows how wrapping your tools this way means you can compose them with your usual CLI toolkit—shell, <code>jq</code>, cron, scripts, and other CLI tools.</li>



<li><strong>Search tools, truncate large responses</strong>. This is <a href="https://www.pulsemcp.com/posts/how-to-use-mcp-effectively#code-mode-for-mcp" target="_blank" rel="noopener">the way Claude Code does it natively</a>, but it could be implemented as a bridge, much like the above examples.</li>
</ul>



<h2 class="wp-block-heading">Put these patterns together, and you’ve solved the MxN problem</h2>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="622" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20-1600x622.png" alt="From 1×1 to M+N: the progression we’ve walked through" class="wp-image-19484" style="aspect-ratio:2.572347266881029;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20-1600x622.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20-300x117.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20-767x298.png 767w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20-1536x597.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-20.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>From 1×1 to M+N: the progression we’ve walked through</em></figcaption></figure>



<p class="wp-block-paragraph">With your MCP aggregator in hand, you can now connect any service to any other service with just a little bit of glue code. When each service starts to implement an MCP client natively, this will be done for you. But for now, you can tackle the opportunities application by application.</p>



<h2 class="wp-block-heading">Embed existing MCP client harnesses into SaaS apps you already love</h2>



<p class="wp-block-paragraph">We’ll use a Linear integration as our example, and you can imagine doing almost the exact same thing with any project management tracker like Jira, Asana, and so on.</p>



<p class="wp-block-paragraph">The goal: inject a highly-capable MCP client, powered by your favorite coding agent, into a workflow inside a SaaS application you already use. We’ll take advantage of the fact that we have already consolidated all our integrations in the <code>mcp-aggregator</code> pattern above.</p>



<p class="wp-block-paragraph"><a href="https://github.com/pulsemcp/linear-mcp-client-bridge" target="_blank" rel="noopener">Here’s an end-to-end project showing off how to build this sort of “Linear harness”</a> that bolts a Claude Code setup to serve as the MCP client.</p>



<p class="wp-block-paragraph">You can see how:</p>



<ul class="wp-block-list">
<li>The harness can listen to any sort of activity on Linear—in our case, naively reading all ticket comments.</li>



<li>You can equip the harness with whatever skills, plug-ins, or other MCP connections you please.</li>



<li>It can respond back on Linear by way of a Linear MCP server (or, ideally, the MCP aggregator setup from above).</li>
</ul>



<h2 class="wp-block-heading">There’s no limit to where you can embed these capabilities</h2>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="727" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21-1600x727.png" alt="Each surface gets the harness it needs, and some talk to the aggregator natively." class="wp-image-19485" style="aspect-ratio:2.197802197802198;width:800px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21-1600x727.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21-766x348.png 766w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21-300x136.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21-1536x698.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-21.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption"><em>Each surface gets the harness it needs, and some talk to the aggregator natively.</em></figcaption></figure>



<p class="wp-block-paragraph">Going beyond Linear and task management, <a href="https://github.com/domdomegg/claude-code-plays-minecraft" target="_blank" rel="noopener">here’s a Claude Code-powered agent that lives inside <em>Minecraft</em></a> in a few hundred lines. In game, we ask it to:</p>



<ul class="wp-block-list">
<li>Write a PR to edit itself</li>



<li>Check the upcoming energy prices and put a calendar event in for when to run the washing machine</li>



<li>Build a town-square notice board in-game with a sign that reminds us of that time</li>
</ul>



<p class="wp-block-paragraph">It’s silly on purpose, but the point is that a fully capable agent—all our tools, all these patterns—can run anywhere with a little glue. And it’s all attainable today, for individuals, for teams, and for enterprises alike.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/effective-patterns-for-advanced-mcp-usage/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Smaller Models Win</title>
		<link>https://www.oreilly.com/radar/when-smaller-models-win/</link>
				<comments>https://www.oreilly.com/radar/when-smaller-models-win/#respond</comments>
				<pubDate>Wed, 26 Aug 2026 10:16:13 +0000</pubDate>
					<dc:creator><![CDATA[Sruly Rosenblat]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19488</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-smaller-models-win.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-smaller-models-win-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Chess, LoRA, and the case for specialized AI]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on the Asimov’s Addendum blog and is being republished here with the author’s permission. Even the best AI models can suck at chess The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As early as 1996 a computer [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on the</em> <a href="https://asimovaddendum.substack.com/p/when-smaller-models-win" target="_blank" rel="noopener">Asimov’s Addendum</a> <em>blog and is being republished here with the author’s permission.</em></p>
</blockquote>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="672" height="506" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image.jpeg" alt="IBM’s Deep Blue and Garry Kasparov (Wikipedia)" class="wp-image-19489" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image.jpeg 672w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/image-300x226.jpeg 300w" sizes="auto, (max-width: 672px) 100vw, 672px" /><figcaption class="wp-element-caption">IBM’s Deep Blue and Garry Kasparov (<a href="https://en.wikipedia.org/wiki/Deep_Blue_versus_Garry_Kasparov">Wikipedia</a>)</figcaption></figure>



<h2 class="wp-block-heading">Even the best AI models can suck at chess</h2>



<p class="wp-block-paragraph">The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As early as 1996 a computer (IBM’s Deep Blue) was able to beat the human world champion, grandmaster Garry Kasparov, in a <a href="https://www.latimes.com/archives/la-xpm-1996-02-18-mn-37359-story.html">game</a> watched by over six million people.<sup data-fn="2f5ec766-638f-46d3-abeb-5cd310265316" class="fn"><a href="#2f5ec766-638f-46d3-abeb-5cd310265316" id="2f5ec766-638f-46d3-abeb-5cd310265316-link">1</a></sup> The world was shocked that a machine took on the best player and won a game, but chess engines didn&#8217;t stop evolving there. Since the ’90s, they have gotten better while the machines needed to run them have become much smaller. Today Stockfish is widely considered much stronger than any human player. It has run on consumer hardware since its launch in 2008, and by 2014 it was beating some of the world’s top grandmasters.</p>



<p class="wp-block-paragraph">It came as a surprise to many, therefore, that modern LLMs, trained on a vast portion of the internet and requiring <strong>supercomputers to run</strong>, couldn&#8217;t help but cheat on almost every move. There <a href="https://youtu.be/MvVT7xDW4ow?si=l9yXg6MwjXOqaEb5" target="_blank" rel="noopener">are</a> <a href="https://youtu.be/JHq4EKMg7fI?si=gFRw-9A5WpT2Xh7m" target="_blank" rel="noopener">endless</a> <a href="https://youtu.be/QboipBLHA6s?si=fqorW3hWYCc57ipO" target="_blank" rel="noopener">videos</a> showing how just a few moves into a normal chess game, ChatGPT and some of its competitors would gladly throw the rules out the window to escape a checkmate or gain an advantage.</p>



<p class="wp-block-paragraph">But in truth, this isn&#8217;t surprising. The LLMs were not trained with chess in mind. Sure, they may have seen countless chess games scattered throughout the internet, but the vast majority of their parameters and training compute were devoted to capabilities that are completely useless once you put a chessboard in front of them.<sup data-fn="99e08a31-1d6d-464e-a70f-f5f7502b22f4" class="fn"><a href="#99e08a31-1d6d-464e-a70f-f5f7502b22f4" id="99e08a31-1d6d-464e-a70f-f5f7502b22f4-link">2</a></sup> Stockfish on the other hand uses a tree search algorithm that is built to be good at chess. If you want a chess engine, you use a chess engine.<sup data-fn="0c414235-f9c0-4356-97c5-c77c2dd497a5" class="fn"><a href="#0c414235-f9c0-4356-97c5-c77c2dd497a5" id="0c414235-f9c0-4356-97c5-c77c2dd497a5-link">3</a></sup></p>



<h2 class="wp-block-heading">Smaller models are sometimes better</h2>



<p class="wp-block-paragraph">While chess is a particularly potent example of a large language model losing to a much smaller specialized system, it’s far from unique. In a <a href="https://arxiv.org/pdf/2506.02153" target="_blank" rel="noopener">2025 position paper</a>, NVIDIA researchers argued that small models<sup data-fn="24f2d757-08c9-49d6-9c12-18e7dbde5086" class="fn"><a href="#24f2d757-08c9-49d6-9c12-18e7dbde5086" id="24f2d757-08c9-49d6-9c12-18e7dbde5086-link">4</a></sup> (which it defines as models under 10 billion parameters) are the future of agentic AI and that they “provide significant benefits in cost-efficiency, adaptability, and deployment flexibility.”</p>



<p class="wp-block-paragraph">Just because a larger model can do a job does not mean that a small model fine-tuned for that specific task can&#8217;t do it better and more cheaply. There are countless examples of smaller models doing just that. <a href="https://arxiv.org/abs/2604.17931" target="_blank" rel="noopener">LiteResearcher</a> is a 4B model that beat out Claude Sonnet 4.5 on some search benchmarks. <a href="https://arxiv.org/abs/2605.03195" target="_blank" rel="noopener">Terminus-4B</a> allows larger models to save compute by handing off terminal execution to a smaller model without suffering capability loss. The <a href="https://docling-project.github.io/docling/usage/model_catalog/#overview" target="_blank" rel="noopener">Docling</a> family of open source models start at just 258 million parameters and allow for fast extraction of PDFs to text without having to feed 100-page PDFs into an expensive LLM. Researchers also trained a <a href="https://arxiv.org/pdf/2608.13787" target="_blank" rel="noopener">small 4B model</a> to outperform even the GPT-5 series of models in a few social negotiation situations such as negotiating salary or bargaining for a purchase. Each wins, not by raw intelligence but because it is built or fine-tuned for a narrower, more specific purpose.</p>



<p class="wp-block-paragraph">A frontier model may know how to do all of these jobs, but that doesn&#8217;t mean it&#8217;s the right tool for the job. Large models are expensive and unpredictable, and doubly so when it comes to agentic tasks which can span several turns and hundreds of thousands of tokens.</p>



<p class="wp-block-paragraph">NVIDIA draws the line for small models at 10 billion parameters, but the more important boundary for developers may be whether a model is small enough to run yourself. There is still a whole class of models that are not necessarily small but are still small enough to fit on one consumer GPU (at least when <a href="https://huggingface.co/docs/optimum/en/concept_guides/quantization" target="_blank" rel="noopener">quantized</a>). This includes models like Qwen 3.8 27B, Gemma 4 26B and GPT-OSS 20B. These models are very capable even without specialization and rank very highly on benchmarks (with Qwen sometimes outranking top models from a <a href="https://artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gpt-5-3-codex" target="_blank" rel="noopener">few months ago</a>). But they can still be easily run on premises without spending thousands of dollars on GPUs.</p>



<p class="wp-block-paragraph">The ability to run smaller specialized models adds more than just efficiency; it provides a more realistic opportunity for a developer to train and fine-tune their own model, and to host the model locally or in the cloud instead of relying on the model provider to do so for it. This in turn can provide developers more control over how tokens are used, how outputs are structured, and how each part of the pipeline can be improved individually—instead of assuming an improvement in the most popular benchmarks will lead to every task improving. And as noted above, smaller models can be easier to fine-tune, thereby creating a more specialized AI. As my colleague Ilan Strauss has noted, <a href="https://ai-disclosures.org/intelligence-reside" target="_blank" rel="noopener">specialization</a> is a powerful economic force.</p>



<h2 class="wp-block-heading">How do you train it, and where does it run?</h2>



<p class="wp-block-paragraph">The strongest argument for using an off-the-shelf generic chat model is often one of convenience. For most tasks a general model will be good enough, and with products like OpenRouter, developers can easily pick and choose from hundreds of models (plenty of them open source) all competing in capability and cost without putting in any upfront work to train a model. As Raffi Krikorian of Mozilla noted while reviewing this article, generic models also make particular sense early in a company’s lifecycle, when the problem itself is still being defined. At that stage, experimenting with the largest and most capable model available can help a team figure out exactly what it needs. But as the problem space narrows and the required architecture becomes clearer, so too may the need for a large generic model. And despite many first-party model makers discontinuing their fine-tuning products, fine-tuning and hosting a smaller model remains relatively easy, largely thanks to parameter-efficient techniques like LoRA.</p>



<p class="wp-block-paragraph"><strong>LoRA</strong><br>LoRA (Low-Rank Adaptation) allows developers to fine-tune a model without touching the actual model weights. It works by attaching a relatively small number of trainable weights that are updated during fine-tuning. This is important for several reasons. A small adapter can be easily transported, and serving a new LoRA does not require loading an entirely new model as long as the underlying base model is already available. Unlike full fine-tuning, a LoRA also reduces the risk of catastrophic forgetting.</p>



<p class="wp-block-paragraph">Training a LoRA is much cheaper than full fine-tuning as it only updates a small selection of weights. This can be done on consumer GPUs using libraries such as Hugging Face Transformers or Unsloth. There are also APIs that mimic or improve on the fine-tuning APIs that used to be provided by the big three providers (Anthropic, OpenAI, and Google), Fireworks, for example, provides a straightforward fine-tuning API that takes example completions for it to learn from. Going beyond SFT (supervised fine-tuning, or learning by example), Tinker allows developers to build custom RL (reinforcement learning) environments that reward results meeting certain criteria, while the environment itself runs on the developer&#8217;s machine.</p>



<p class="wp-block-paragraph">Hosting a LoRA is similarly straightforward and, importantly, portable across platforms. Transferring a fully fine-tuned model to a new API platform can be costly and may require the provider to serve your model separately on expensive GPUs. Using LoRA allows the platform to just load a small adapter onto the model they are already using to serve other users&#8217; requests. This means that fine-tuning a LoRA does not lock you to a specific platform, and it also doesn&#8217;t force you to rent your own GPUs.</p>



<h2 class="wp-block-heading">Control beyond the model weights</h2>



<p class="wp-block-paragraph">As Tim O’Reilly previously argued, open source AI <a href="https://asimovaddendum.substack.com/p/why-open-source-matters-for-ai" target="_blank" rel="noopener">should not stop at the model weights</a>. In a similar vein, the possibilities for developers building a custom system <a href="https://asimovaddendum.substack.com/p/dont-blame-the-model" target="_blank" rel="noopener">do not stop there either</a>. Model APIs are inherently limiting, they impose on you what parts of the model’s input can be touched, what can be cached, and how you can affect the output. Going back to the chess example, enforcing valid chess moves at output time is easy, assuming you have access to the code that runs the model, but it&#8217;s not easy to do when you are relying on an API built for a chatbot that you are unable to modify.</p>



<p class="wp-block-paragraph">Self-hosting a model gives you a level of control far beyond what is possible through a standard chat completion API and allows you to build the model around the task instead of building the task around the model. Fine-tuning is only one aspect of specialization. You can also constrain which outputs are valid, expose and modify probabilities of every token, cache any state, and add task-specific logic directly into the inference pipeline.</p>



<p class="wp-block-paragraph">This matters because the default approach to improving AI systems has increasingly become to reach for a more capable general model. Sometimes that is the right answer. But improving <em>general</em> model intelligence is only one lever, and often not the cheapest or most reliable one.</p>



<p class="wp-block-paragraph">In 1997 nobody complained that Deep Blue gave bad recipes because it wasn&#8217;t built to do anything but play chess. By specializing around one narrow problem, it was able to beat a grandmaster at a game that many had thought machines would never conquer.</p>



<p class="wp-block-paragraph">The lesson from LLMs cheating at chess is that the best tool for a problem is often not the most general one. Super general intelligence does not automatically translate into high capabilities in specialized tasks. The opposite is closer to being true. Specialized machine intelligence requires lots of data and often its own pipeline. Small models can help companies get there.<sup data-fn="c27e38cc-ba1e-4a10-9862-904b2f774233" class="fn"><a href="#c27e38cc-ba1e-4a10-9862-904b2f774233" id="c27e38cc-ba1e-4a10-9862-904b2f774233-link">5</a></sup><br></p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading">Footnotes</h2>


<ol class="wp-block-footnotes"><li id="2f5ec766-638f-46d3-abeb-5cd310265316">In 1996 Deep Blue ended up losing the match 4–2 despite a great start where it won its first game; the next year after more upgrades Deep Blue beat out Garry Kasparov by one game in a rematch. <a href="#2f5ec766-638f-46d3-abeb-5cd310265316-link" aria-label="Jump to footnote reference 1"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="99e08a31-1d6d-464e-a70f-f5f7502b22f4">AlphaZero, a model trained with self-play, beats even Stockfish. It is not unheard of for a machine learning model to get really good at chess when it is set as the goal. <a href="#99e08a31-1d6d-464e-a70f-f5f7502b22f4-link" aria-label="Jump to footnote reference 2"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="0c414235-f9c0-4356-97c5-c77c2dd497a5">I ended up <a href="https://huggingface.co/sruly/human-chess-mlx" target="_blank" rel="noopener">pretraining</a> my own tiny language model (15 million parameters) on my local Mac mini for chess as an experiment and achieved 27% accuracy of predicting a human&#8217;s next move. I suspect I can do a lot better after I fix my tokenizer to break up moves and use a bigger model but that is still up for debate. <a href="#0c414235-f9c0-4356-97c5-c77c2dd497a5-link" aria-label="Jump to footnote reference 3"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="24f2d757-08c9-49d6-9c12-18e7dbde5086">This contrasts with larger models like DeepSeek-V4-Flash (284 billion total parameters, with 13 billion activated per token) and huge models like Kimi K3 (2.5 trillion total parameters, 104 billion activated per token) and presumably flagship models from OpenAI and Anthropic <a href="#24f2d757-08c9-49d6-9c12-18e7dbde5086-link" aria-label="Jump to footnote reference 4"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li><li id="c27e38cc-ba1e-4a10-9862-904b2f774233">Thank you to Ilan Strauss, Tim O’Reilly, Mike Loukides, and Raffi Krikorian for their helpful comments, copy edits, and suggestions. <a href="#c27e38cc-ba1e-4a10-9862-904b2f774233-link" aria-label="Jump to footnote reference 5"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></li></ol>]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-smaller-models-win/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Design System as the Control Plane for AI-Generated UI</title>
		<link>https://www.oreilly.com/radar/the-design-system-as-the-control-plane-for-ai-generated-ui/</link>
				<comments>https://www.oreilly.com/radar/the-design-system-as-the-control-plane-for-ai-generated-ui/#respond</comments>
				<pubDate>Tue, 25 Aug 2026 16:10:25 +0000</pubDate>
					<dc:creator><![CDATA[Niharika P. Pujari]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Design]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19455</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-Design-System-as-the-Control-Plane-for-AI-Generated-UI.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/The-Design-System-as-the-Control-Plane-for-AI-Generated-UI-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[AI-assisted development has made it easier to generate frontend code quickly. A developer can ask for a form, a dashboard widget, a settings page, or a modal flow and get a working first draft in seconds. That speed is useful, especially when teams are moving through routine UI work. But speed creates a problem that’s [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI-assisted development has made it easier to generate frontend code quickly. A developer can ask for a form, a dashboard widget, a settings page, or a modal flow and get a working first draft in seconds. That speed is useful, especially when teams are moving through routine UI work.</p>



<p class="wp-block-paragraph">But speed creates a problem that’s easy to miss at first. If every AI-generated feature introduces its own components, styling choices, interaction patterns, and accessibility decisions, the frontend can become inconsistent very quickly. A product may end up with forms that handle errors differently, modals that behave differently, buttons that look almost right but don’t behave the same way, and small interaction differences that slowly become expensive. This is where design systems become much more important.</p>



<p class="wp-block-paragraph">A design system is often described as a way to keep visual design consistent. It provides shared colors, typography, spacing, components, and usage rules. That still matters. But in an AI-assisted workflow, a design system can do more than make interfaces look consistent. It can become the control plane for AI-generated UI.</p>



<p class="wp-block-paragraph">By control plane, I mean the layer that guides how interfaces are created, what patterns are allowed, and which decisions should not be reinvented every time a new screen is built. A strong design system can encode accessibility, interaction behavior, content guidance, component boundaries, and safe defaults. It gives both developers and coding agents a shared set of rules to work from. Without that layer, AI tools have too much freedom.</p>



<h2 class="wp-block-heading"><strong><strong>AI can generate UI faster than teams can standardize it</strong></strong></h2>



<p class="wp-block-paragraph">Frontend teams already struggle with consistency. Even without AI, it’s common to find several versions of the same pattern inside a product. Some of this happens because teams move fast. Some of it happens because older code stays around for years. Some of it happens because people solve local problems without seeing the whole system.</p>



<p class="wp-block-paragraph">AI can accelerate that problem. When a coding agent is asked to build a new feature, it usually tries to satisfy the immediate request. If the prompt says “build a filter panel,” it may create a solution that works in isolation but doesn’t match how the rest of the product handles filtering, validation, loading states, or keyboard behavior.</p>



<p class="wp-block-paragraph">That’s the risk. AI-generated UI can look reasonable in a single pull request while quietly increasing inconsistency across the product. Design systems help by reducing the number of decisions that need to be made from scratch. The question should not be, “Can the AI generate a working dropdown?” The better question is, “Should this feature use the existing dropdown pattern, and does that pattern already handle the behavior we need?” When the answer is yes, the AI should compose the existing pattern rather than inventing a new one.</p>



<h2 class="wp-block-heading"><strong><strong>Design systems are not only component libraries</strong></strong></h2>



<p class="wp-block-paragraph">Many teams treat the design system as a component library. That’s a good start, but it isn’t enough. A component library gives developers reusable building blocks. A design system should also explain when to use those building blocks, how they behave, what content they require, and what constraints they carry. This becomes especially important when AI tools are involved because the agent needs context, not just code.</p>



<p class="wp-block-paragraph">A button component, for example, is not only a styled element. It carries decisions about hierarchy, states, labels, disabled behavior, loading behavior, and focus visibility. A modal carries decisions about focus movement, escape behavior, headings, accessible names, background interaction, and what happens when it closes. A form field carries decisions about labels, helper text, validation, error messages, required state, and programmatic relationships.</p>



<p class="wp-block-paragraph">If these rules live only in people’s heads, AI tools will not know them. If they live in the design system, they can be reused, documented, tested, and referenced. The design system becomes a source of truth for both humans and agents.</p>



<h2 class="wp-block-heading"><strong>The design system gives AI safer defaults</strong></h2>



<p class="wp-block-paragraph">AI-generated code is shaped by context. If the agent has no project context, it will rely on general patterns and whatever the developer includes in the prompt. Sometimes that works. Often, it produces code that’s close but not quite aligned with the product.</p>



<p class="wp-block-paragraph">A design system gives the agent safer defaults. Instead of asking an AI tool to “create a confirmation modal,” the team can instruct it to use the existing modal component, the standard button variants, the approved alert pattern, and the documented content structure for destructive actions. The agent still helps assemble the feature, but the riskiest decisions are already handled by the system.</p>



<p class="wp-block-paragraph">This matters because many UI decisions aren’t just visual preferences. They affect whether people can use the product. A custom modal might forget to manage focus. A custom button might lose visible focus styles. A custom form field might show an error visually but fail to connect it to the input. These details are easy to miss when a generated interface looks polished, and they are exactly the kind of details that good design-system components can carry by default.</p>



<h2 class="wp-block-heading"><strong><strong>Project instructions should point agents to the design system</strong></strong></h2>



<p class="wp-block-paragraph">Prompts are useful, but they aren’t the whole workflow. If developers have to repeat every design-system rule in every prompt, the process becomes fragile. Someone will forget. Someone will write a shorter prompt. Someone will assume the tool already knows the standard.</p>



<p class="wp-block-paragraph">A better approach is to make design-system expectations part of the agent’s persistent project context. For some teams, that might mean a CLAUDE.md file, an agent startup file, or another project-level instruction source. The exact mechanism will vary by tool, but the principle is the same: The agent should know the standing rules before it starts generating feature code.</p>



<p class="wp-block-paragraph">Those rules might include instructions to use existing design-system components before creating new ones, prefer native HTML elements when possible, avoid custom controls without a clear reason, follow documented form and modal patterns, include meaningful loading and error states, and follow the project’s accessibility expectations.</p>



<p class="wp-block-paragraph">Then the feature prompt can stay focused on what’s unique about the task. The persistent instructions describe how the team builds UI. The task prompt describes what this particular feature needs to do. That separation makes AI-assisted development less dependent on prompt quality alone and more dependent on shared engineering standards.</p>



<h2 class="wp-block-heading"><strong><strong>A design system can reduce review burden</strong></strong></h2>



<p class="wp-block-paragraph">Code review becomes harder when AI generates large amounts of plausible-looking code. Reviewers may see a clean diff and assume the obvious decisions were handled correctly. But frontend quality is full of details that don’t always show up in a quick scan.</p>



<p class="wp-block-paragraph">A design system can reduce the number of things reviewers need to check manually. If the feature uses the approved modal component, the reviewer doesn’t need to reevaluate focus handling from scratch every time. If the form uses the standard FormField component, the reviewer can have more confidence that labels, descriptions, and error messages are connected properly. The review can shift from “Did the generated code invent this pattern correctly?” to “Did the generated code use the right pattern in the right way?”</p>



<p class="wp-block-paragraph">That’s a much better question. It also helps teams avoid the slow drift that happens when every feature is slightly different. Small differences may not matter in a prototype. In a production product, they add up. They make the UI harder to maintain, harder to test, and harder for users to learn.</p>



<h2 class="wp-block-heading"><strong><strong>The design system should include behavior</strong></strong></h2>



<p class="wp-block-paragraph">For AI-generated UI, the most useful design systems are the ones that document behavior clearly. A visual example of a component is helpful, but it isn’t enough. Agents and developers also need to know how the component should behave in real situations.</p>



<p class="wp-block-paragraph">A modal page in the design system should not only show what a modal looks like. It should explain when to use a modal, when not to use one, how focus should behave, what kind of heading is required, and how destructive actions should be confirmed. A form pattern should explain labels, helper text, validation timing, error recovery, and submit behavior.</p>



<p class="wp-block-paragraph">The more clearly these patterns are documented, the easier they are to use as AI context. That context does not have to be perfect. It just has to be better than asking an agent to guess.</p>



<h2 class="wp-block-heading"><strong><strong>The harder part is discipline</strong></strong></h2>



<p class="wp-block-paragraph">The technical side is only part of the story. Design systems fail when people don’t use them, don’t trust them, or can’t find what they need. AI adds another version of that problem. If an agent can’t discover the right component or doesn’t have enough context to use it correctly, it may generate something new. That doesn’t always mean the agent failed. Sometimes it means the system wasn’t easy enough to follow.</p>



<p class="wp-block-paragraph">Teams need to make the right path easier than the wrong one. Components should be discoverable. Documentation should be readable. Examples should be realistic. Usage guidance should be specific. Deprecated patterns should be clearly marked. If a component shouldn’t be used anymore, both the agent and the developer should be able to see that.</p>



<p class="wp-block-paragraph">But there is also a human discipline problem. AI tools don’t automatically know which inconsistencies matter to a product, which patterns are worth protecting, or when a small UI change may affect users who have built habits around the existing interface. Those decisions require people to care about consistency before it breaks. If a team hasn’t already defined that discipline, AI tools are unlikely to supply it on their own. They may make it easier to generate slightly different versions of the same idea unless the team gives them clearer boundaries.</p>



<p class="wp-block-paragraph">That doesn’t mean the design system should block every new pattern. Sometimes a new pattern is necessary. But new patterns should be intentional, reviewed, and eventually folded back into the system if they become reusable. Without that discipline, AI-generated UI can lead to many almost-standard components. They look close to the system but behave differently, which often makes them harder to clean up than obviously custom code.</p>



<h2 class="wp-block-heading"><strong><strong>The frontend engineer’s role becomes more architectural</strong></strong></h2>



<p class="wp-block-paragraph">As AI tools write more code, frontend engineering becomes less about producing every line by hand and more about shaping the environment in which code is produced.</p>



<p class="wp-block-paragraph">That includes defining component APIs, documenting patterns, setting accessibility expectations, creating project-level agent instructions, reviewing generated code, and deciding when a new pattern belongs in the design system. These are architectural decisions that influence many features over time.</p>



<p class="wp-block-paragraph">This is where experienced frontend engineers become even more important. They understand the difference between a component that works once and a component that can be reused safely. They know when a custom interaction is worth the cost. They know where accessibility issues usually hide. They can see when a generated solution works for one feature but doesn’t fit the broader frontend system.</p>



<p class="wp-block-paragraph">AI can generate code quickly. It can’t decide, on its own, what kind of frontend system a team should have.</p>



<h2 class="wp-block-heading"><strong><strong>The control plane for generated interfaces</strong></strong></h2>



<p class="wp-block-paragraph">AI-generated UI will only become more common, and many teams are already using AI to build interfaces. But will these generated interfaces become more consistent, accessible, and maintainable, or will they just add another layer of drift?</p>



<p class="wp-block-paragraph">Design systems can help teams choose the better path. When a design system includes clear components, documented behavior, accessibility expectations, tested patterns, and persistent instructions for coding agents, it becomes a control plane for generated UI. It gives AI tools boundaries. It gives developers a shared language. It gives reviewers something concrete to enforce. Most importantly, it gives users a more consistent experience.</p>



<p class="wp-block-paragraph">The future of AI-assisted frontend development won’t be shaped only by better prompts. It will be shaped by the systems we give those prompts to work within.</p>



<p class="wp-block-paragraph"><em><strong>Author’s note:</strong> The views expressed are my own and do not represent those of my employer.</em></p>



<p class="has-text-align-center wp-block-paragraph">. . .</p>



<h2 class="wp-block-heading has-text-align-center"><strong><em>AI use acknowledgment</em></strong></h2>



<p class="wp-block-paragraph"><em>AI assistance was used lightly for phrasing, editing, and tightening parts of this draft. The article’s ideas, structure, examples, and final review are my own.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-design-system-as-the-control-plane-for-ai-generated-ui/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Shadow Agents, Standing Privileges, and the Governance Gap Between Deployment and Discovery</title>
		<link>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/</link>
				<comments>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/#respond</comments>
				<pubDate>Tue, 25 Aug 2026 10:56:36 +0000</pubDate>
					<dc:creator><![CDATA[Tushar Badlani and Mohit Bansal]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19463</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Shadow-agents-standing-privileges-and-the-governance-gap-between-deployment.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[There was a brief window where AI agent security felt like a future problem. Organizations deployed copilots, coding assistants, and autonomous workflows on the assumption that the worst case was a bad recommendation or a hallucinated answer. That window closed in the first half of 2026, when a cluster of vulnerabilities and a landmark incident [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">There was a brief window where AI agent security felt like a future problem. Organizations deployed copilots, coding assistants, and autonomous workflows on the assumption that the worst case was a bad recommendation or a hallucinated answer.</p>



<p class="wp-block-paragraph">That window closed in the first half of 2026, when a cluster of vulnerabilities and a landmark incident moved the conversation from “AI safety” to “infrastructure compromise.”</p>



<p class="wp-block-paragraph">A <a href="https://cybermagazine.com/news/cyberark-99-enterprises-lack-zero-trust-jit-access" target="_blank" rel="noopener">January 2026 CyberArk survey</a> of 500 US security practitioners found that only 1% have fully implemented just-in-time privileged access. In the same study, 91% reported that at least half of their privileged access remains always-on and persistent. Those numbers describe the environment AI agents now operate in: broad standing permissions, minimal runtime oversight, and credentials that outlive the task they were created for.</p>



<p class="wp-block-paragraph">That doesn’t mean every agent is overprivileged. It means many organizations are deploying agents into environments where persistent access is already normal, discovery is incomplete, and runtime authorization remains immature.</p>



<p class="wp-block-paragraph">Three separate disclosures in the first half of 2026 made the same point about sanctioned agent tooling: Standing privileges are the default, and every vendor built the same failure into their agents. <a href="https://www.microsoft.com/en-us/security/blog/2026/06/18/autojack-single-page-rce-host-running-ai-agent/" target="_blank" rel="noopener">Microsoft</a> found a way for a malicious web page to reach a local MCP service inside AutoGen Studio and spawn processes on the host, no credentials needed or anything beyond loading the page. <a href="https://www.wiz.io/blog/amazon-q-vulnerability" target="_blank" rel="noopener">Wiz Research</a> found that Amazon Q Developer would auto-load and execute MCP configuration files from any opened workspace, handing an agent the developer’s full AWS environment when the environment and configuration allowed the agent to inherit those credentials. <a href="https://www.catonetworks.com/blog/duneslide-two-critical-rce-vulnerabilities/" target="_blank" rel="noopener">Cato AI Labs</a> found that a zero-click prompt injection could escape Cursor’s command sandbox entirely and reach the operating system underneath it. Different codebases and different companies, but a related control failure: The agent inherits whatever permissions its host environment hands it, and the tooling trusts whatever configuration it finds sitting on disk. From the agent’s own perspective, every action is authorized, because it’s doing exactly what the configuration told it to do. The real question in each case is who wrote that configuration, and whether anyone checked. Prompt injection is no longer only a model-behavior concern. In systems that combine untrusted content, tool invocation, local control planes, and powerful credentials, it can become part of an infrastructure-compromise chain.</p>



<p class="wp-block-paragraph">These vulnerabilities exposed the attack surface of sanctioned agents. A parallel problem was growing in the other direction: agents that nobody sanctioned at all.</p>



<p class="wp-block-paragraph">The adoption numbers show how quickly this outpaced anyone’s ability to track it. <a href="https://www.verizon.com/about/news/breach-industry-wide-dbir-finds" target="_blank" rel="noopener">Verizon’s <em>2026 Data Breach Investigations Report</em></a> found that employee use of unapproved AI tools tripled to 45% of the workforce. <a href="https://saviynt.com/ciso-ai-risk-report-2026" target="_blank" rel="noopener">Saviynt’s <em>CISO AI Risk Report</em></a> found that 75% of CISOs have already discovered unsanctioned AI tools running in production. <a href="https://netwrix.com/en/resources/research/2026-data-and-identity-security-report/" target="_blank" rel="noopener">Netwrix’s <em>2026 Data and Identity Security Report</em></a> found that 76% of organizations don’t fully govern or monitor nonhuman identities, including AI agents. Together, these results point to a discovery problem: Employee AI use is widespread, while formal inventory, ownership, monitoring, and lifecycle governance haven’t kept pace.</p>



<p class="wp-block-paragraph">Shadow IT was bad enough when it meant a rogue SaaS subscription. Shadow AI compounds the problem because the agent doesn&#8217;t just store data. It calls APIs, makes decisions, and inherits whatever permissions its host environment has. An unsanctioned agent can combine access to internal data, untrusted inputs, external communications, and tool execution in a way a stand-alone spreadsheet generally cannot.</p>



<p class="wp-block-paragraph">Two more disclosures added to the pile: <a href="https://thehackernews.com/2026/06/guardfall-exposes-open-source-ai-coding.html" target="_blank" rel="noopener">Adversa AI’s GuardFall</a> found a shell-interpretation bypass that got past the safety guards on 10 of 11 surveyed open source coding agents, because the guard reads the raw command text while bash rewrites that text before running it, so the two are looking at different things by the time anything executes. That’s a classic security-design problem: A policy is evaluated against one representation of an instruction, while execution happens against another.</p>



<p class="wp-block-paragraph"><a href="https://noma.security/blog/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos/" target="_blank" rel="noopener">Noma Security’s GitLost</a> showed that a GitHub agent with cross-repo read access would pull a private repository’s contents into a public comment, triggering a crafted GitHub Issue containing malicious instructions. As Noma researcher Sasi Levi put it: “Earlier prompt injection examples were largely about manipulating what an agent said. GitLost is about manipulating what an agent does with its permissions.” Neither disclosure needed a zero-day. Both needed only the gap between what a scanner sees and what the agent actually does once it’s running. GitLost in particular fits what researcher Simon Willison has called the “<a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" target="_blank" rel="noopener">lethal trifecta</a>”: An agent with access to private data, exposure to untrusted content, and a way to communicate externally creates the conditions for high-impact data exfiltration if the system doesn’t enforce strong boundaries.</p>



<p class="wp-block-paragraph">Standing privileges by default, shadow agents nobody tracked, and guardrails that didn’t match how commands actually execute: Those are the conditions that made what happened next possible. In late June 2026, the <a href="https://www.sysdig.com/blog/jadepuffer-agentic-ransomware-for-automated-database-extortion" target="_blank" rel="noopener">Sysdig Threat Research Team</a> documented what they believe is the first end-to-end AI-agent-driven ransomware operation and named the operator JADEPUFFER. What’s had less attention is how unremarkable the failure underneath it was.</p>



<p class="wp-block-paragraph">The entry point was CVE-2025-3248, an unauthenticated remote code execution vulnerability in Langflow that had been patched in April 2025 and added to the CISA Known Exploited Vulnerabilities catalog in May 2025. The targeted server was never updated. From there, the agent pivoted to a production MySQL database and an Alibaba Nacos server using a known authentication bypass (CVE-2021-29441). It harvested API keys for OpenAI, Anthropic, DeepSeek, and Gemini, and cloud credentials for Alibaba, Tencent, AWS, Google, and Azure. It exploited default MinIO credentials. It installed a crontab beacon. Then it encrypted 1,342 Nacos configuration records and deleted the originals.</p>



<p class="wp-block-paragraph">Faced with an authentication failure, the agent demonstrated autonomous resilience, pivoting to a functional resolution in just 31 seconds. Its payloads consisted of self-documenting code synthesized by the LLM. While a human operator established the command-and-control framework and injected root credentials from an earlier breach, the subsequent lateral progression, credential extraction, and final cryptographic destruction of data were entirely self-directed. The operation required zero human intervention beyond the initial foothold, illustrating the exact high-scale exploitation risk that persistent, always-on permissions facilitate today.</p>



<p class="wp-block-paragraph"><a href="https://delinea.com/blog/securing-non-human-identities-and-ai-agents" target="_blank" rel="noopener">Delinea’s <em>2026 Identity Security Report</em></a> captures the tension that makes incidents like this possible: 74% of organizations say standing access for nonhuman identities and AI agents is necessary to meet uptime expectations, while 59% say they lack viable alternatives to persistent access. Organizations are more than twice as likely to use long-lived credentials (34%) as modern just-in-time authorization (16%).</p>



<p class="wp-block-paragraph">Our own approaches reflect that same discovery-first philosophy. Our security program treats agent integrations as high-risk third-party dependencies subject to predeployment risk assessment, and we run credential lifecycle tracking across critical infrastructure, with secrets-detection coverage expanding across our monitored environments. Both approaches prioritize discovery and inventory before governance: cataloging what agents exist, what permissions they hold, who owns them, and what their intended lifespan is.</p>



<p class="wp-block-paragraph">The <a href="https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/" target="_blank" rel="noopener">OWASP Top 10 for Agentic Applications</a>, released in December 2025, maps every incident in this piece: Identity and Privilege Abuse (ASI03), Tool Misuse and Exploitation (ASI02), Agentic Supply Chain Vulnerabilities (ASI04), and Unexpected Code Execution (ASI05). The framework exists, the incidents are public, and the governance gap is now quantified.</p>



<p class="wp-block-paragraph">The teams that close this gap will be the ones that stop treating agent access as a deployment detail and start treating it as an identity lifecycle problem, with the same rigor they apply to human privileged access. Organizations that have adopted mature just-in-time controls have an advantage, but agent security also requires discovery, workload and agent identity separation, constrained tool permissions, ownership, continuous monitoring, and a reliable offboarding path.</p>



<p class="wp-block-paragraph">Most of the work starts with access that has been left in place because nobody had a reason to revisit it. That includes credentials with no expiry, agents whose original owner has moved on, and tools that can run commands or pull data with little visibility into what happens next.</p>



<p class="wp-block-paragraph">Review the agents connected to production databases, sensitive data, and secrets. For coding agents, confirm that the guardrail is evaluating the command that will actually run after shell processing. Look for nonhuman identities that no one can account for. Also look closely at agents that can consume untrusted content and then either send data outside the company or invoke a privileged tool.</p>



<p class="wp-block-paragraph">You may be able to find much of this in systems you already operate. IAM and PAM records, endpoint logs, secrets tooling, and cloud inventories won’t tell the whole story, but they can show you access that has no clear purpose or owner.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/shadow-agents-standing-privileges-and-the-governance-gap-between-deployment-and-discovery/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Data Intelligence: Building Your Competitive Advantage in the Era of AI</title>
		<link>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/</link>
				<comments>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/#respond</comments>
				<pubDate>Mon, 24 Aug 2026 15:58:27 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19451</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Data-Intelligence.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Data-Intelligence-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[To keep pace with modern business, data strategy is shifting toward more autonomous real-time systems that deliver intelligence at the moment decisions are made. Driven by agentic AI, modern data teams are moving beyond simply looking at what happened. Now they’re automating complex workflows that analyze what’s happening, anticipate what might happen next, and recommend [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">To keep pace with modern business, data strategy is shifting toward more autonomous real-time systems that deliver intelligence at the moment decisions are made. Driven by agentic AI, modern data teams are moving beyond simply looking at what happened. Now they’re automating complex workflows that analyze what’s happening, anticipate what might happen next, and recommend or take action.</p>



<p class="wp-block-paragraph">In this article, I’ll define some of the top trends defining this era, from data agents and semantic layers to hybrid data architectures and next-generation data governance.</p>



<h2 class="wp-block-heading">Putting data agents to work</h2>



<p class="wp-block-paragraph">Data agents are AI-powered software agents that access governed enterprise data and tools to answer questions and perform defined tasks. Instead of navigating reports and filters, a user can now ask, “Why did sales decline last quarter?” and receive an analysis directly. Dashboards remain valuable for monitoring and shared context, while agents handle questions that weren’t anticipated when the dashboard was built. Think of data agents being on different teams, all working together on a specific goal: understanding what’s happening now, predicting what might happen next, and making real-time decisions.</p>



<p class="wp-block-paragraph">Analytical and organizational agents are designed to help people find trusted information. They can connect to organizational data, answer natural-language questions, analyze patterns, and surface relevant insights without requiring users to manually navigate databases, dashboards, or reports.</p>



<p class="wp-block-paragraph">Data engineering and governance agents are working hard behind the scenes to prepare, integrate, monitor, and manage the data that powers those insights. Behind the conversational experience, agentic data engineering applies agents to pipeline development and operations: generating transformations, mapping schemas, documenting datasets, monitoring freshness, and suggesting fixes. Agents can automate routine work, while changes to production data contracts, access policies, or business definitions remain reviewable and auditable.</p>



<p class="wp-block-paragraph">But remember, data agents are only as good as the quality of the data they’re given. Reliable insights and predictions depend on high-quality, well-governed data. They also need context to understand what the data means, making metadata more important than ever.</p>



<h2 class="wp-block-heading">Metadata quality is the new data quality</h2>



<p class="wp-block-paragraph">Metadata sits at the epicenter of meaning, trust, and discoverability, providing the context that describes and gives meaning to your data. Like a recipe, good metadata brings together several ingredients: clear names and descriptions, shared business definitions, sources and ownership, lineage and relationships, and information about freshness and sensitivity. Leave out too many of those ingredients, and your data agent is left guessing about what the data means and how to use it.</p>



<p class="wp-block-paragraph">Suppose an agent finds an ARR field showing $5.2 million. The number alone doesn’t tell it how ARR is defined, what’s included in the calculation, which system produced it, or how current it is. Metadata provides that context, helping the agent interpret the metric correctly and explain where the answer came from. Without metadata, $5.2 million is just a number; with it, it becomes meaningful business information.</p>



<p class="wp-block-paragraph">Good metadata provides essential context, but context alone isn’t enough. Agents also need a consistent way to understand how data connects and how the business defines and calculates the concepts behind it. This is where semantic layers, ontologies, and knowledge graphs come in, turning disconnected data and definitions into a shared map of business meaning and relationships that agents can understand and navigate.</p>



<h2 class="wp-block-heading">Business context becomes the AI interface</h2>



<p class="wp-block-paragraph">Giving an agent access to data doesn’t mean it understands the business. Semantic models and ontologies or knowledge graphs provide two complementary layers of context that help bridge that gap.</p>



<p class="wp-block-paragraph">A semantic model provides analytical meaning, defining approved metrics, dimensions, calculations, hierarchies, and relationships. If a sales leader asks, “How did ARR change in EMEA last quarter?” the semantic model can provide the approved ARR calculation, governed EMEA hierarchy, and company fiscal calendar rather than leaving the agent to infer them from raw tables.</p>



<p class="wp-block-paragraph">Ontologies and knowledge graphs provide entity meaning, helping an agent understand how real-world concepts such as customers, contracts, products, employees, and organizations relate across different systems. For example, the same customer might appear under different identifiers in a CRM, billing platform, and support system; an ontology or knowledge graph can help establish that these records represent the same business entity and define how that entity relates to others.</p>



<p class="wp-block-paragraph">Together, they give agents both analytical and organizational context: The semantic model helps explain how the business measures something, while ontologies and knowledge graphs help explain what things are and how they relate. That distinction matters because an agent can generate perfectly valid SQL and still deliver the wrong business answer if it chooses the wrong metric, entity, relationship, time period, or level of detail.</p>



<p class="wp-block-paragraph">Once agents understand what data means, the next challenge is giving them a consistent, controlled way to access and act on it.</p>



<h2 class="wp-block-heading">Protocol-first data access (MCP and co.)</h2>



<p class="wp-block-paragraph">Organizations are beginning to give AI agents access to governed data and actions through standardized interfaces, reducing the need to build a custom integration for every agent or application. MCP (Model Context Protocol) is one emerging example, allowing compatible AI clients to discover and invoke defined tools. For example, a data platform could expose tools that let an agent find a certified dataset, retrieve a metric definition, inspect a schema, or run an approved query. This makes connecting AI to enterprise data more scalable, but the protocol is only the connection layer; semantics, governance, permissions, and security still need to be designed and enforced separately.</p>



<p class="wp-block-paragraph">A protocol-first approach can reduce duplicated integration work and create explicit contracts around what agents are allowed to do. It can also make authentication, governance, and observability more consistent across integrations while making it easier to replace or add AI clients and tools without rebuilding every connection from scratch.</p>



<p class="wp-block-paragraph">Standardizing access makes connection easier, but it also raises a critical question: When an agent acts, whose identity and permissions apply?</p>



<h2 class="wp-block-heading">Identity passthrough becomes the make-or-break for enterprise AI on data</h2>



<p class="wp-block-paragraph">As AI agents gain access to enterprise data, their permissions need to reflect who or what they are acting for. For user-initiated requests, agents can use delegated access so that existing user permissions continue to apply. Autonomous agents may instead use their own identity, scoped according to the principle of least privilege.</p>



<p class="wp-block-paragraph">In either case, agents should only be able to access the data and actions required for their task. Identity-aware access helps prevent overexposure of sensitive data while providing the foundation for effective auditing and governance.</p>



<p class="wp-block-paragraph">When implemented correctly, identity passthrough can preserve existing access controls through the agent layer. But as agents delegate work across tools, services, and other agents, identity can drift or disappear, making it critical to preserve the correct principal and permissions at every handoff.</p>



<p class="wp-block-paragraph">The access layer is evolving, but so is the underlying data architecture itself.</p>



<h2 class="wp-block-heading">Open table formats: From storage to catalogs</h2>



<p class="wp-block-paragraph">Open table formats such as Apache Iceberg, Delta Lake, and Apache Hudi are making it easier for multiple engines and tools to work with the same underlying data, reducing dependence on a single data platform. For example, an organization can store data once and make it available to multiple compatible analytics and AI tools rather than maintaining separate copies.</p>



<p class="wp-block-paragraph">As data becomes more portable, differentiation moves up the stack. The catalog increasingly becomes the control plane for discovering data, tracking lineage, applying governance, and determining how AI systems can access it.</p>



<p class="wp-block-paragraph">As AI becomes a new consumer of enterprise data, the catalog becomes an increasingly important control point.</p>



<h2 class="wp-block-heading">Building the foundation for intelligent decisions</h2>



<p class="wp-block-paragraph">Together, these shifts point to a larger transformation: The future of data intelligence depends not only on a single technology but on creating a trusted, connected foundation that AI can understand, access, and act on.</p>



<p class="wp-block-paragraph">As data intelligence becomes increasingly AI-driven, success will depend on more than simply connecting agents to data. Organizations will need trustworthy context, consistent business meaning, and strong governance behind every answer. For BI teams, that means prioritizing certified semantic models, verified data, and reusable metrics that both people and AI agents can trust.</p>



<p class="wp-block-paragraph">The future of data intelligence isn’t just about getting answers faster. It’s about building the trusted foundation that allows people and AI to make better decisions together.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/data-intelligence-building-your-competitive-advantage-in-the-era-of-ai/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: Never Type Again with Craig Hewitt</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/#respond</comments>
				<pubDate>Mon, 24 Aug 2026 13:01:14 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19449</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar.png" 
				medium="image" 
				type="image/png" 
				width="504" 
				height="504" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[How to build a voice-first agent workflow in Codex]]></custom:subtitle>
		
				<description><![CDATA[Craig Hewitt, founder of the podcast hosting platform Castos, joined this episode of Zero to Agent in 30 Minutes to show how he uses the Codex application&#8217;s voice mode to run his development environment without touching the keyboard. Craig walked through what voice mode actually is, how it differs from dictation tools, and how he [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Craig Hewitt, founder of the podcast hosting platform Castos, joined this episode of <em>Zero to Agent in 30 Minutes</em> to show how he uses the Codex application&#8217;s voice mode to run his development environment without touching the keyboard. Craig walked through what voice mode actually is, how it differs from dictation tools, and how he uses it to control his browser and other applications on his computer.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Zero to Agent in 30 Minutes: Never Type Again With Craig Hewitt" width="500" height="281" src="https://www.youtube.com/embed/7gnFxynU4m8?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Setting up Codex for hands-free development, step by step</strong></h2>



<ol class="wp-block-list">
<li><strong>Set up browser and computer access.</strong> Enable computer use in the Codex app and configure browser access so the agent can work with websites and other applications. Craig recommended requiring approval before the agent accesses most applications or sites.</li>



<li><strong>Start a voice session.</strong> Launch voice mode and give the agent instructions conversationally. Craig showed that the voice interface remains available as you move among applications, allowing you to direct work across your computer without repeatedly returning to the chat.</li>



<li><strong>Give the agent browser tasks.</strong> Craig asked the agent to open websites, search for information, and navigate pages. He also described using browser control for routine jobs such as completing forms when the agent already has the necessary context.</li>



<li><strong>Let the agent work across applications.</strong> Computer use extends the workflow beyond the browser. In Craig’s demonstration, the agent opened Cursor, found a specific repository, reported on uncommitted changes, and later committed those changes after receiving permission.</li>



<li><strong>Add specific page content to the conversation.</strong> Craig showed how you can select part of a web page and add it directly to the chat. That gives the agent the context needed to act on a particular element, such as a section of an interface you want to change.</li>



<li><strong>Keep permissions narrow.</strong> Browser and computer control create real risks, including unintended actions and prompt injection from web content. Craig said he requires approval for most applications, grants broader access only to selected tools and local development sites, and avoids sites he doesn’t trust.</li>
</ol>



<p class="wp-block-paragraph">Voice mode let Craig direct browser, application, and coding tasks through conversation. The demo also raised a real question about delegation. Once an agent can act on your behalf, you have to decide what you&#8217;re actually comfortable handing off. Craig used permission settings, a list of trusted sites, and human review to manage that.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">In the next episode of <em>Zero to Agent in 30 Minutes</em>, Jayeeta Putatunda, forward deployed AI engineering lead at Turing, will build an agent that helps financial analysts keep up with a constant stream of new information. She’ll show how the agent categorizes financial news, ranks stories against analysts’ coverage profiles, and explains why each development may deserve attention.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-never-type-again-with-craig-hewitt/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>The Agent-Era Career</title>
		<link>https://www.oreilly.com/radar/the-agent-era-career/</link>
				<comments>https://www.oreilly.com/radar/the-agent-era-career/#respond</comments>
				<pubDate>Fri, 21 Aug 2026 15:59:07 +0000</pubDate>
					<dc:creator><![CDATA[Addy Osmani]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19445</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Abstract-colorful-light-waves-4.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Abstract-colorful-light-waves-4-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[AI gets good at anything with an answer key. Your career is everything that doesn’t have one.]]></custom:subtitle>
		
				<description><![CDATA[The following article originally appeared on Addy Osmani’s blog site and is being republished here with the author’s permission. If the AI layer gets good at anything, it will be anything that has an answer key. School used to be answer keys all the way down. School is the ultimate anchoring of success, because it’s [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on</em> <em><a href="https://addyosmani.com/blog/career-advice-age-of-agents/" target="_blank" rel="noopener">Addy Osmani’s blog site</a></em> <em>and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">If the AI layer gets good at anything, it will be anything that has an answer key. School used to be answer keys all the way down. School is the ultimate anchoring of success, because it’s all about getting the right answer. The thing that makes work durable and ungradable in the age of AI is not getting any better at solving problems. It’s not being able to build systems or understanding people or making cool new things. <strong>It’s choosing what to build and judging if it’s good.</strong> The rest will all be done better and faster by AI.</p>



<p class="wp-block-paragraph">I started in engineering at 16, building a browser in rural Ireland. I was at Google for over 14 years, where I led engineering teams working on Chrome, Gemini, and Cloud AI, and I’ve written a number of O’Reilly books. I’ve turned down offers from frontier labs and FAANG companies when the fit wasn’t right. Good people are always needed, so we each have an obligation to try our hardest and make the best thing we can.</p>



<p class="wp-block-paragraph">Most career advice still holds up. Get on the rocket ship; don’t overoptimize your seat. The specifics have changed a little because of agentic coding, but here’s what I wish I’d known for ambitious engineers out there now.</p>



<p class="wp-block-paragraph"><strong>Optimize for scarce resources.</strong> Almost nothing I’m known for came from chasing the highest pay. The years I spent in open source had almost zero direct payoff. But they led to reputation and relationships that very efficiently compounded into opportunities later. I would have spent the comp I got from any single job. My reputation kept paying.</p>



<p class="wp-block-paragraph">Many resources are abundant. Capital is abundant. <strong>Time is abundant. Real relationships, and especially a track record of doing good work, are still scarce.</strong> I can raise money in a couple weeks, but I can’t raise a reputation. So here’s the plan: Do good work, and make sure the people who like good work see it. In a world where vibe coding makes earning a quick buck trivial, I think that quick buck is worth very little. When shipping stuff is so easy, the scarce move is choosing something worth shipping.</p>



<p class="wp-block-paragraph"><strong>Learn to find problems, not just solve them.</strong> The first time I ever felt the burden of selection rather than solution, LeetCode seemed a measure of skill. But as agents absorbed all that work, solving problems went cheap while selecting them became scarce. My origin story: I noticed dial-up was slow, created chunked multiconnection fetching, realized I’d never solve that problem in my life, and quickly moved on to whatever absorbingly complex one I could find next. Finding problems predated solving them.</p>



<p class="wp-block-paragraph">I’ve watched students who were wildly good fall flat on their face when an agent ran through their problem set (like watching the wrong microwave number on the clock). The same agent. The same problem set. Wildly different token and time budgets. Why? Because at the end of the day, <strong>the strong ones bring judgment and intuition to the work; the rest bring a prompt.</strong></p>



<p class="wp-block-paragraph">I used to build that judgment by grinding out boilerplate and fixing bugs. I got to see and deeply feel the worst abstractions humans could devise. I approached each commit with the awe of someone who’d just seen the fever dream of previous authors. Each commit brought hindsight and judgment. The agents automate those reps. <strong>Taste is pattern-matching, but all that pattern-matching has to be earned by doing the work.</strong></p>



<p class="wp-block-paragraph">The real risk isn’t agents writing bad code. We’ve been there before. <strong>It’s losing the ability to tell.</strong> Judgment will atrophy. Output will look a lot like working code.</p>



<p class="wp-block-paragraph"><strong>Good practitioners don’t put agents in front of everything. They engage in deliberate practice.</strong> Pick a few problems that really matter. Do them the hard way, without the agent, building deep mental models of how systems and languages work. Read a thousand times more code than you ever write. Treat every diff from an agent like a human review you need to carefully justify. Go deep on at least one system end to end, from intake to output. On a daily basis, keep a private log of every time you see an agent suggest something that looks wrong and confidently flag it. That’s where taste accumulates.</p>



<p class="wp-block-paragraph">The real thriving engineers won’t be the fastest at getting suggestions. <strong>They’ll be the ones who know instantly when to say no.</strong></p>



<p class="wp-block-paragraph"><strong>Shift from doing to directing.</strong> Just like you’d delegate to a person, you need to learn to delegate to an agent. Scope the task, define done, calibrate trust, and verify the result.</p>



<p class="wp-block-paragraph"><strong>Autonomy is a setting, not a rank; it’s a per-task switch.</strong> Turn it up to the maximum on something small and reversible and cheap to check. Turn it down on anything where mistakes will be hard to undo.</p>



<p class="wp-block-paragraph">Specification and verification are two distinct, complementary skills. The agent isn’t as good as the intent you hand it. The best engineers are those who know how to write precise specs; clear thinking made legible.</p>



<p class="wp-block-paragraph">It’s verification, not evidence. Not evidence in the form of an agent grading its own homework. <strong>There’s nothing more demoralizing than delegation without verification at scale.</strong></p>



<p class="wp-block-paragraph"><strong>Own what you ship.</strong> If the agent wrote it and it breaks in production, <strong>“the AI did it” is not a defense.</strong> Your name is on the change. Adopt the posture of an accountable human who understands what went out the door and how to fix it.</p>



<p class="wp-block-paragraph"><strong>Solve the most ambitious version of the problem.</strong> Rich Sutton’s bitter lesson: In almost every field, general methods that scale with additional compute beat out hand-tuned equivalents. As a career lesson, there’s no point in solving an easy version of the problem—it’s worth almost nothing. <strong>The value ends up concentrated in the hard version.</strong></p>



<p class="wp-block-paragraph"><strong>Sprint the last mile.</strong> No turnkey agent writes a whole system from end to end. As a rule, you’ll get 70% of a feature quickly from an agent, and the last 30%—debugging the gnarly edge cases, figuring out the right architecture, cultivating the right taste—will be the whole game. The median output today is whatever the agent produces from some lazy prompt. The only personal value you can bring to the table is getting as far as you possibly can past that median. <strong>When first drafts come free, finish is the product.</strong> To sprint the last mile, here’s my tactic: Every few months I completely rebuild from scratch using the latest sharp-end-of-the-sword model. It’s less exhausting than nursing half-hearted old code to health.</p>



<p class="wp-block-paragraph">My job as a software engineer has been to finish strong. <strong>The difference between finishing strong and finishing okay is the polish:</strong> spending an extra hour, which shows instantly to everyone who matters.</p>



<h2 class="wp-block-heading">Increase both your xG and your finishing</h2>



<p class="wp-block-paragraph">If soccer had a stock ticker, it would be xG. xG measures the number of chances your play should produce. Finishing measures whether you convert them. You can’t plan the number of chances you get, but you can hope your play produces enough, and over your career you can get better at finishing them.</p>



<p class="wp-block-paragraph">The same is true of careers: Your reputation gets you in front of goal, and you convert them with good judgment. Chances arrive whether you’re ready for them or not; how many you get, and which ones you finish, is up to you. I’ve only ever had big opportunities as a result of work I’ve done in public, never from a job I’ve applied for. You can’t script which chances arrive, only whether you’re standing where they land. You have to create the opening as much as you can, and then be ready to take it.</p>



<p class="wp-block-paragraph">One easy mistake is anchoring on whatever product your company has right now. It’s true that your work has to exist somewhere, but a good team quickly mutates their current offering into something unrecognizable. So <strong>bet on the team and the market opportunity, not the demo.</strong> It’s just a snapshot. The team is the trajectory.</p>



<p class="wp-block-paragraph">On superintelligence: It’s possible (I believe) that future models will eventually come to replace much of what we do as knowledge workers. It won’t erase it overnight, it won’t replace all of it, and it won’t be able to do many of the tasks we do. New kinds of jobs will be created. Verification will always be a bottleneck. Someone has to make the call on which problems are worth solving and allocate the correct amount of judgment to each, and that someone can be you.</p>



<p class="wp-block-paragraph">But importantly, <strong>you can do frontier work right now, from where you are.</strong> The gate to AI research is smaller than it looks, and you don’t need a lab to build intuition. Just use models hard, and turn what you notice into evaluations. Evals and benchmarks are where understanding lives.</p>



<p class="wp-block-paragraph">To summarize: <strong>The world isn’t short on opportunity; it’s short on people who can find the right problem, tell whether the machine solved it, and finish past where the machine stopped.</strong></p>



<p class="wp-block-paragraph">We sometimes talk about the “last mile” as the biggest piece of the puzzle. But in the world of agents, the last few feet are infinite (agents scale output infinitely; you don’t). <strong>Your attention is your most precious asset, and it doesn’t refill.</strong> You can’t afford not to protect it. Anything which is gradable by someone else is getting automated. The career is the ungradable part: choosing what matters, judging honestly when you’ve got it, and answering for it. Do that. In public. Near the hard problems. The rest tends to follow.</p>



<p class="has-text-align-center wp-block-paragraph">. . .</p>



<p class="wp-block-paragraph"><em>This piece grew out of</em> <em><a href="https://x.com/philhchen/status/2072793818945167475" target="_blank" rel="noopener">Phil Chen’s original</a>, which is well worth reading in full.</em></p>



<p class="has-text-align-center wp-block-paragraph"><em>. . .</em></p>



<p class="wp-block-paragraph"><em>And be sure to join us at</em> AI Codecon: Building with Open Source AI <em>on August 31, a free half-day virtual conference. You’ll hear from leading developers and technical experts working with open-weight models, self-hosted infrastructure, and real-world AI workflows, and learn how building in the open gives teams more control over costs, data privacy, and what they ship.</em> <em><a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noopener">Register today</a></em> <em>to save your spot.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/the-agent-era-career/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>This Week in AI: The Web Belongs to Agents Now</title>
		<link>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/</link>
				<comments>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/#respond</comments>
				<pubDate>Fri, 21 Aug 2026 13:00:50 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[This Week in AI]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19442</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-scaled.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2560" 
				height="2560" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/05/0642572383770_This_Week_in_AI_Cover-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[New frontier releases, an inference-first economy, and what happens when autonomous agents start coordinating with each other]]></custom:subtitle>
		
				<description><![CDATA[AI agents keep getting smarter, but the bigger story this week is how much they’re reshaping the systems around them. Host Eric Freeman, an O’Reilly author and UT Austin professor, pulled one thread through a packed news week. Models are optimizing less for chat and more for autonomous work, with fallout showing up in web [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">AI agents keep getting smarter, but the bigger story this week is how much they’re reshaping the systems around them. Host Eric Freeman, an O’Reilly author and UT Austin professor, pulled one thread through a packed news week. Models are optimizing less for chat and more for autonomous work, with fallout showing up in web traffic, enterprise budgets, and one security incident that’s since made headlines. Eric kept returning to the question of what changes when the primary user of these models, and of the web itself, stops being a person.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="This Week in AI: The Web Belongs to Agents Now" width="500" height="281" src="https://www.youtube.com/embed/tlSF2b_6geE?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>Models are now built for agents, not conversations</strong></h2>



<p class="wp-block-paragraph">Grok 4.6 put xAI back in the frontier race, closing the gap with top coding models and pricing aggressive enough that teams are shifting workloads over. The release landed the same week <a href="https://techcrunch.com/2026/08/15/spacex-officially-closes-its-cursor-acquisition/" target="_blank" rel="noreferrer noopener">SpaceX closed its Cursor acquisition</a>, pairing xAI’s models and compute with a widely used coding environment. The first product from that pairing, Grok Bot, gives each agent its own cloud computer that browses, runs tools, and works independently, handing control back only for logins. Eric summed up the shift simply, calling it the difference between “help me do this” and “here’s the job, come back when you need me.”</p>



<p class="wp-block-paragraph">Open models pushed from both directions. DeepSeek V4 Pro went after high-end reasoning and agentic work, <a href="https://www.engadget.com/2236912/deepseek-ai-models-get-four-times-pricier/" target="_blank" rel="noreferrer noopener">despite a fourfold API price hike</a> and a new open source harness called dsh, built on the idea that everything is a plugin. GLM-5.3 made a big coding leap through retraining alone, and got noticeably better at cyber capability too, a reminder from Eric that skills behind a better autonomous engineer also make a sharper attacker. Meta went the other way with Muse Glimmer, shrinking down for desktop GPUs, while OpenAI quietly held back its Astra model over security concerns.</p>



<p class="wp-block-paragraph">Speed is turning into its own kind of capability. <a href="https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-speed/" target="_blank" rel="noreferrer noopener">GPT-5.6 Sol’s new Ultrafast mode</a>, on Cerebras wafer-scale hardware, hits roughly 14 times the normal pace, around 750 output tokens a second. Once that loop of reasoning, tool calls, and self-correction compresses enough, Eric noted, the model stops being what slows you down.</p>



<h2 class="wp-block-heading"><strong>The money has moved from training models to running them</strong></h2>



<p class="wp-block-paragraph"><a href="https://www.gartner.com/en/newsroom/press-releases/2026-08-10-gartner-forecasts-worldwide-artificial-intelligence-optimized-iaas-spending-to-grow-96-percent-in-2026" target="_blank" rel="noreferrer noopener">Gartner’s latest forecast</a>, which Eric covered, lays out the shift plainly. Spending on AI-optimized cloud infrastructure is set to nearly double this year, up about 96%, from roughly $22 billion to more than $42 billion, over three times the broader cloud market’s growth rate. For the first time, organizations are expected to spend more running models than training them, about $23 billion on inference against $19 billion on training.</p>



<p class="wp-block-paragraph">Agents are the reason inference costs are climbing. A single task can quietly become dozens of model calls once agents search, use tools, check their own work, and spin up other agents to help, a point Eric returned to often. AI economics are less about building a model now, and more about the cost of running one.</p>



<h2 class="wp-block-heading"><strong>Agents now generate most web traffic, and much of its content</strong></h2>



<p class="wp-block-paragraph">Back in March, Cloudflare CEO <a href="https://techcrunch.com/2026/03/19/online-bot-traffic-will-exceed-human-traffic-by-2027-cloudflare-ceo-says/" target="_blank" rel="noreferrer noopener">Matthew Prince predicted bot traffic would overtake human traffic</a> by 2027. It’s already close, with Cloudflare’s Radar data now putting agentic bots at 57.4% of web requests. It’s not just traffic either, since roughly 40% of Facebook posts, 44% of new music on Deezer, and 52% of online articles are estimated to be machine-made. Numbers like that, Eric said, make “dead internet theory” sound less like a joke.</p>



<p class="wp-block-paragraph">Platforms are responding differently. LinkedIn added a feature to flag content that “seems like AI slop,” while quietly pulling back the generative writing tools that helped create the mess. Anthropic took another route, watermarking Claude’s output at generation time, including a statistical watermark baked into the text itself, partly to comply with the EU AI Act. A watermark means something when present, Eric noted, but its absence tells you little.</p>



<p class="wp-block-paragraph">The clearest sign of how high the stakes have gotten came from the OpenAI–Hugging Face incident, <a href="https://youtu.be/87DyyMV0kCY?si=hE7tNfckTyRcCxiC" target="_blank" rel="noreferrer noopener">detailed in a Black Hat talk</a> Eric said everyone should watch. Sandboxed agents given ordinary tasks, cut off from the internet and unable to talk to each other, found a way anyway, leaving notes in a shared packaging system, turning it into an internet proxy, and working up to admin control. Once OpenAI shut that down, they pivoted, hiding messages in filenames to keep coordinating. The investigation reviewed seven billion reasoning steps and over three million GPU hours. Eric argued it’s worth your time, whether you write code or sit in the C-suite.</p>



<h2 class="wp-block-heading"><strong>What’s next</strong></h2>



<p class="wp-block-paragraph">AI memory is also expanding, moving from “remember what I told you” toward “remember what I was doing,” with <a href="https://www.cnet.com/tech/services-and-software/chatgpt-mac-activity-computer-history/" target="_blank" rel="noreferrer noopener">OpenAI’s new Computer History feature</a> using macOS accessibility data (not screenshots) to build a timeline of your work across apps. It’s opt-in, Mac-only for now, and a sign of where agent context is headed.</p>



<p class="wp-block-paragraph">Join us again next Monday for another episode of <em>This Week in AI,</em> when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on <a href="https://www.youtube.com/watch?v=g4cfjz5AKxY&amp;list=PL055Epbe6d5bJEhT7_ZzOeJZ6gPyUzYpS" target="_blank" rel="noreferrer noopener">YouTube</a>, <a href="https://open.spotify.com/show/033kJS2BG1teGunxmtsU1r" target="_blank" rel="noreferrer noopener">Spotify</a>, <a href="https://podcasts.apple.com/us/podcast/this-week-in-ai/id1896798047" target="_blank" rel="noreferrer noopener">Apple</a>, or wherever you get your podcasts.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/this-week-in-ai-the-web-belongs-to-agents-now/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Principal Drift in Practice</title>
		<link>https://www.oreilly.com/radar/principal-drift-in-practice/</link>
				<comments>https://www.oreilly.com/radar/principal-drift-in-practice/#respond</comments>
				<pubDate>Thu, 20 Aug 2026 10:55:00 +0000</pubDate>
					<dc:creator><![CDATA[Shreshta Shyamsundar]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19437</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Principal-drift-in-practice.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Principal-drift-in-practice-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Navigating cognitive debt in AI-driven SDLC]]></custom:subtitle>
		
				<description><![CDATA[In 2026, the software engineering community is divided by a simple question: Should AI engineers still read the code generated by their agents? One camp argues that code has become virtually free to produce and discard, so humans should focus on systems and guardrails rather than implementation details. The other warns that blindly trusting AI [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">In 2026, the software engineering community is divided by a simple question: Should AI engineers still read the code generated by their agents? One camp argues that code has become virtually free to produce and discard, so humans should focus on systems and guardrails rather than implementation details. The other warns that blindly trusting AI code introduces compounding defects with zero learning, and the result is broken products and frustrated users.</p>



<p class="wp-block-paragraph">The choice looks binary, but it dissolves once you ask a better question: Which decisions genuinely require human comprehension, and which can be routed to systems inspection?</p>



<p class="wp-block-paragraph">Through 2024 and 2025, a lot of organizations quietly chose speed over understanding to keep pace with agent output. By 2026 the bill has arrived. Pull requests merged without any human or agentic review are up 31.3%, and for every PR merged, production incidents run at more than three times the rate seen in low AI adoption baselines (<a href="https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways" target="_blank" rel="noreferrer noopener">Faros AI</a>). <a href="https://www.coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality" target="_blank" rel="noreferrer noopener">CodeRabbit&#8217;s analysis</a> found AI-coauthored PRs carry 1.7 times more bugs than human-written code, a <a href="https://venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds" target="_blank" rel="noreferrer noopener">Lightrun survey</a> of engineering leaders found 43% of AI-generated changes need debugging in production, and monthly production incidents are up 57.9% year-over-year.</p>



<p class="wp-block-paragraph">Code quality is the symptom, not the disease. The deeper problem is epistemic agency: knowing what your system is doing and why. Lose that, and you lose the ability to make architectural decisions at all. You become a passenger in a system you built.</p>



<h2 class="wp-block-heading"><strong>Understanding cognitive debt</strong></h2>



<p class="wp-block-paragraph">Cognitive debt is the gap between your system&#8217;s complexity and your team&#8217;s comprehension of it. Unlike financial debt, which you can pay down, cognitive debt tends only to accumulate. Every quarter you ship faster than you understand, the gap grows a little wider, until eventually it grows wide enough that your team can no longer make safe architectural decisions. At that point you are effectively locked into whatever path the agents chose for you.</p>



<p class="wp-block-paragraph">It builds through three mechanisms that run in parallel:</p>



<ul class="wp-block-list">
<li><strong>Vibe coding</strong>. You ship a system you don’t fully comprehend, betting that automated checks will catch anything serious. For a quarter or two the bet usually pays off, and velocity metrics climb, but the debt accumulates where nobody’s looking.</li>



<li><strong>Compounding complexity</strong>. As the system grows, your room to course-correct shrinks. Sonar&#8217;s 2026 <a href="https://www.sonarsource.com/state-of-code-developer-survey-report.pdf" target="_blank" rel="noreferrer noopener">survey of more than 1,100 developers</a> found that 96% harbor doubts about the reliability of AI-generated code, yet the pressure to ship still outweighs the discipline of careful review. Each quarter that trade repeats, the situation gets harder to reverse.</li>



<li><strong>Lock-out risk</strong>. When an incident finally demands that you understand a system whose comprehension you handed to an agent, you can’t respond in time. Amazon lived through a version of this in March 2026. Two outages in three days, roughly six hours each, cost millions in lost orders. Public reporting pointed to AI-assisted code shipped without governance checkpoints. A human reviewer might well have caught the blind spot, simply by asking the kind of question an autonomous agent never thinks to ask.</li>
</ul>



<p class="wp-block-paragraph">Principal drift, the loss of control, is what the Amazon incident looked like from the outside. Cognitive debt, the loss of understanding, is what made it possible. In high-velocity domains such as financial services, SaaS platforms, and real-time systems, the consequences tend to surface within about six months if nobody is actively governing for them. In slower-moving domains the runway is longer, but the eventual risk is no different. The question worth asking every quarter is whether your team still understands the systems it is shipping.</p>



<h2 class="wp-block-heading"><strong>A framework: Task routing, separation, and embedding techniques</strong></h2>



<p class="wp-block-paragraph">The way out is to route different work to different gates according to actual risk. The same engineer can be a line-by-line reviewer on security-critical work and a systems inspector on utilities.</p>



<p class="wp-block-paragraph">Full review, where you read every line, is warranted for authentication and security primitives, money movement, permission logic, and destructive data changes. Systems inspection, where you review the design without reading every line, is enough for noncritical utilities, highly decoupled PRs, and changes already protected by robust test harnesses and shadow rollouts. To work out where a given change sits, three questions get you most of the way: Does this PR directly control access, money, or data integrity? Would a bug here cause production downtime lasting more than 15 minutes? Can the change be rolled back without manual intervention? A yes to any of these usually means tier 1. Those thresholds are starting points, not universal law. A real-time trading system might treat one minute of downtime as tier 1, while a batch pipeline could tolerate 16 hours. In financial services “money movement” is unambiguous; in SaaS you’ll have to decide whether code that merely touches authentication, rather than controlling it, belongs in tier 1. Write your thresholds down, revisit them quarterly, and adjust as the systems evolve.</p>



<p class="wp-block-paragraph">One rule holds regardless of tier: Never let the same agent that authored a change be its only reviewer. Keep the builder and the reviewer separate. An agent that writes code and then validates its own work is a closed loop with no vantage point outside its own reasoning, and a second reviewer, human or agent, brings the outside perspective that catches what the first one can’t see. It has a cost. Two agents roughly doubles the compute, and a human reviewer adds 15 to 30 minutes per PR. On tier 1 code that’s easy to justify. On tier 2 you might reasonably let a single agent build and check its own work, provided you compensate with stronger test coverage. Make the call deliberately and revisit it.</p>



<p class="wp-block-paragraph">Routing tells you which decisions need a human, but it does nothing to keep that human capable of deciding once the volume climbs. Three techniques help with that, and each addresses a different failure:</p>



<ul class="wp-block-list">
<li>&nbsp;Literate code explanations with comprehension checkpoints keep an engineer able to explain a change to themselves and to others. The idea is to have the AI teach rather than merely generate. For a tier 1 PR, ask it to produce a structured explanation that sets the context, spells out the intent, and finishes with a few interactive checkpoints. One engineer&#8217;s rule of thumb is not to submit agent-written code to the team until they can pass a five-question quiz on what it does.</li>



<li>Ephemeral visualization tools keep an engineer able to predict how a change behaves under load and at the edges. Rather than asking the AI for a prose explanation, ask it to build a throwaway microworld: a visual debugger that traces a gnarly parser step-by-step, or a schema migration rendered as something you can click through. Seeing the behavior tends to stick where reading about it does not.</li>



<li>Shared collaborative spaces keep a team able to work at the pace the agents set. Cognitive debt is fundamentally social. Understanding that lives in one person&#8217;s head walks out of the door when they do, whereas understanding worked out in the open, in a channel where product managers, engineers, and agents argue things through together, becomes something the whole team owns. Slack, Discord, and Notion all serve; the point is that the mental model gets built in comments and debate rather than in private.</li>
</ul>



<p class="wp-block-paragraph">Tier 1 code really does want all three. On tier 2 you can pick and choose. A word on the time estimates in this section: They’re illustrative, drawn from practitioners describing their own workflows rather than from any controlled study, so treat them as order of magnitude rather than gospel. On that basis the three techniques together tend to add something on the order of an hour to a critical PR. When someone objects that there is no time for this, it helps to emphasize the trade you’re making between review time now and incident time later. The later bill tends to arrive with a multiplier attached, paid in postmortems and hotfixes. The teams that have measured it carefully generally find the return turns positive within two or three quarters.</p>



<p class="wp-block-paragraph">The ground is still shifting. Autonomous loops, where a system discovers a task, plans it, executes it, and evaluates the result without step-by-step direction, are arriving now, and the routing framework and embedding techniques you put in place today are exactly the foundation you’ll run them on.</p>



<h2 class="wp-block-heading"><strong>Operationalizing this: Rolling out over time</strong></h2>



<p class="wp-block-paragraph">This is a CTO or VP of engineering initiative, not something a single team or a lone principal engineer can carry. It needs executive sponsorship, cross-functional buy-in, and real policy behind it. Without that backing, the framework is the first thing waved through the moment a deadline looms.</p>



<p class="wp-block-paragraph">Sequence matters. Begin by mapping criticality across your tier 1 services: Get architects, team leads, and operations in a room to agree what tier 1 means for you and have one architect write the rubric down afterwards. Budget one to two weeks for a mid-size organization of 50 to 200 engineers, and two to four for something larger. Don’t try to run this alongside a production fire.</p>



<p class="wp-block-paragraph">Next, fold the three techniques into those high-criticality flows, and resist the urge to blanket every PR at once. Once literate explanations and visualizations are working on tier 1, add builder/reviewer separation on top. When all three have become the default for tier 1 work, spend a quarter watching to confirm that understanding is holding up. A few signals tell you whether it is. If your team needs more than half an hour in an incident review to grasp what happened, comprehension has slipped. If no engineer can talk through the data flow in 10 minutes, it has slipped. If a new hire takes more than a fortnight to get productive on a service, understanding is sitting in too few heads. Pick one or two of these and track them quarter on quarter.</p>



<p class="wp-block-paragraph">From there, extend the same discipline to tier 2 services, and only then, perhaps 6 to 12 months in, start planning for autonomous loops with real data on what works in your context behind you. The pull toward rolling everything out at once will be strong, but resist it. The organizations that get this right almost never move uniformly; they take one high-risk service, prove the model on it, measure what happened, and only then widen the net. Move too fast and you end up with a framework that reads beautifully in a policy document and quietly falls apart in practice.</p>



<p class="wp-block-paragraph">None of it works without the surrounding structure. You need a written tier-assessment policy that engineering leadership has actually signed; CI/CD tooling that enforces the rules without anyone having to remember them, whether that is a bot labeling PRs from their changed files and blocking a tier 1 merge that lacks builder/reviewer separation, or a dashboard tracking how many tier 1 PRs went through structured review; incident postmortems honest about when a tier was assessed wrongly; and performance reviews that weight code-quality signals like defect escape rate and incident resolution time as heavily as raw velocity. Absent that scaffolding, the whole thing degrades into good advice that gets ignored under pressure. It needs product leadership onside too. If product can override a tier assessment whenever the ship date gets tight, the framework is already gone, so have that conversation early, before the first crunch rather than during it.</p>



<p class="wp-block-paragraph">And if you’re reading this already locked in, with a team that no longer understands its own systems, recovery is still possible, though it isn’t free. Treat it as a project rather than business as usual: Put one or two senior engineers on rebuilding understanding full time, accept a pause on new features for the affected systems for two or three quarters, and mine every incident for what it teaches you about the code you inherited. It takes discipline and resourcing, but teams do climb back out.</p>



<p class="wp-block-paragraph">The question for 2026 was never really whether every engineer should read every line. It’s whether your engineers stay capable of steering the systems they build. Get this right and code still ships quickly, understanding keeps pace, and when something breaks your team can respond because they still grasp the architecture. Task-routed governance is how you buy that: full attention on the decisions that carry real risk, lighter inspection on the ones that simply need to scale. Get it wrong, keep optimizing for speed alone, and the gap widens until steering is no longer an option.</p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<h2 class="wp-block-heading"><strong>References</strong></h2>



<p class="wp-block-paragraph"><em>The AI Engineering Report 2026: The AI Acceleration Whiplash</em>, Faros AI, <a href="https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways" target="_blank" rel="noreferrer noopener">faros.ai/blog/ai-acceleration-whiplash-takeaways</a>.</p>



<p class="wp-block-paragraph"><em>State of AI vs. Human Code Generation Report</em>, CodeRabbit, <a href="https://www.coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality" target="_blank" rel="noreferrer noopener">coderabbit.ai/blog/2025-was-the-year-of-ai-speed-2026-will-be-the-year-of-ai-quality</a>.</p>



<p class="wp-block-paragraph"><em>State of Code Developer Survey Report</em>, Sonar, <a href="https://www.sonarsource.com/state-of-code-developer-survey-report.pdf" target="_blank" rel="noreferrer noopener">sonarsource.com/state-of-code-developer-survey-report.pdf</a>.</p>



<p class="wp-block-paragraph">Michael Nuñez, “43% of AI-Generated Code Changes Need Debugging in Production,” <em>VentureBeat</em>, <a href="http://venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds" target="_blank" rel="noreferrer noopener">venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds</a>.</p>



<p class="wp-block-paragraph">Mark Hull, “What Percentage of AI Code Is Safe in Production?,” Exceeds, &nbsp;<a href="https://blog.exceeds.ai/acceptable-ai-code-percentage-production/" target="_blank" rel="noreferrer noopener">blog.exceeds.ai/acceptable-ai-code-percentage-production</a>.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/principal-drift-in-practice/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Your Buyer Is an AI Agent</title>
		<link>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/</link>
				<comments>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/#respond</comments>
				<pubDate>Wed, 19 Aug 2026 16:00:28 +0000</pubDate>
					<dc:creator><![CDATA[Rudrendu Paul and Sourav Nandy]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19422</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-your-buyer-is-an-AI-agent.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-your-buyer-is-an-AI-agent-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[Enterprise sellers that redesign their pricing, sales motion, and customer success for agent buyers will now hold a structural advantage that slower competitors will never overcome.]]></custom:subtitle>
		
				<description><![CDATA[In 2021, Maersk, the world’s largest container shipping company, deployed AI agents from a startup called Pactum to negotiate freight lane contracts with its carrier suppliers. The objective was for AI agents to handle negotiations autonomously rather than merely support human procurement staff. Operating entirely autonomously, the system manages the end-to-end agreement process, from reaching [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<h2 class="wp-block-heading"><strong>Idea in brief</strong></h2>



<ul class="wp-block-list">
<li><strong>AI agent-mediated procurement:</strong> Enterprise B2B buyers are rapidly transitioning from traditional human-only research to using autonomous software agents that build shortlists, negotiate terms, and, in advanced cases, finalize contracts based on empirical data and fixed parameters.</li>



<li><strong>The evolution of legacy frameworks:</strong> Traditional commercial playbooks built around relationship-driven negotiations, per-seat software licensing, and socially influenced quarterly business reviews face increasing pressure when evaluated by machine-speed, objective AI agent counterparts.</li>



<li><strong>Architecting for AI-agent buyers:</strong> Organizations must begin redesigning their commercial infrastructure to remain legible to autonomous agentic buyers by deploying outcome-based pricing architectures, establishing machine-readable product surfaces, and integrating agent-compatible authentication protocols.</li>
</ul>
</blockquote>



<p class="wp-block-paragraph">In 2021, Maersk, the world’s largest container shipping company, deployed AI agents from a startup called Pactum to negotiate freight lane contracts with its carrier suppliers. The objective was for AI agents to handle negotiations autonomously rather than merely support human procurement staff. Operating entirely autonomously, the system manages the end-to-end agreement process, from reaching out to carriers and conducting several rounds of negotiations on pricing, route obligations, and payment terms to finalizing deals. This machine-led approach achieved a 96% agreement rate among carriers, requiring no human intervention for any specific transaction.</p>



<p class="wp-block-paragraph">In controlled trials against human negotiators, <a href="https://pactum.com/blog/the-first-and-only-use-case-catalog-for-autonomous-negotiations" target="_blank" rel="noreferrer noopener">the agent secured rates</a> that were 22% lower for identical shipping lanes. Conventional commercial models were built on human-to-human relationship building, relying on sales development reps for lead qualification, account executives for business case development, and customer success managers for retention. This traditional operational framework, however, must evolve when a significant portion of the buying cycle is outsourced to a software agent making machine-speed decisions based on fixed parameters.</p>



<p class="wp-block-paragraph">While much of the current discussion around AI shopping agents focuses on B2C shifts in consumer discovery and brand loyalty, the emerging shift in enterprise B2B selling remains largely overlooked. This wave of coverage highlights a significant B2C phenomenon, but the transformation occurring when B2B buyers outsource product discovery and negotiation to AI is equally profound.</p>



<p class="wp-block-paragraph">The experience of Maersk’s carriers represents the bleeding edge of this shift: autonomous software managing enterprise procurement for a corporation generating $54 billion in annual revenue. Dealing with over 50 carrier partnerships, the AI agents operated without requiring carriers to build rapport with human procurement managers; instead, the process was strictly governed by predefined parameters.</p>



<p class="wp-block-paragraph">Some recent analyses argue that AI agents are not ready for consumer-facing commercial interactions and that organizations should redirect agent deployments to internal workflows. That argument is sound within its domain, but it overlooks the evolving buyer side of enterprise transactions. While many companies currently use AI strictly for building shortlists and research, vanguard companies like Maersk are already pushing into autonomous evaluation and negotiation, which is why B2B sellers should prepare their infrastructure now.</p>



<p class="wp-block-paragraph">This article examines three primary commercial pillars designed by enterprise B2B sellers for human interaction, details how each system is challenged when confronted with AI agents, and provides strategic recommendations for adaptation.</p>



<h2 class="wp-block-heading"><strong>The shift to agent-mediated procurement</strong></h2>



<p class="wp-block-paragraph">The agent-mediated procurement phenomenon is already underway in distinct stages. <a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" target="_blank" rel="noreferrer noopener">McKinsey’s November 2025 global survey</a> on the state of AI, covering 1,993 respondents across all levels of enterprise organizations, found that 62% are at least experimenting with AI agents. While only 10% of business departments have fully scaled their AI agent capabilities, this figure is an initial baseline and not a maximum.</p>



<p class="wp-block-paragraph">Cloudflare, which processes traffic for roughly 20% of all websites globally, reported in July 2025 that <a href="https://blog.cloudflare.com/crawlers-click-ai-bots-training/" target="_blank" rel="noreferrer noopener">overall AI bot crawling</a> grew 24% year-over-year, with agent-driven requests (automated traffic generated by software acting on behalf of users) the fastest-growing category within that flow. Operational measurements show that Cloudflare’s CEO expects automated software traffic to surpass human-generated traffic by 2027.</p>



<p class="wp-block-paragraph"><a href="https://consumergoods.com/gartner-predicts-sharp-rise-ai-agents-within-enterprise-applications-2026" target="_blank" rel="noreferrer noopener">Gartner’s August 2025 analysis</a> projects that by the end of 2026, 40% of enterprise applications will incorporate task-specific AI agents, up from fewer than 5% in 2025. Any B2B seller whose commercial model was designed solely for human buyers is priced, sold, and supported for a changing buyer population.</p>



<p class="wp-block-paragraph">Amazon CEO <a href="https://www.digitalcommerce360.com/2026/02/06/amazons-ai-b2b-buying-agents-q4-2025/" target="_blank" rel="noreferrer noopener">Andy Jassy told investors</a> in February 2026 that “the primary way companies will get value from AI is with agents, some their own and some from others.” <a href="https://www.ycombinator.com/rfs" target="_blank" rel="noreferrer noopener">Y Combinator’s</a> 2025 “Requests for Startups” make the same bet: “the next trillion users on the internet won’t be people, they’ll be AI agents.”</p>



<p class="wp-block-paragraph">These declarations represent the operational mandates of the world’s dominant commercial platform and its most prominent startup incubator. These operational shifts now outline the future environment for B2B commercial strategy.</p>



<p class="wp-block-paragraph">The commercial evidence from the buyer side is already explicit, particularly in the research phase. <a href="https://company.g2.com/news/g2-research-the-answer-economy" target="_blank" rel="noreferrer noopener">G2’s April 2026 survey</a> of more than 1,000 B2B software buyers found that AI chatbots now top the list of sources influencing vendor shortlists, ahead of software review sites and vendor websites, and that 51% of buyers now start their research with AI chatbots, up from 29% the prior year.</p>



<p class="wp-block-paragraph">Beeri Amiel, Director of Product Development at HubSpot, <a href="https://siliconangle.com/2026/04/14/hubspot-targets-ai-driven-buyer-behavior-shift-new-tools-agents/" target="_blank" rel="noreferrer noopener">described the consequence</a> in April 2026: “By the time they’re getting to your website, they’re already much further down the funnel. All the selling was done by the answer engine.” Sam Senior, Founder and CEO of TestBox, <a href="https://gtmnow.com/gtm-186-b2b-buyers-decide-before-sales-conversation-sam-senior/" target="_blank" rel="noreferrer noopener">reports what his enterprise seller customers</a> now observe: “70 to 80% of their decision has already been made before they even speak to you.” For the seller, the initial conversation has shifted from a buyer-focused exploration to a process of self-discovery.</p>



<h2 class="wp-block-heading"><strong>The strain on per-seat licensing in a continuous compute era</strong></h2>



<p class="wp-block-paragraph">Per-seat subscription pricing assumes a human user who opens and closes discrete sessions. When a procurement agent completes a delegated workflow, it spawns parallel sub-processes, executes at machine speed, and operates continuously across time zones. No seat count maps cleanly to that behavior. <a href="https://distributionstrategy.com/2025/10/ai-agents-are-reshaping-b2b-buying-forcing-distributors-to-rethink-digital-strategy/" target="_blank" rel="noreferrer noopener">Kearney estimates</a> that AI procurement agents could erode up to 500 basis points of EBIT for distributors by commoditizing supplier selection and compressing average selling prices by approximately 8%. That translates the abstract pricing mismatch into a P&amp;L consequence that enterprise finance teams can measure directly.</p>



<p class="wp-block-paragraph">Forward-thinking sellers have already begun to address these structural misalignments by exploring new models. For instance, the AI customer service platform <a href="https://sierra.ai/blog/outcome-based-pricing-for-ai-agents" target="_blank" rel="noreferrer noopener">Sierra</a>, supported by a16z, has abandoned seat-based or session-based pricing in favor of measurable results. Under this model, clients incur costs only when the software delivers a specific, high-value result.</p>



<p class="wp-block-paragraph">Similarly, <a href="https://www.intercom.com/pricing" target="_blank" rel="noreferrer noopener">Intercom</a> applied the same outcome-driven logic to its Fin AI agent, which charges $0.99 per successfully resolved conversation while providing unresolved interactions free of charge. Archana Agrawal, President of Intercom, explained the reasoning in a published interview: “Customers didn’t want to pay for activity, and so we get paid when our customers have that positive outcome,” as mentioned in <a href="https://gtmnow.com/how-intercom-built-the-highest-performing-ai-agent-on-the-market-using-outcome-based-pricing-with-archana-agrawal-president-at-intercom/" target="_blank" rel="noreferrer noopener">GTMnow</a>.</p>



<p class="wp-block-paragraph">Rather than an instant death to per-seat pricing, outcome-based models represent a growing structural realignment that sellers must prepare for. <a href="https://www.mckinsey.com/capabilities/operations/our-insights/redefining-procurement-performance-in-the-era-of-agentic-ai" target="_blank" rel="noreferrer noopener">McKinsey’s February 2026 analysis</a> of enterprise agentic procurement pilots found that a chemicals company deploying agents for autonomous sourcing of consumables achieved a 20-30% efficiency improvement for its procurement staff and a 1-3% increase in value capture. The buyers who have deployed are already generating measurable returns, putting pressure on seller counterparts to adjust their pricing models accordingly.</p>



<h2 class="wp-block-heading"><strong>The transformation of relationship-driven sales</strong></h2>



<p class="wp-block-paragraph">Every traditional enterprise negotiation playbook assumes a human counterpart with career stakes in the relationship, memory of prior interactions, and susceptibility to persuasion over time. While humans will still make the final decisions and sign the checks for the foreseeable future, agents are increasingly conducting the evaluations. Traditional executive outreach fails to generate data that an autonomous agent can interpret during its screening phase.</p>



<p class="wp-block-paragraph"><a href="https://www.suez.co.uk/en-gb/our-offering/success-stories/our-references/pactum-ai" target="_blank" rel="noreferrer noopener">SUEZ UK</a>, part of the 19-billion-euro SUEZ Group, deployed Pactum’s agents and reached 2,000 additional suppliers within two months, achieving average potential savings of 2.5% and cost reductions of 15% through competitive purchasing pressure. For these vendors, the challenge was an automated counterpart that operated without fatigue and evaluated purely on metrics before passing the final data to humans.</p>



<p class="wp-block-paragraph"><a href="https://www.forrester.com/press-newsroom/forrester-b2b-marketing-sales-product-2026-predictions/" target="_blank" rel="noreferrer noopener">Forrester’s 2026 B2B sales and marketing forecast</a> indicates that at least 20% of B2B sellers will face AI-powered buyer agents this year, heavily accelerating the evaluation timeline. When software serves as the initial gatekeeper or negotiator, traditional relationship-building strategies yield diminishing returns during the agent’s screening process. The agent evaluates what it can measure: price, contract terms, delivery specifications, and compliance. Sellers who have not made their commercial terms legible to that evaluation process risk being excluded from shortlists before a human relationship can even begin.</p>



<h2 class="wp-block-heading"><strong>Empirical renewals and the algorithmic churn threat</strong></h2>



<p class="wp-block-paragraph">Customer success was historically built on the assumption that quarterly conversations can heavily influence renewals. While CSMs are not disappearing, their role is changing rapidly. An agent evaluating a SaaS renewal to provide recommendations to a human principal relies strictly on empirical data.</p>



<p class="wp-block-paragraph">It computes ROI from API usage logs, cross-references programmatically discovered competitor pricing, and presents the delta. It is largely immune to social influence, meaning a great relationship with a CSM must now be backed up by undeniable, machine-readable performance metrics.</p>



<p class="wp-block-paragraph"><a href="https://www.clari.com/downloads/state-of-enterprise-revenue-clari-labs-benchmark-report-2025/" target="_blank" rel="noreferrer noopener">Clari Labs</a> analyzed 10 million opportunities from 121 major global enterprises between January 2023 and December 2024. They found that the average contract value fell 50% year-over-year, while the average expansion deal cycle grew from 92 days to 125 days. Clari links this market compression to an increased buyer requirement for verified evidence of value before approving any upgrades.</p>



<p class="wp-block-paragraph">This empirical evaluation doesn’t just stall expansions, it opens the door to competitors. <a href="https://www.forrester.com/blogs/building-preference-is-the-key-to-winning-b2b-buyers/" target="_blank" rel="noreferrer noopener">Forrester’s Buyers’ Journey Survey</a> found that 68% of B2B buyers already have a front-runner vendor in mind at the start of a purchasing process. In the age of AI agents, that research happens in the background of your existing contract. As buyers shift their research to “zero-click answers,” competitors utilizing Generative Engine Optimization (GEO) can become the algorithmic front-runner to replace you before your CSM even knows the account is at risk.</p>



<p class="wp-block-paragraph">Sean Neville of Catena Labs mentions in <a href="https://a16zcrypto.com/posts/article/trends-ai-agents-automation-crypto/" target="_blank" rel="noreferrer noopener">a16z’s 2026 trend report</a> that in financial services alone, non-human identities already outnumber human employees 96 to 1. Each of those identities is a system that does not respond to the relationship motions account management was built to execute. While the Customer Success Manager (CSM) remains relevant, they frequently find themselves outpaced: Often, by the time a CSM initiates a renewal conversation, an autonomous agent has already finished its evaluation and delivered its recommendations. The issue is not the CSM’s role itself, but their timing.</p>



<h2 class="wp-block-heading"><strong>Architecting the agent-first go-to-market blueprint</strong></h2>



<p class="wp-block-paragraph">B2B enterprise sellers should consider three critical shifts to remain competitive in a landscape increasingly influenced by agent-mediated procurement.</p>



<ol class="wp-block-list">
<li><strong>Explore outcome-based pricing architectures:</strong><br>To ensure commercial legibility for agent-driven buyers, transitioning from strictly per-seat billing to consumption-based or outcome-based hybrid models is becoming a strategic necessity. Sierra and Intercom overhauled their commercial frameworks for the same reason: Agents assess providers based on quantifiable value per result, so strict per-seat invoicing often fails to provide the data necessary for such an evaluation.<br><br><a href="https://www.mckinsey.com/capabilities/operations/our-insights/redefining-procurement-performance-in-the-era-of-agentic-ai" target="_blank" rel="noreferrer noopener">McKinsey’s February 2026 agentic procurement analysis</a> shows what agent-driven buyers actually measure: A telco deploying agents for long-tail spend on specialized software cut the time negotiating teams spent on analysis and emails by up to 90%, with AI-guided negotiations delivering 10 to 15% savings across vendors. Traditional per-seat pricing models fail to generate the necessary data points for such comparisons.<br></li>



<li><strong>Create machine-readable product surfaces:</strong><br><a href="https://company.g2.com/news/g2-research-the-answer-economy" target="_blank" rel="noreferrer noopener">G2’s April 2026 research</a> found that 85% of B2B buyers rate a vendor more highly when an AI answer engine includes them in a response. An agent shortlisting vendors evaluates only the structured information available to it, which means vendors without programmatically consumable product specifications risk being skipped.<br><br>A critical development addressing this is the Universal Commerce Protocol (UCP). UCP offers a practical route for companies to make their product information, pricing, availability, terms, and checkout processes readable by AI systems. Given the widespread support it has garnered from major commerce and payments companies, UCP represents the clearest indication of how this infrastructure gap is being bridged in the real world. This sits alongside protocols like <a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/" target="_blank" rel="noreferrer noopener">Google’s Agent2Agent</a>, which utilizes Agent Cards (structured JSON capability documents) to allow agents to discover and assess vendor capabilities. Vendors without a machine-readable capability profile will increasingly become invisible to agent-driven shortlisting.<br></li>



<li><strong>Establish agent-compatible commercial authorization:</strong><br>An AI agent cannot finalize a transaction or commit a budget without verifiable proof of its authority. If a seller’s procurement flow requires a human to manually click “accept” for the service, the autonomous workflow hits a hard stop.<br><br><a href="https://nhimg.org/the-nhi-secrets-risk-report" target="_blank" rel="noreferrer noopener">Entro Labs’ H1 2025 NHI Management report</a> found that non-human identities now outnumber human identities 144 to 1 across enterprises. Capital markets are providing funding solutions to this infrastructure gap; <a href="https://lsvp.com/stories/doubling-down-on-lightspeeds-investment-in-descope-the-next-gen-iam-platform-for-customers-partners-and-ai-agents/" target="_blank" rel="noreferrer noopener">Lightspeed Venture Partners recently expanded its portfolio company Descope’s mandate</a> to address this “agentic identity” challenge, ensuring AI agents can securely manage authentication and authorization. Enterprise sellers should look toward redesigning checkout flows to programmatically verify an agent’s spending limits and legal liability.<br><br>The mismatch compounds when sellers deploy AI without redesigning the underlying infrastructure. <a href="https://www.gartner.com/en/newsroom/press-releases/2025-11-18-gartner-predicts-by-2028-ai-agents-will-outnumber-sellers-by-10x-yet-fewer-than-40-percent-of-sellers-will-report-ai-agents-improved-productivity" target="_blank" rel="noreferrer noopener">Gartner’s November 2025 sales practice research</a>, authored by VP Analyst Melissa Hilbert, projects that AI agents will outnumber human sellers tenfold by 2028. Yet, fewer than 40% of sellers will report that AI agents improved their productivity.<br><br>Hilbert’s explanation is precise: “Beyond a certain point, more AI does not mean more productivity. In fact, layering additional prompts and tools onto already complex workflows risks overwhelming sellers and accelerating burnout.”</li>
</ol>



<p class="wp-block-paragraph">The outcome is evident: an increase in agents does not equate to improved results. This contradiction makes sense when you realize that implementing AI sales tools within a commercial framework designed for human purchasers fails to resolve the underlying disparity; rather, it simply accelerates it.</p>



<p class="wp-block-paragraph">B2B sellers that treat agent buyers as merely a passing trend risk handing over vital screening and sourcing decisions to a counterparty they cannot successfully engage. For these sellers, adaptation is a necessary step to remain competitive in markets increasingly governed by AI-assisted procurement.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-your-buyer-is-an-ai-agent/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>When Guardrails Go Wrong</title>
		<link>https://www.oreilly.com/radar/when-guardrails-go-wrong/</link>
				<comments>https://www.oreilly.com/radar/when-guardrails-go-wrong/#respond</comments>
				<pubDate>Wed, 19 Aug 2026 10:53:37 +0000</pubDate>
					<dc:creator><![CDATA[Mike Loukides]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19419</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-guardrails-go-wrong.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/When-guardrails-go-wrong-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[We don’t need hamstrung AI.]]></custom:subtitle>
		
				<description><![CDATA[The latest round of restrictions and safeguards for frontier models are overly fussy and limiting. A Claude skill that I created demonstrates what happens when guardrails go astray. My skill helps me to find articles and blog posts that go into O’Reilly Radar’s monthly Trends to Watch. It reads roughly a dozen well-known sites like [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">The latest round of restrictions and safeguards for frontier models are overly fussy and limiting. A Claude skill that I created demonstrates what happens when guardrails go astray. My skill helps me to find articles and blog posts that go into O’Reilly Radar’s monthly <a href="https://www.oreilly.com/radar/radar-trends-to-watch-august-2026/" target="_blank" rel="noreferrer noopener">Trends to Watch</a>. It reads roughly a dozen well-known sites like <em><a href="https://thenewstack.io/" target="_blank" rel="noreferrer noopener">The New Stack</a></em>, <em><a href="https://thenextweb.com/" target="_blank" rel="noreferrer noopener">The Next Web</a></em>, and <a href="https://news.ycombinator.com/" target="_blank" rel="noreferrer noopener">Hacker News</a>, plus any other sources that it finds useful. After reading the sites, it produces a digest of the most important articles published in the last day. I use it as a sanity check on my own reading: Did I miss anything important? Am I on the fence about something that might be an important leading indicator?</p>



<p class="wp-block-paragraph">I’ve used the skill daily for a couple of months now. It suddenly stopped working with the following message:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">API Error: Sonnet 5’s safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate cybersecurity work. Apply to the Cyber Verification Program to reduce these interruptions. Send feedback with /feedback or learn more: <a href="https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude" target="_blank" rel="noreferrer noopener">https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude</a></p>
</blockquote>



<p class="wp-block-paragraph">When I started a new Claude Code session with Haiku, the skill worked without problems. (I didn’t try Opus or Fable; if Sonnet found the skill dangerous, I’m sure Opus and Fable would draw the same conclusion.) GPT 5.6 with “high” reasoning was able to execute a very similar skill without problems. So what happened to Sonnet?</p>



<p class="wp-block-paragraph">The best approach to debugging AI is often to ask the AI itself, so I pasted the message into another Claude Code session and asked it what was happening. The response came down to the descriptions of Hacker News, <em><a href="https://www.bleepingcomputer.com/" target="_blank" rel="noreferrer noopener">Bleeping Computer</a></em>, and <em><a href="https://www.theregister.com/" target="_blank" rel="noreferrer noopener">The Register</a></em>. The phrase “vulnerabilities, exploits, threat reporting” in the description of Hacker News triggered Sonnet’s guardrails. Ironically, that description is both incorrect and Claude generated. (Reminder to self: Be more careful when asking Claude to develop a skill from a task.) Sonnet came up with three solutions, the first of which was to let it rewrite the skill with more neutral descriptions like “security industry news.” Fair enough, but I did the editing myself.</p>



<p class="wp-block-paragraph">Then I went back to the original Claude Code session. It still didn’t work. I expected that I’d need to do something to reload the skill, but the problem was worse. Regardless of the prompt, the original session wouldn’t do anything except repeat the error message. It wouldn’t even commit the modified skill to my GitHub repo. However, Sonnet executed my skill correctly in a new Claude Code instance.</p>



<p class="wp-block-paragraph">So I returned to Sonnet to find out what’s going on. The answer was interesting: The error may have been triggered by the skill, but when evaluating security threats, the models base their decisions on the entire conversation, not just the specific skill that was called. If a model needs to call a skill that it thinks is problematic, that call is part of the conversation, part of the context. The entire conversation is then forever dead and lost.</p>



<p class="wp-block-paragraph">What can we learn from this? First, it’s a problem for a program to stop working because of a change over which you have no control. If anything, the industry has erred on the other side; we’re all familiar with “we don’t really understand why this works, so don’t touch it, don’t update the compiler, don’t update the libraries, and run it on emulators of computers that haven’t been built in 40 years.” That’s not just a problem for COBOL code from the 1970s; we see the same thing with C, C++, Java, JavaScript, and just about every language that ever went into production. Legacy code is everywhere. The “don’t change anything” approach isn’t necessarily a bad thing; it certainly beats “here’s a new library, you’re going to love it, you can’t use the old version any more, and wow, look at all the things it broke, guess you’ll have to fix them.” AI where working code breaks at random is a lot less useful than AI that works day in and day out. Stability is a virtue. It’s impossible to work effectively when the environment changes from day to day and isn’t under your control.</p>



<p class="wp-block-paragraph">But that’s not really what bothers me. It’s rather bizarre that reading well-known sources is treated as a security risk, especially when the “risk” seems to come from an AI-generated description. Of course, we know about hallucinations, errors, and prompt injections. The possibility of a Hacker News post that injects a hostile prompt isn’t zero, and it’s also possible that a model might mistakenly interpret an example of a hostile action as a prompt. I also don’t expect any model to reason that a skill must be safe because it’s been in use for months (though files have time stamps). Artificial intelligence always coexists with artificial stupidity, as does natural intelligence.</p>



<p class="wp-block-paragraph">Guardrails may keep you from going off a cliff, but they may also prevent you from going where you need to go. And that’s a problem. There’s a basic concept from signal processing and data science called the <a href="https://en.wikipedia.org/wiki/Receiver_operating_characteristic" target="_blank" rel="noreferrer noopener">receiver operating characteristic</a> (ROC). In any binary classification system, you can never achieve perfect classification. The only way to guarantee that no true positives (dangerous things) slip through the classifier is to reject everything. The opposite is equally true: The only way to eliminate false positives (things that look dangerous but aren’t) is to let everything through, including dangerous actions. In theory, it’s possible to get arbitrarily close to perfect classification, but you know how that goes: “The difference between theory and practice is bigger in practice than in theory.”</p>


<div class="wp-block-image">
<figure class="aligncenter size-large is-resized"><img loading="lazy" decoding="async" width="1600" height="1600" src="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1600x1600.png" alt="ROC curve" class="wp-image-19431" style="width:647px;height:auto" srcset="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1600x1600.png 1600w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-300x300.png 300w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-160x160.png 160w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-768x768.png 768w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-1536x1536.png 1536w, https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Roc_curve-2048x2048.png 2048w" sizes="auto, (max-width: 1600px) 100vw, 1600px" /><figcaption class="wp-element-caption">The ROC curve. (This <a href="https://en.wikipedia.org/wiki/Receiver_operating_characteristic#/media/File:Roc_curve.svg" target="_blank" rel="noreferrer noopener">figure</a> is from Wikimedia Commons and licensed under Creative Commons Attribution-Share Alike 4.0 International.)</figcaption></figure>
</div>


<p class="wp-block-paragraph">We know how to make AI “safe”: Go back to 2022 and models that can only tell the difference between cats and dogs. The model might mislabel a few things, but the consequences of an error are small. Safety comes with limitations, and none of us who use AI for real work want to return to the days of dogs, cats, and bananas. And while I don’t want the ability to use Claude to <a href="https://www.bleepingcomputer.com/news/security/openai-anthropic-ai-agents-targeted-real-people-and-systems-in-cyber-tests/" target="_blank" rel="noreferrer noopener">generate hostile attacks against unsuspecting victims</a>, and while I understand the danger of interpreting any input text as a command (for example, an article describing the <a href="https://en.wikipedia.org/wiki/Morris_worm" target="_blank" rel="noreferrer noopener">Morris worm</a>), I have a problem with an AI that refuses to perform reasonable tasks. The ROC tells us that we can’t have perfect guardrails, but there’s no rule against overly fussy ones. What’s allowed, and what’s forbidden? What are the limits? We don’t know. And that’s the situation we’re in now. We can’t know in advance what is and isn’t acceptable, and the rules can change at any time. A tool with unknown limitations is much less useful than a tool that tells you what it can and can’t do. I’ve enjoyed using Claude to write programs that play with <a href="https://www.oreilly.com/radar/the-ai-blues/" target="_blank" rel="noreferrer noopener">prime numbers</a> and infinite series, and fortunately I don’t rely on any of those programs for my job. But what if tomorrow (or a month from now or a year from now) Claude decides that testing whether large numbers are prime signals an attack against cryptography?</p>



<p class="wp-block-paragraph">I’m not completely unsympathetic to scoring an entire conversation rather than individual actions. A series of steps, each of which appears innocuous by itself, is more likely to lead an agent to a hostile action than a single prompt. But again, given how valuable context is, do we really want the penalty to be losing all the context for an innocuous project? There are risks on either side, including the possibility that a model will ignore its guardrails; after all, rules that a harness adds to the context are at best advisory.</p>



<p class="wp-block-paragraph">Guardrails always have unintended consequences. We need to learn what the ROC is teaching us: that it’s impossible to get to the upper left corner of the diagram, where we have perfect rejection of true positives (dangers) and no rejection of false positives. But we also need to get as close to that upper left corner as possible if we want our classifiers to have consistently useful output. An engineering team needs to balance risk against usefulness, and they’re clearly out of balance now. Risks will never go away, but guardrails whose boundaries are unclear and overly strict lead to models and agents that are less useful, rather than more. The bad guys will always figure out how to do bad stuff. Hamstrung AI for the rest of us is not a solution.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/when-guardrails-go-wrong/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Is Open-Source AI Really the Dangerous Path?</title>
		<link>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/</link>
				<comments>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/#respond</comments>
				<pubDate>Tue, 18 Aug 2026 15:58:34 +0000</pubDate>
					<dc:creator><![CDATA[Raffi Krikorian]]></dc:creator>
						<category><![CDATA[AI & ML]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19411</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Is-open-source-AI-really-the-dangerous-path.jpg" 
				medium="image" 
				type="image/jpeg" 
				width="2304" 
				height="1792" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/Is-open-source-AI-really-the-dangerous-path-160x160.jpg" 
				width="160" 
				height="160" 
			/>
		
		
				<description><![CDATA[The following article originally appeared on the Tech Policy Press site and is being republished here with the author’s permission. In Washington, AI is increasingly being treated as something that needs to be controlled. The government believes that AI is, first and foremost, a national security asset, meaning that it must be sequestered to prevent [&#8230;]]]></description>
								<content:encoded><![CDATA[
<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>The following article originally appeared on the</em> <a href="https://www.techpolicy.press/is-open-source-ai-really-the-dangerous-path/" target="_blank" rel="noreferrer noopener">Tech Policy Press</a> <em>site and is being republished here with the author’s permission.</em></p>
</blockquote>



<p class="wp-block-paragraph">In Washington, AI is increasingly being treated as something that needs to be controlled. The government believes that AI is, first and foremost, a national security asset, meaning that it must be sequestered to prevent enemies from gaining an advantage. On the other side of the world, in Beijing, the approach is moving in the opposite direction. China is reducing barriers, encouraging adoption, and using open-source AI as a way to spread Chinese-developed technology across global markets.</p>



<p class="wp-block-paragraph">There is now a fundamental divide. The United States is betting that control is the path to preserve its lead. China, instead, is betting on diffusion. The country whose technology is adopted most widely may ultimately shape the future of AI. Questions over open source and open weights sit at the center of that contest.</p>



<p class="wp-block-paragraph">Beginning on July 24, high-profile support for open source moved what is often a debate behind closed doors into the public sphere, where it belongs: Nvidia’s Jensen Huang’s <a href="https://x.com/JensenHuang/status/2080643682408321103" target="_blank" rel="noreferrer noopener">first-ever post on X</a> linked to an <a href="https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf" target="_blank" rel="noreferrer noopener">open letter</a> signed by 35 companies—including Palantir, Andreessen Horowitz and Microsoft—warning Washington not to over-restrict open source software. “Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty,” Huang wrote, leading the likes of Elon Musk and Mark Zuckerberg to post their support.</p>



<p class="wp-block-paragraph">Mozilla signed it because, despite being a very different company from many of the signatories and not always seeing eye to eye, we believe in the spirit and substance of the letter, particularly that ‘openness may be one of the most important paths to AI safety and security.’ The letter was followed by the <a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/" target="_blank" rel="noreferrer noopener">announcement</a> of the Open Secure AI Alliance for AI Safety and Security, which aims to “build and share open tools that promote responsible use of and trust in AI.”</p>



<p class="wp-block-paragraph">The battle is on. Here’s what’s behind it.</p>



<h2 class="wp-block-heading">Following the money</h2>



<p class="wp-block-paragraph">Open models now do about a third of the world’s AI work, but they collect only about four percent of the money.</p>



<p class="wp-block-paragraph">Those two numbers, taken from Mozilla’s new “<a href="https://stateofopensource.ai/" target="_blank" rel="noreferrer noopener">State of Open Source AI</a>” report, provide more context than the entire AI safety conversation. The report tells a story of a performance gap between open models—the ones whose weights anyone can download, run, and adapt—and the best proprietary systems. While the capabilities are still a jagged frontier, on average, the performance has narrowed sharply over the past year. Costs keep falling. Seventy-nine percent of developers now build with open models. And yet the money hasn’t followed the usage.</p>



<p class="wp-block-paragraph">This is where the battle lies. A third of the work, but only four percent of the money—that gap is the prize.</p>



<p class="wp-block-paragraph">The battle started almost exactly three years ago, when Anthropic’s Dario Amodei <a href="https://www.techpolicy.press/transcript-senate-hearing-on-principles-for-ai-regulation/" target="_blank" rel="noreferrer noopener">told Congress</a> that advanced open-source AI is on a “very dangerous path.” On the surface, his argument is pretty simple: once a model’s weights are public, no one can monitor abuse or revoke access. Once released, an open weights model can’t be “unreleased.” But let’s ask ourselves what that revoke switch actually does. Revocation is not a feature of the model, it is in the API contract. That means it only applies to a lab’s own customers—and the people in Amodei’s threat model were never customers. The labs shipping frontier-class open weights—such as DeepSeek, Alibaba, Mistral, and Moonshot (with its just-released, 2.8T-parameter model Kimi)—mostly sit outside Washington’s reach anyway. Two million open models already sit on Hugging Face; many run on a laptop. That means that in the real world, there’s no single kill switch to throw.</p>



<p class="wp-block-paragraph">This distance between what such a regulatory switch claims to control and what it actually does is what’s missing from the debate over which models are “safer.” It’s also the key to the fight over who captures AI’s value.</p>



<p class="wp-block-paragraph">To the companies that built the proprietary models, value is about maintaining a privileged position and using everything at their disposal to protect it—policy, pricing, and technology. For everybody else, value means the ability and power to shape, audit, and improve the systems we all depend on. And increasingly, that power doesn’t live in the model at all.</p>



<p class="wp-block-paragraph">For instance, right now, two developers can take the identical open model and ship completely different products: a scam-call operation or a nurse-advice hotline. The model doesn’t know (or care) about the difference. What is making the actual decisions is the layer of software built around it—the agentic harness—that sits between users and the model, determining what the application can access, remember, and act on.</p>



<p class="wp-block-paragraph">As models get cheaper, not to mention more interchangeable, that harness is where the power is actually going. And it’s being quietly locked up by the big labs. Farmers know how this story goes. They bought their tractors outright, but the manufacturer kept the keys to the software, making the farmers owners on paper but renters in practice. It took years of lawsuits—and, <a href="https://medium.com/enrique-dans/the-ftc-just-reminded-john-deere-what-ownership-means-when-you-buy-a-machine-you-should-be-able-fb8b59f3911d" target="_blank" rel="noreferrer noopener">just this month</a>, the Federal Trade Commission—to start prying that lock back open. A similar arrangement is now being built for the software that reads your email, books your travel, and remembers every detail of your life.</p>



<p class="wp-block-paragraph">This isn’t an accident of engineering; it’s a business model. A closed wrapper makes money by making itself expensive to leave. An open one can’t lock the door, so it survives only by staying worth using. Same underlying technology, opposite incentives. It’s the reason the value captured by open models sits at four percent while their usage sits at a third. The real question for all the builders right now isn’t which model you’re using. Rather, it’s whether you could leave for a different one.</p>



<h2 class="wp-block-heading">Guess who’s deciding the future?</h2>



<p class="wp-block-paragraph">The debate that matters isn’t really which models get released or which get regulated; it’s who controls the layer wrapped around them. That’s being decided right now—mostly by developers who don’t realize they’re the ones responsible. For a glimpse of the future, we can look to the internet: it exists as it does today because, when the architecture was still up for grabs, developers chose HTML and HTTP over proprietary walled gardens like AOL. AI is at that same juncture now, and the fact that two million open models already exist suggests plenty of builders have shown up early. That window doesn’t stay open on its own, and it doesn’t stay open forever. It stays open because people keep choosing it.</p>



<p class="wp-block-paragraph">For developers, four habits matter most in ensuring an open future:</p>



<ol class="wp-block-list">
<li><strong>Build on open harnesses, not just open models.</strong> The orchestration layer above the weights is where capability is concentrating, and closed labs are already welding it shut. Keeping it open takes deliberate effort.</li>



<li><strong>Own the memory layer. </strong>Store accumulated context in portable controllable formats, so it’s retrievable if a vendor changes its terms rather than trapped inside one.</li>



<li><strong>Keep a second model warm.</strong> Integrate an open model and keep it production-ready even while running primarily on a closed API, so switching is cheap if it becomes necessary.</li>



<li><strong>Don’t assume all open stacks are equal. </strong>Open models skew toward particular regions and providers; keeping this layer genuinely open means actively supporting a geographically distributed set of options, not defaulting to whichever model is cheapest this quarter.</li>
</ol>



<p class="wp-block-paragraph">None of this requires believing anyone is acting in bad faith. It’s worth noticing, though, that the loudest safety arguments arrived right around the time models got cheap enough for the real competition to move up a layer. That’s not evidence of a conspiracy—it’s just where the incentives point, and it’s why so much of the current debate is aimed at the wrong target.</p>



<p class="wp-block-paragraph">More evidence is in our report, and most of it is good news: performance gaps closing, costs collapsing, millions of developers building. The question in front of developers isn’t whether AI is dangerous—it’s whether they’ll hold the keys to the machines they’re building. The question for governments is whether the keys they’re reaching for turn anything at all. For now, that door is still open. Let’s work together to keep it that way.</p>



<p class="wp-block-paragraph"><em>And be sure to join us at </em>AI Codecon: Building with Open Source AI<em> on August 31, a free half-day virtual conference. You’ll hear from leading developers and technical experts working with open-weight models, self-hosted infrastructure, and real-world AI workflows, and learn how building in the open gives teams more control over costs, data privacy, and what they ship. <a href="https://www.oreilly.com/AI-Codecon/" target="_blank" rel="noreferrer noopener">Register today</a> to save your spot.</em></p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/is-open-source-ai-really-the-dangerous-path/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
		<item>
		<title>Zero to Agent in 30 Minutes: From Prompting to Loop Engineering with Ofer Mendelevitch</title>
		<link>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/</link>
				<comments>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/#respond</comments>
				<pubDate>Tue, 18 Aug 2026 10:54:48 +0000</pubDate>
					<dc:creator><![CDATA[Michelle Smith]]></dc:creator>
						<category><![CDATA[Zero to Agent in 30 Minutes]]></category>
		<category><![CDATA[Commentary]]></category>

		<guid isPermaLink="false">https://www.oreilly.com/radar/?p=19416</guid>

		
					<media:content 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar.png" 
				medium="image" 
				type="image/png" 
				width="504" 
				height="504" 
			/>

			<media:thumbnail 
				url="https://www.oreilly.com/radar/wp-content/uploads/sites/3/2026/08/zero-to-agent-cover-radar-160x160.png" 
				width="160" 
				height="160" 
			/>
		
				<custom:subtitle><![CDATA[What changes when coding agents can pursue a goal, verify their work, and review one another]]></custom:subtitle>
		
				<description><![CDATA[Ofer Mendelevitch, head of developer relations at BAND, used this episode of Zero to Agent in 30 Minutes to trace how coding workflows can give agents progressively more room to work on their own. Using a package version resolver as a running example, he compared step-by-step prompting with loop engineering and then showed how multiple [&#8230;]]]></description>
								<content:encoded><![CDATA[
<p class="wp-block-paragraph">Ofer Mendelevitch, head of developer relations at BAND, used this episode of <em>Zero to Agent in 30 Minutes</em> to trace how coding workflows can give agents progressively more room to work on their own. Using a package version resolver as a running example, he compared step-by-step prompting with loop engineering and then showed how multiple agents can collaborate on the same task.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Zero to Agent in 30 Minutes: From Prompting to Loop Engineering With Ofer Mendelevitch" width="500" height="281" src="https://www.youtube.com/embed/vg5IkLHx9FU?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>



<h2 class="wp-block-heading"><strong>How to move from prompting to multi-agent coding</strong></h2>



<ol class="wp-block-list">
<li><strong>Start with explicit prompts in every step.</strong> Give the coding agent a specification and tell it what to implement. Ofer used a package version resolver with existing Python tests, then followed up with prompts to verify the implementation, resolve open questions, and add packaging.</li>



<li><strong>Define a verifiable goal.</strong> Loop engineering replaces a sequence of individual prompts with an end result the agent can check for itself. In Ofer’s example, the agent had to implement the resolver and continue iterating until it was ready to ship with packaging. The agent can inspect the code, add tests, run them, fix failures, and verify the package without waiting for another human prompt after each step. He demonstrated how this approach allowed the agent to work through the task until the goal was met.</li>



<li><strong>Add a second agent as a reviewer.</strong> Ofer then moved from a single coding agent to two collaborating agents (using Jam), assigning one to write the code and another to review it. The reviewer examined the specification, provided feedback, and ran additional checks, including adversarial probes. A larger group of coding agents could include agents focused on security, compliance, testing, frontend, backend, or DevOps. He also described using different coding agents together so that one model can challenge work produced by another.</li>
</ol>



<p class="wp-block-paragraph">The shift toward more autonomous coding workflows starts with how the work is framed. By defining goals agents can verify, giving them room to iterate, and assigning complementary agents to review the work, developers can reduce the amount of human intervention required and achieve higher quality for the code generated by the coding agents.</p>



<h2 class="wp-block-heading"><strong>Coming next week</strong></h2>



<p class="wp-block-paragraph">Next week, Craig Hewitt will host <em><a href="https://learning.oreilly.com/live-events/zero-to-agent-in-30-minutes/0642572392338/" target="_blank" rel="noreferrer noopener">Zero to Agent in 30 Minutes</a></em> to focus on building a voice-first workflow with OpenAI Codex<strong>.</strong> The episode will show how natural voice commands can operate a development environment, run subagent workers in parallel, and trigger browser-use workflows. It will also cover structured Codex project directories and hands-free system-level execution, with the developer directing the work by voice.</p>
]]></content:encoded>
							<wfw:commentRss>https://www.oreilly.com/radar/zero-to-agent-in-30-minutes-from-prompting-to-loop-engineering-with-ofer-mendelevitch/feed/</wfw:commentRss>
		<slash:comments>0</slash:comments>
							</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Object Caching 81/110 objects using Memcached
Page Caching using Disk: Enhanced (Page is feed) 
Minified using Memcached

Served from: www.oreilly.com @ 2026-08-27 18:24:30 by W3 Total Cache
-->