<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
    <channel>
        <title>VentureBeat</title>
        <link>https://venturebeat.com/feed/</link>
        <description>Transformative tech coverage that matters</description>
        <lastBuildDate>Mon, 27 Jul 2026 11:10:41 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <copyright>Copyright 2026, VentureBeat</copyright>
        <item>
            <title><![CDATA[VentureBeat Research: Where enterprise AI agent governance hasn't caught up]]></title>
            <link>https://venturebeat.com/technology/venturebeat-research-where-enterprise-ai-agent-governance-hasnt-caught-up</link>
            <guid isPermaLink="false">39OjqkS473sAWD2bJI3iUR</guid>
            <pubDate>Fri, 24 Jul 2026 19:34:29 GMT</pubDate>
            <description><![CDATA[<p>Enterprises deployed AI agents ahead of the controls needed to manage them — and they did it knowingly. That is the central finding across the five parallel surveys VentureBeat Research fielded in June, spanning every layer of the agentic stack. Now those enterprises are retrofitting to catch up with their own standards, and they are budgeting for it: In each of the five control layers we measured, 57 to 68% of enterprises plan to switch vendors or add new ones within 12 months, and roughly a third, depending on the layer, plan to move within the quarter.</p><p><a href="https://venturebeat.com/category/resources">VentureBeat Research</a> measured the five controls an enterprise has to build before it can trust an agent: identity, evaluation, cost telemetry, the context layer, and orchestration. Identity governs which agent is allowed to do what, under whose credentials. Evaluation determines whether the agent&#x27;s work is any good. Cost telemetry tracks what each agent costs to run. The context layer supplies the business data and definitions agents draw on when they answer. And the orchestration control plane coordinates multi-step agent work. Each of our five reports measures one of those controls.</p><p><b>Most deployed &quot;agents&quot; are chatbots wearing the label.</b> Seventy-one percent of enterprises said a quarter or fewer of their deployed &quot;agents&quot; can complete multi-step work on their own; only 10% said true agents are the majority of what they run. These respondents are positioned to know: 81% recommend or decide AI purchases at their companies. A single-prompt chatbot with a human reading every answer needs none of the controls the other four reports measure. A true multi-step agent needs all of them — and most enterprises can&#x27;t say which one they&#x27;ve deployed. <i>(Full findings: </i><a href="https://venturebeat.com/resources/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents"><i>Agentic Orchestration report.</i></a><i>)</i></p><p><b>Autonomy is outrunning trust in the evaluations that gate it.</b> Two-thirds of enterprises either already allow an agent to push a code or system change to production on automated evaluation results alone, with no human review, or are actively engineering toward that within 12 months. Only 5% fully trust the evaluations that would make that call — and half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. Before removing human review from any workflow, test evaluations against production outcomes rather than internal benchmarks. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway"><i>Agent Reliability &amp; Evals report</i></a><i>.)</i></p><p><b>Companies that let agents share credentials get hit more often.</b> Sixty-nine percent of companies let at least some of their agents share credentials — multiple agents operating under one API key or service account. Organizations that allow credential sharing anywhere experienced a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (nine of 22) at companies where every agent has its own scoped identity. The fix is scoped identity for every agent, starting with the ones that touch production systems. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials"><i>Agentic Security &amp; Identity report</i></a><i>.)</i></p><p><b>The most expensive hardware in the building runs at half capacity or less.</b> More than eight in 10 enterprises that run their own GPUs reported utilization of 50% or less, and only 44% rigorously track what their AI compute actually costs and returns. The number worth chasing first isn&#x27;t more GPUs — it&#x27;s the utilization and per-workload cost of the ones already running. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs"><i>AI Infrastructure &amp; Compute report</i></a><i>.)</i></p><p><b>Agents answer confidently from data nobody governs.</b> Fifty-seven percent of enterprises traced a confident, wrong agent answer in the past six months to their own missing or inconsistent business context — wrong metrics, stale definitions, absent documents — and most saw it happen more than once. Governing the definitions agents answer from — metrics and entities first — has to come before scaling the agents that depend on them. <i>(Full findings: </i><a href="https://venturebeat.com/resources/the-ai-context-gap-enterprise-ai-organizations-have-a-trust-problem-not-a-retrieval-problem-and-most-are-still-building-the-fix"><i>Context Layers / RAG report</i></a><i>.)</i></p><p>No layer has an entrenched incumbent: The defaults today are the built-in tools that ship with the big AI platforms enterprises already use. Switching intent runs highest in orchestration itself, where 68% plan to adopt, add, or replace platforms within 12 months and 34% within the quarter. Our surveys did not ask which direction that money moves — toward the platforms&#x27; built-in tools or toward the specialists challenging them — and that open question is the next four quarters of this market.</p><hr/><p><b>About this research</b> </p><p><a href="https://venturebeat.com/category/resources">VentureBeat Research</a> fielded five parallel surveys in June 2026 under its VB Pulse program: Agentic Orchestration (101 respondents), Agent Reliability &amp; Evals (157), Agentic Security &amp; Identity (107), AI Infrastructure &amp; Compute (107), and Context Layers / RAG (101) — 573 qualified respondents in total, all at organizations with 100 or more employees. Samples are self-selected, and some findings should be read directionally; each report carries its full methodology note. What the pattern supports more strongly than any single percentage is the direction: every survey, independently, points the same way. VentureBeat produces both this research and <a href="https://venturebeat.com/vbtransform2026">VB Transform</a>, the conference where these reports debuted.</p>]]></description>
            <author>mmarshall@venturebeat.com (Matt Marshall)</author>
            <category>Technology</category>
            <category>VB Transform</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/1u0FLsMdYSkijWN8fnF7M4/860ed6aa6bada2f76f6311ecb1cd29bc/Gemini_Generated_Image_g5fa0gg5fa0gg5fa.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows]]></title>
            <link>https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows</link>
            <guid isPermaLink="false">4YLayWmjH3TlUhOPU8GMlC</guid>
            <pubDate>Fri, 24 Jul 2026 17:00:00 GMT</pubDate>
            <description><![CDATA[<p><a href="https://www.anthropic.com/">Anthropic</a> released Claude <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> on Friday, a model the company says delivers nearly all the intelligence of its top-of-the-line Claude <a href="https://www.anthropic.com/claude/fable">Fable 5</a> at half the cost — a launch that signals how the AI race is shifting from raw capability to the economics of daily use.</p><p>The model, available immediately on all of Anthropic&#x27;s platforms, is priced at $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor, <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a>. It becomes the new default model on <a href="https://support.claude.com/en/articles/11049741-what-is-the-max-plan">Claude Max</a>, Anthropic&#x27;s premium consumer tier, and the strongest model available on <a href="https://support.claude.com/en/articles/8325606-what-is-the-pro-plan">Claude Pro</a>.</p><p>The positioning is deliberate. Anthropic is not claiming <a href="http://anthropic.com/news/claude-opus-5">Opus 5 </a>is its smartest model — that distinction still belongs to <a href="https://www.anthropic.com/claude/fable">Fable 5</a>, and rival systems retain an edge in certain domains. Instead, the company is making a subtler argument that may matter more to enterprise buyers: that the most economically important AI work happens in a middle band of difficulty, where near-frontier intelligence delivered efficiently and cheaply beats frontier intelligence delivered expensively.</p><p>&quot;Opus 5 as your daily driver, the model you hand complex work to and review when it&#x27;s done,&quot; an Anthropic spokesperson said in an interview with VentureBeat, describing how the company&#x27;s lineup now stratifies. &quot;Fable 5 for your most ambitious work, the days-long autonomous projects nothing could take on before... Sonnet 5 for work you run at scale, where speed and cost per call decide what ships. Haiku 4.5 for subagents and instant answers.&quot;</p><h2><b>How Claude Opus 5 benchmark results stack up against Fable 5 and rival AI models</b></h2><p>On paper, the results are striking. Anthropic says <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> sets new state-of-the-art marks on coding and knowledge-work evaluations including <a href="https://www.frontierbench.ai/announcement">Frontier-Bench</a> and <a href="https://artificialanalysis.ai/evaluations/gdpval-aa">GDPval-AA</a>. On <a href="https://www.frontierbench.ai/announcement">Frontier-Bench v0.1</a>, an agentic terminal coding benchmark, Opus 5 scores 43.3 percent — more than double Opus 4.8&#x27;s 18.7 percent and well ahead of Fable 5&#x27;s 33.7 percent — at a lower cost per task, according to the company. On <a href="https://arcprize.org/arc-agi/3">ARC-AGI 3</a>, an evaluation of novel problem-solving, Anthropic reports Opus 5 scored three times as high as the next best model. On <a href="https://github.com/xlang-ai/OSWorld-V2">OSWorld 2.0</a>, a computer-use benchmark, the company says the model surpasses Fable 5&#x27;s best result at just over a third of the cost.</p><p>The numbers come with honest caveats that are themselves notable in an industry prone to superlatives. Anthropic acknowledges <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> remains behind <a href="https://www.anthropic.com/claude/mythos">Mythos 5</a>, a competing model, on cybersecurity tasks and biology research, and an OpenAI-family model still leads on one agentic coding benchmark.</p><p>The more revealing caveat came from Anthropic itself, when asked where <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> still falls short of <a href="https://www.anthropic.com/claude/fable">Fable 5</a>. The spokesperson&#x27;s answer amounted to a candid admission about what benchmarks do and don&#x27;t capture.</p><p>&quot;The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it&#x27;s strongest. What those evals don&#x27;t measure is duration,&quot; the spokesperson told VentureBeat. &quot;One way to put it: Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark.&quot;</p><p><a href="https://www.anthropic.com/claude/fable">Fable 5</a>, by contrast, &quot;is for the longest, most autonomous jobs, where the model has to stay coherent across many connected steps over hours or days with dense source material,&quot; the spokesperson said, advising customers to &quot;run both on a representative workload, one bounded task and one long-horizon job.&quot; That framing — bounded tasks versus long-horizon autonomy — may become the defining axis of model differentiation in 2026, as benchmarks saturate and the hardest remaining problems involve sustained, multi-day agentic work rather than discrete puzzles.</p><h2><b>Why token efficiency is becoming the real battleground for enterprise AI spending</b></h2><p>Threaded through the launch is a theme Anthropic clearly wants buyers to absorb: <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> doesn&#x27;t just score well, it scores well per dollar. The model ships with an adjustable &quot;effort&quot; setting that lets customers trade intelligence for speed and token savings, and Anthropic&#x27;s charts emphasize performance at a given cost rather than peak performance alone.</p><p>Early customers echoed the point with unusual specificity. Harvey, the legal AI company, said Opus 5 achieved similar performance to Opus 4.8&#x27;s maximum-reasoning mode &quot;while generating 26% fewer tokens on average,&quot; according to Niko Grupen, its head of applied research. Richard Pham of Fundamental Research Lab said that on hard financial-modeling tasks, the model averaged nine percentage points higher accuracy &quot;while using roughly one-third fewer turns and tool calls and 60% less time.&quot;</p><p>Wade Foster, chief executive of Zapier, said Opus 5 topped his company&#x27;s AutomationBench leaderboard &quot;without spending more tokens than prior Claude models,&quot; running a full churn-prevention workflow from start to finish. &quot;Previous models didn&#x27;t pass; Opus 5 hit 100%,&quot; he said. Scott Wu, chief executive of Cognition, the company behind the Devin coding agent, said that on FrontierCode 1.1, &quot;Claude Opus 5 approaches Fable-level performance at half the cost,&quot; with particular strength in debugging and root-cause analysis.</p><p>The efficiency emphasis reflects commercial reality. Enterprise AI spending is no longer experimental, and inference costs — the price of actually running these models at scale — have become a board-level line item. </p><p>Anthropic&#x27;s business skews heavily toward API and enterprise usage; according to a February 2026 analysis by <a href="https://research.contrary.com/company/anthropic">Contrary Research</a>, Claude held roughly 40 percent of the enterprise large language model market by usage as of late 2025, and Claude Code alone had reached about $1 billion in annualized revenue. For a company whose customers pay by the token, a model that does more with fewer tokens is not a nice-to-have. It is the product.</p><h2><b>Self-verifying AI agents and what they mean for the hidden costs of automation</b></h2><p>Beyond the numbers, Anthropic is selling a behavioral story: that <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> verifies its work and iterates until it succeeds. The company offered several examples from testing that read like small parables of machine stubbornness.</p><p>In one <a href="https://www.frontierbench.ai/announcement">Frontier-Bench</a> task, the model was asked to reconstruct a machine part as a 3D CAD model from a drawing it was intentionally given no way to view. Rather than fail, Anthropic says, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels — and did so repeatedly, while no competing model solved the task in five attempts. In another case, given a real bug in a popular open-source package manager, the model found the root cause and fixed an edge case the community&#x27;s own patch had missed; a competing model patched only the symptom and declared victory. An engineer at a trading firm, the company says, used Opus 5 to build a market data feed for a new exchange in a single session and, finding no live feed to validate against, watched the model build its own test harness to check its parsing code.</p><p>Customers described similar behavior in the wild. Cristian Rivera, a staff software engineer at Stripe, said he gave the model &quot;a chief-of-staff role over my dev environments&quot; for a weekend: &quot;it built its own monitor, drove each box, and pulled me in only for the judgment calls.&quot;</p><p>This is the capability enterprises actually care about, and it is worth dwelling on why. The gap between a model that produces plausible output and one that verifies its output is the gap between a demo and a deployable system. Most of the hidden cost of enterprise AI today is human review — engineers checking the machine&#x27;s work. A model that reliably checks its own work compresses that cost, which is precisely why customers keep citing fewer turns, fewer passes, and less time rather than higher raw scores.</p><h2><b>Inside Anthropic&#x27;s safety strategy: capability gaps, classifiers, and model fallbacks</b></h2><p>The launch also showcases Anthropic&#x27;s increasingly intricate approach to safety — one that now involves deliberately not teaching its models certain skills. The company says its automated behavioral audit found Opus 5 to be its most aligned model to date, scoring 2.3 on overall misaligned behavior, lower than <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a>, <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a>, or <a href="https://www.anthropic.com/claude/fable">Fable 5</a>, with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse.</p><p>On the capability side, Anthropic says it intentionally avoided training <a href="http://anthropic.com/news/claude-opus-5">Opus 5</a> on cyber tasks, as it did with Opus 4.8. The model improved on them anyway — a side effect of general capability gains — and now nearly matches Mythos 5 at finding software vulnerabilities. But it remains far behind at exploiting them: on Anthropic&#x27;s OSS-Fuzz evaluation, Opus 5 identified vulnerabilities at a 79.4 percent rate, close to Mythos 5&#x27;s 80 percent, but succeeded at developing exploits in only 4 challenges versus Mythos 5&#x27;s 13. That asymmetry — strong at defense-relevant discovery, weak at offense-relevant exploitation — appears to be by design, and the safeguards follow the same logic. Anthropic expects Opus 5&#x27;s cyber classifiers to intervene about 85 percent less often than Fable 5&#x27;s.</p><p>When a classifier does trigger, requests in <a href="http://claude.ai">Claude.ai</a>, <a href="https://code.claude.com/docs/en/overview">Claude Code</a>, and <a href="https://claude.com/product/cowork">Claude Cowork</a> fall back to <a href="https://www.anthropic.com/news/claude-opus-4-8">Opus 4.8</a> by default — raising an obvious question: if a request is too risky for one model, why is it acceptable for another? &quot;The model it falls back to has lower capability levels making the risk of harmful use lower as well,&quot; the spokesperson said, adding that &quot;there is a message that lets the user know when this occurs and is visible in the chat.&quot;</p><p>The logic is defensible, but it reveals how AI safety actually works in 2026: risk is not a property of the question alone, but of the question multiplied by the capability of the system answering it. On biology, the calculus runs the other way. Opus 5 is now Anthropic&#x27;s most capable generally available model for scientific research — scoring 10.2 percentage points higher than Opus 4.8 on the company&#x27;s internal chemistry benchmark — though the spokesperson acknowledged that &quot;Mythos 5 remains the stronger model for long-horizon, open-ended work like autonomous drug design campaigns.&quot;</p><h2><b>The business stakes behind the launch: a $380 billion valuation and massive compute bets</b></h2><p>The launch lands at a moment of extraordinary commercial momentum — and extraordinary obligations — for Anthropic. Reuters reported in February that the company was valued at <a href="https://www.reuters.com/technology/anthropic-valued-380-billion-latest-funding-round-2026-02-12/">roughly $380 billion</a> in its latest funding round, following a period in which, per Contrary Research&#x27;s analysis, its annualized revenue climbed from about $1 billion at the end of 2024 to a projected $9 billion by the end of 2025, with internal targets reportedly <a href="https://research.contrary.com/company/anthropic">reaching $20 to $26 billion for 2026</a>. Those targets are underwritten by enormous infrastructure commitments, including a <a href="https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships">reported $30 billion Azure compute deal</a> alongside arrangements with Google Cloud and Nvidia — spending that only pencils out if enterprises keep expanding usage.</p><p>That is the context in which Opus 5&#x27;s pricing strategy makes sense. Holding the price at Opus 4.8 levels while roughly doubling performance on key agentic benchmarks is effectively a steep price cut per unit of capability, designed to widen the funnel of workloads that are economical to automate. Every task that was marginal at Opus 4.8&#x27;s cost-per-success becomes viable at Opus 5&#x27;s — and every viable task is recurring token revenue.</p><p>The regulatory backdrop has grown more complex as well. A U.S. judge gave final approval this week to <a href="https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/">Anthropic&#x27;s $1.5 billion copyright settlement with book authors</a>, Reuters reported, closing a chapter of litigation over the company&#x27;s early training data. And in June, Reuters, citing Axios, reported that the U.S. government had moved to <a href="https://www.reuters.com/technology/us-blocks-foreign-access-anthropics-most-advanced-ai-models-axios-reports-2026-06-13/">block foreign access </a>to Anthropic&#x27;s most advanced models — a reminder that frontier AI is now entangled with export policy in ways that shape which customers can buy what.</p><p>Also shipping Friday: a Fast mode running at roughly 2.5 times default speed at twice the base price, automatic fallback routing on the API, and mid-conversation tool changes that no longer invalidate the prompt cache — a small feature that agent developers may appreciate more than any benchmark. Consistent with prior Opus models, Opus 5 carries no data retention requirements for general access, a point the spokesperson flagged unprompted for customers with &quot;a hard zero data retention requirement.&quot; Developers can access the model as claude-opus-5 on the <a href="https://platform.claude.com/login?returnTo=%2F%3F">Claude API</a> starting today.</p><p>Two questions will determine whether the bet pays off: whether <a href="http://anthropic.com/news/claude-opus-5">Opus 5&#x27;s efficiency claims </a>survive contact with production workloads at scale, and whether enterprises embrace a world where safety classifiers, not users, sometimes decide which model answers. But the deeper message of Friday&#x27;s launch is that the AI industry&#x27;s center of gravity has moved. For three years, the labs competed on what their best model could do on its best day. With Opus 5, Anthropic is competing on something less glamorous and far more lucrative: what a very good model can do every day, for half the price. In a market where the frontier keeps moving, Anthropic is wagering that the real fortune lies just behind it.</p><p>
</p>]]></description>
            <author>michael.nunez@venturebeat.com (Michael Nuñez)</author>
            <category>Orchestration</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/1Yq5FNKJmmxeMO98fIwpRM/d8ae9af717a5770950856aff8f2ec3e6/Opus-5-Hero.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI]]></title>
            <link>https://venturebeat.com/infrastructure/microsoft-launches-new-in-house-ai-models-it-says-cut-costs-up-to-89-versus-openai</link>
            <guid isPermaLink="false">2GnXy6jBrKjztwcfU2nQGd</guid>
            <pubDate>Thu, 23 Jul 2026 23:37:05 GMT</pubDate>
            <description><![CDATA[<p><a href="https://microsoft.ai/">Microsoft AI</a> released two new in-house models into public preview on Wednesday — <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5-Pro</a>, its highest-fidelity image generator to date, and <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a>, a speech model built for high-volume enterprise workloads — while publishing production data that amounts to the company&#x27;s most aggressive argument yet that it can power its own products without leaning on OpenAI&#x27;s frontier models.</p><p>The announcement, made by <a href="https://microsoft.ai/">Microsoft AI&#x27;s Superintelligence team</a>, lands roughly a year after the company committed to building purpose-built models internally, and it arrives with an unusual level of specificity about where those models now run: <a href="https://www.bing.com/">Bing</a>, <a href="https://www.microsoft.com/en-us/microsoft-365/powerpoint">PowerPoint</a>, <a href="https://www.microsoft.com/en-us/microsoft-365/onedrive/online-cloud-storage">OneDrive</a>, <a href="https://www.microsoft.com/en-us/dynamics-365">Dynamics 365</a>, <a href="https://excel.cloud.microsoft/en-us/">Excel</a>, <a href="https://github.com/features/copilot">GitHub Copilot</a>, and <a href="https://azure.microsoft.com/en-us">Azure</a>. The message to enterprise buyers — and, implicitly, to OpenAI — is that Microsoft&#x27;s homegrown models are no longer research projects. They are production infrastructure serving millions of users.</p><p>&quot;Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models,&quot; the company wrote in its announcement blog.</p><h2><b>How MAI-Image-2.5-Pro and MAI-Voice-2-Flash stake out opposite ends of the AI cost curve</b></h2><p>The two new releases occupy opposite ends of what Microsoft calls the quality-speed-cost curve, and the positioning is deliberate. <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5-Pro</a> targets the premium tier: hero imagery, detailed editing, and precise in-image text rendering — the last of which has long been a notorious weak spot for image generation models. Microsoft priced the model at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. The base <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5</a> model recently launched at <a href="https://microsoft.ai/news/introducing-mai-image-2-5/">No. 2 for image editing on Arena</a>, the community leaderboard that has become a de facto scoreboard for generative media.</p><p>The creative industry appears to be taking notice. Rob Reilly, global chief creative officer at advertising giant WPP, called the Pro model &quot;a strong leap forward for GenMedia tools&quot; in a statement included in Microsoft&#x27;s announcement, adding that &quot;Microsoft has firmly established itself among the leaders in generative AI.&quot;</p><p><a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a> goes the other direction. First previewed at Microsoft&#x27;s <a href="https://news.microsoft.com/build-2026/">Build conference</a>, Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It is designed for the unglamorous but enormous market of high-volume voice — call centers, voice agents, and real-time speech applications where latency and cost-per-call matter more than marginal gains in expressiveness. Together, the two models reflect a strategy of building families of models rather than a single flagship, because, as the company put it, a creative studio chasing maximum fidelity has very different needs from a customer service operation handling millions of calls a day.</p><h2><b>Microsoft&#x27;s production metrics show in-house models cutting GPU costs by up to 89%</b></h2><p>The model launches are arguably less newsworthy than the deployment metrics Microsoft attached to them — numbers that read like a systematic case for swapping out third-party frontier models across its product portfolio. </p><p><a href="https://explore.microsoft.com/en-us/bing/features/bing-image-creator?form=MA13FV">Bing Image Creator </a>now runs entirely on <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5</a>, end to end, marking the first time the consumer image tool is fully in-house. In PowerPoint, Microsoft says MAI-Image-2.5 reduces GPU costs by up to 84% compared with GPT-Image-2, OpenAI&#x27;s image model. In OneDrive, where MAI-Image-2.5 is now the default for key image-editing scenarios, the company reports a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.</p><p>On the voice side, <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a> now powers Dynamics 365 Contact Center — the platform used by customers including T-Mobile and EasyJet — where Microsoft claims GPU cost reductions of up to 89%. The model is also integrated into Azure Voice Live for developers building speech-to-speech agents.</p><p>Perhaps the most consequential deployment sits in healthcare. Microsoft&#x27;s <a href="https://www.microsoft.com/en-us/health-solutions/clinical-workflow/dragon-copilot">Dragon Copilot</a>, used by 170,000 medical providers and responsible for processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft says internal evaluations show a 50% relative reduction in both transcription and language-identification error rates across most languages — a meaningful claim in a domain where transcription errors can propagate directly into clinical notes.</p><h2><b>Inside the &#x27;hill-climbing&#x27; strategy that lets small models beat GPT-5.6 in Excel</b></h2><p>In a companion post published the same day, Microsoft detailed the methodology behind these results — what it calls its &quot;<a href="https://microsoft.ai/news/hill-climbing-mai-models-for-github-copilot-and-excel/">hill-climbing machine</a>,&quot; an integrated flywheel of data, models, and the product &quot;harness&quot; that surrounds them.</p><p>The clearest example is <a href="https://microsoft.ai/news/introducingmai-code-1-flash/">MAI-Code-1-Flash</a>, the lightweight coding model launched in GitHub Copilot in June. Microsoft says the model achieves an approximately 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. Developer retention tells a similar story: users were 6% more likely to return across multiple days than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.</p><p>Then Microsoft did something more interesting. It took the MAI-Code-1-Flash checkpoint and further <a href="https://microsoft.ai/news/hill-climbing-mai-models-for-github-copilot-and-excel/">trained it inside an Excel reinforcement learning environment</a>, teaching a coding model the tools and workflows of spreadsheet knowledge work. The result, according to production user feedback, is a model on par with GPT-5.6 for the most common Excel tasks — while being small enough to run on Nvidia&#x27;s older H100 and even A100 GPUs rather than requiring the latest-generation accelerators.</p><p>That hardware detail deserves emphasis. Every major AI company is fighting for allocation of cutting-edge chips, and a model that delivers frontier-adjacent quality on two-generation-old silicon fundamentally changes the deployment economics. It also frees the newest hardware — including Microsoft&#x27;s now-operational GB200 cluster — for training rather than serving.</p><h2><b>Satya Nadella&#x27;s &#x27;frontier diffusion&#x27; manifesto redraws the OpenAI relationship</b></h2><p>Microsoft CEO Satya Nadella framed the announcements in a lengthy post on X titled &quot;<a href="https://x.com/satyanadella/status/2080329851127669104">Frontier Diffusion &amp; Control</a>,&quot; which functions as something close to a strategic manifesto. &quot;We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs,&quot; Nadella wrote, adding that Microsoft is &quot;beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.&quot;</p><p>Translated from executive prose: capabilities that were state-of-the-art a year ago are now table stakes, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product usage. Why pay frontier prices for a frontier model when a user just wants to reformat a spreadsheet column?</p><p>Nadella was careful to note that &quot;frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI&quot; — but he also articulated a pointed principle of model independence, arguing that a company&#x27;s evaluations &quot;should continue to hill climb even when any given model has been removed.&quot; </p><p>“Keeping the harness, memory, context, and skills outside the model, he argued, is what gives Microsoft control. The subtext is hard to miss. Reuters reported in April that Microsoft’s <a href="https://www.reuters.com/legal/litigation/microsoft-end-exclusive-license-openais-technology-2026-04-27/">exclusive license to OpenAI’s technology</a> had been revised into a non-exclusive arrangement, and The Information reported last September that Microsoft had <a href="https://www.theinformation.com/articles/microsoft-buy-ai-anthropic-shift-openai">begun incorporating Anthropic models</a> into some products. Wednesday’s announcement completes the triangle: Microsoft as orchestrator, with its partners’ frontier models as interchangeable components and its own models absorbing an ever-larger share of routine traffic.”</p><h2><b>Developers cheer cheaper task-specific models while skeptics question Microsoft&#x27;s track record</b></h2><p>The response online captured both the appeal and the skepticism surrounding the strategy. &quot;I love when people use small models for niche tasks,&quot; wrote one X user, <a href="https://x.com/mavihsk/status/2080330529547993252">@mavihsk</a>, responding to Nadella&#x27;s post. &quot;Why do I have to use the all-knowing model just to change my field in Excel?&quot; Another user, <a href="https://x.com/nabu_lines/status/2080343512780837226">@nabu_lines</a>, distilled the pitch neatly: &quot;cost and performance both improve when you stop overusing the biggest model.&quot;</p><p>Others were less charitable about Microsoft&#x27;s execution track record. &quot;Microsoft is the worst when it comes to listening to user feedback,&quot; wrote designer <a href="https://x.com/designedbyabin/status/2080332368301412434">@designedbyabin</a>, arguing the company &quot;will lose the AI race because they repeatedly failed to understand user needs.&quot; And one user, <a href="https://x.com/tokenoverflow/status/2080386145712824694">@tokenoverflow</a>, offered a drier critique of the model-independence pitch: &quot;i want it keep hill climbing after removing microsoft.&quot;</p><p>The skeptics raise a fair point. Microsoft&#x27;s self-reported metrics — accept rates, save rates, GPU savings — come from its own internal evaluations, not independent benchmarks, and the company chooses which comparisons to publish.</p><p>But the strategy&#x27;s logic does not depend on any single number. Nadella&#x27;s framing that software now has &quot;<a href="https://x.com/satyanadella/status/2080329851127669104">real marginal cost for the first time</a>&quot; explains why Microsoft is obsessive about tokens, GPUs, and serving costs: when AI features run on every keystroke across a billion-user product portfolio, an 84% GPU cost reduction is not an optimization. It is the difference between a viable business and a money pit.</p><h2><b>Why Microsoft is turning its internal AI playbook into an Azure product</b></h2><p>The final piece of the strategy is that Microsoft is selling the playbook, not just the models. Nadella explicitly positioned the hill-climbing approach as &quot;a template for every other AI native, SaaS, or Enterprise company,&quot; and Microsoft is packaging the toolchain through Foundry and what it calls Frontier Tuning — letting enterprises train specialized models against their own proprietary evaluations and reinforcement learning environments. That turns Microsoft&#x27;s internal cost-cutting exercise into an Azure product, and it gives enterprise customers a reason to run their AI workloads on Microsoft&#x27;s cloud even if the models themselves come from elsewhere.</p><p>The company&#x27;s emphasis on models trained &quot;on clean, traceable, enterprise-grade data, without distillation from third-party models&quot; serves the same commercial end. In an industry facing mounting scrutiny over training data provenance, Microsoft is betting that enterprise buyers — and courts — will care where model capabilities come from. Microsoft says it is now extending the hill-climbing approach to <a href="https://copilot.microsoft.com/">Copilot Chat</a>, <a href="https://outlook.live.com/mail/">Outlook</a>, and <a href="https://www.microsoft.com/en-us/microsoft-365/powerpoint">PowerPoint</a>, and both new models are available in public preview through <a href="https://azure.microsoft.com/en-us/products/ai-foundry">Microsoft Foundry</a> and the <a href="https://playground.microsoft.ai/">MAI Playground</a>. &quot;None of this is an endpoint,&quot; the company wrote. &quot;We&#x27;re just getting started.&quot;</p><p>Seven years ago, <a href="https://www.cnbc.com/2024/08/10/rise-of-openai-microsofts-13-billion-artificial-intelligence-bet.html">Microsoft bet more than $13 billion</a> that OpenAI would build the future of AI. Wednesday&#x27;s announcement suggests the company has since learned a cheaper lesson: the future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it ordinary.</p>]]></description>
            <author>michael.nunez@venturebeat.com (Michael Nuñez)</author>
            <category>Infrastructure</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/21LP7BYkYtNcUVP0RAuTq2/d6a298425c240d52da4e211885164c27/Nuneybits_Vector_art_of_Microsoft_logo_rerouting_AI_traffic_fro_0889d4d3-b42c-4858-a9dc-f77546ce9c6c.webp?w=300&amp;q=30" length="0" type="image/webp"/>
        </item>
        <item>
            <title><![CDATA[Agentic coding goes hands-free as OpenAI brings GPT-Live's full duplex voice control to Codex and ChatGPT on the desktop ]]></title>
            <link>https://venturebeat.com/orchestration/agentic-coding-goes-hands-free-as-openai-brings-gpt-lives-full-duplex-voice-control-to-codex-and-chatgpt-on-the-desktop</link>
            <guid isPermaLink="false">6vdl1xVT8AqHbtY9AbK3Bi</guid>
            <pubDate>Thu, 23 Jul 2026 21:17:00 GMT</pubDate>
            <description><![CDATA[<p>Two weeks after debuting its <a href="https://venturebeat.com/technology/openai-launches-gpt-live-a-full-duplex-voice-upgrade-that-lets-chatgpt-talk-more-like-a-person">more naturalistic GPT-Live audio AI model</a> with full-duplex capabilities (listening and speaking at the same time), OpenAI is bringing it directly into developer workflows. </p><p>The company announced that <a href="https://x.com/OpenAI/status/2080378182469857576">GPT-Live now powers the ChatGPT desktop application</a> on macOS and Windows, integrating directly with agentic systems like Codex and ChatGPT Work (which are separate experiences available in the ChatGPT desktop app). </p><p>When OpenAI initially launched GPT-Live on July 8, 2026, it introduced a continuous audio model capable of listening and speaking simultaneously—eliminating rigid turn-taking while delegating complex reasoning to background models like GPT-5.5. </p><p>Today&#x27;s release expands that conversational layer to technical tasks, enabling software engineers to orchestrate multi-threaded coding jobs, review pull requests, and debug applications using natural voice commands.</p><p>As such, it could usher in a new era of &quot;hands free&quot; software development and even live, in-person group coding parties for <a href="https://openai.com/index/codex-for-knowledge-work/">the more than 10 million weekly active users</a> across Codex and <a href="https://venturebeat.com/orchestration/openai-introduces-chatgpt-work-a-cloud-based-ai-agent-that-manages-tasks-across-email-slack-and-calendars">ChatGPT Work</a>. Codex, of course, is the name given to OpenAI&#x27;s models and harness focused on coding, but which the company has this year expanded into a more <a href="https://venturebeat.com/technology/openai-drastically-updates-codex-desktop-app-to-use-all-other-apps-on-your-computer-generate-images-preview-webpages">general productivity platform. </a>An OpenAI spokesperson told VentureBeat this is the first time voice activation</p><p>OpenAI posted a <a href="https://youtu.be/E0ZMOschrTU?si=WWc8fZ2o0UtxrDFk">promotional video</a> showing some of its employees, Codex developer experience engineer Jason Liu and Codex technical staffer Guinness Chen, speaking to the same ChatGPT desktop app session in the same room, each issuing different instructions and conversing with the same model. </p><div></div><h2><b>New capabilities unlocked</b></h2><p>At its core, this integration relies on decoupling the real-time voice layer from the underlying execution engines.</p><p>While GPT-Live maintains fluid conversation—inserting natural verbal acknowledgments like &quot;got it&quot; without interrupting the user—it passes heavy computational workloads to background reasoning models. </p><p>On macOS, the desktop application incorporates &quot;Appshots&quot; and screen context features, allowing ChatGPT Voice to analyze the frontmost window alongside local files, codebase structures, and active plugins.</p><p>This architecture creates a pair-programming dynamic where developers talk through problems conversationally while agents execute tasks asynchronously. </p><p>Rather than manually stopping coding sessions to type detailed instructions or switch windows, developers direct the system hands-free. </p><p>The full-duplex engine dynamically decides when to speak, pause, or invoke tools, maintaining conversational state even as background agents process complex code modifications.</p><h2><b>Directing coding and complex builds with your voice alone</b></h2><p>The central operational capability in this update centers on multi-task execution across Codex and ChatGPT Work environments. </p><p>Software engineers can initiate multiple concurrent task threads from a single spoken prompt. For instance, a developer preparing to ship a feature can instruct the system to investigate an open authentication bug, review a pending API migration pull request, and generate missing unit tests simultaneously.</p><p>The desktop application coordinates these actions across disparate contexts, tracing issues through Slack conversations, GitHub repositories, and local codebases.</p><p>Developers can also verbally convert design mockups into working code, splitting tasks across frontend, backend, and testing layers. </p><p>With support for multi-folder projects (build 26.715) and remote execution via iOS, engineers can check task progress, answer agent prompts, and redirect active jobs without switching applications or managing individual processes line by line.</p><h2><b>Proprietary license</b></h2><p>OpenAI’s voice-enabled desktop release operates under a proprietary, commercial enterprise model. Access is restricted to paid subscribers across Plus, Pro, Business, Enterprise, and Education plans.</p><p>For individual developers and corporate engineering departments, this commercial structure means the model weights, voice processing pipelines, and agent state architectures remain fully closed. </p><p>Organizations cannot modify or self-host the underlying systems. Furthermore, tasks initiated via ChatGPT Voice consume standard usage allocations directly from existing Codex and ChatGPT Work plan quotas, treating voice-triggered actions identically to standard agentic workloads.</p><h2><b>Community reactions</b></h2><p>Developer communities immediately noted the implications of bringing continuous full-duplex voice to autonomous coding workflows. </p><p>Reacting to the build 26.715 release announcement—which details voice integration and multi-folder project support—AI Insider journalist <a href="https://x.com/ChrisGPT/status/2080375250139693293">@ChrisGPT noted on X</a>: &quot;Today OpenAI will release voice and remote guidance for codex ! One step closer to personal AGI&quot;. </p><p>Early technical feedback highlights widespread enthusiasm for orchestrating complex agentic tasks hands-free, particularly when stepping away from the workstation or managing build pipelines remotely.</p>]]></description>
            <author>carl.franzen@venturebeat.com (Carl Franzen)</author>
            <category>Orchestration</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/4j8qxL7Wc3EDLHQPUQjpUD/90a64f52fbda243e304a6e28dc556320/ChatGPT_Image_Jul_23__2026__03_54_24_PM.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start]]></title>
            <link>https://venturebeat.com/technology/black-forest-labs-launches-flux-3-capable-of-generating-images-and-20-second-video-with-audio-but-in-limited-release-to-start</link>
            <guid isPermaLink="false">7qbvNYBarUcjvFiVZ9ivde</guid>
            <pubDate>Thu, 23 Jul 2026 17:58:07 GMT</pubDate>
            <description><![CDATA[<p>Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with <a href="https://bfl.ai/blog/flux-3">today&#x27;s launch of FLUX 3</a>, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.</p><p>The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface. </p><p>That distinction is central to the company&#x27;s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company&#x27;s words, &quot;that can perceive, predict, and act across physical and digital environments.&quot; This release marks BFL&#x27;s first public video generation model. </p><div></div><p>FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a <a href="https://tally.so/r/44d9NX">gated &quot;Early Access&quot; program now</a>, to which anyone can apply, but which BFL must approve. </p><p>There is presently no public access through BFL&#x27;s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including <a href="https://venturebeat.com/technology/anthropic-says-its-most-powerful-ai-cyber-model-is-too-dangerous-to-release">Anthropic</a> and <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov">OpenAI</a>, though those were ostensibly for security concerns and due to government request. </p><p>What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.</p><p>Another big notable omission: FLUX 3 is <i>not</i> launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as &quot;open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction&quot; — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.</p><p>But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company&#x27;s commitment, but it is disappointing given the role open weights have played in FLUX&#x27;s adoption thus far. </p><h2><b>Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption</b></h2><p>BFL has published several benchmark comparisons, but they&#x27;re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability. </p><p>In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google&#x27;s Gemini Omni Flash in 52%.</p><p>One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a &quot;preliminary evaluation of an early FLUX 3 candidate&quot; — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.</p><p>Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.</p><p>Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0&#x27;s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.</p><p><a href="https://venturebeat.com/technology/googles-gemini-omni-flash-hits-the-api-turning-enterprise-video-production-into-a-conversation">Gemini Omni Flash</a>, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL&#x27;s own measurement, the two are indistinguishable on 10-second text-to-video quality. </p><p>Google&#x27;s advantage in that matchup is that Omni is generally available via Google&#x27;s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.</p><p>One regional wrinkle matters for a German company&#x27;s home market. Editing <i>uploaded</i> video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.</p><p>Here&#x27;s a rough guide for enterprises considering which video models to rely upon: </p><table><tbody><tr><td><p><b>Model</b></p></td><td><p><b>Max single-generation duration</b></p></td><td><p><b>Max resolution</b></p></td><td><p><b>Key constraints</b></p></td><td><p><b>Price per 10-second clip (720p)</b></p></td><td><p><b>Price per 10-second clip (1080p)</b></p></td><td><p><b>Price per 10-second clip (4K)</b></p></td></tr><tr><td><p>FLUX 3 Video </p></td><td><p><b>20 seconds </b></p></td><td><p>Not stated; evaluations run at 720p </p></td><td><p>Early access; no published SLA or pricing </p></td><td><p>Not announced </p></td><td><p>Not announced </p></td><td><p>Not announced </p></td></tr><tr><td><p>HappyHorse 1.1 </p></td><td><p>15 seconds </p></td><td><p>1080p </p></td><td><p>No 4K; closed weights </p></td><td><p>Not published (v1.0 reseller rate is ~$1.82) </p></td><td><p>Not published (v1.0 reseller rate is ~$3.12) </p></td><td><p>n/a </p></td></tr><tr><td><p>Veo 3.1 </p></td><td><p>Per-second billing </p></td><td><p><b>4K</b> </p></td><td><p><b>Supports clip extension; preview </b></p></td><td><p>$4.00 </p></td><td><p>$4.00 </p></td><td><p>$6.00 </p></td></tr><tr><td><p>Veo 3.1 Fast </p></td><td><p>Per-second billing </p></td><td><p><b>4K </b></p></td><td><p>Preview </p></td><td><p>$1.00 </p></td><td><p>$1.20 </p></td><td><p><b>$3.00 </b></p></td></tr><tr><td><p>Veo 3.1 Lite </p></td><td><p>Per-second billing </p></td><td><p>1080p </p></td><td><p>No 4K, no clip extension; preview </p></td><td><p><b>$0.50 </b></p></td><td><p><b>$0.80 </b></p></td><td><p>n/a </p></td></tr><tr><td><p>Gemini Omni Flash </p></td><td><p>10 seconds (3s minimum) </p></td><td><p>720p at 24 FPS </p></td><td><p>Preview abd no EU access</p></td><td><p>$1.00 </p></td><td><p>n/a </p></td><td><p>n/a </p></td></tr></tbody></table><h2><b>One architecture for media generation and physical action</b></h2><p>FLUX 3 builds on <a href="https://venturebeat.com/technology/black-forest-labs-new-self-flow-technique-makes-training-multimodal-ai">Self-Flow</a>, BFL&#x27;s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026. </p><p>The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.</p><p>&quot;We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,&quot; said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. &quot;True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.&quot;</p><p>He put the case more bluntly elsewhere in the announcement: &quot;You can&#x27;t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.&quot;</p><p>BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.</p><p>For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.</p><p>For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.</p><h2><b>What FLUX 3 Video can actually do</b></h2><p>The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation. </p><p>Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching <a href="https://developers.openai.com/api/docs/guides/video-generation">OpenAI&#x27;s discontinued Sora model.</a></p><p>The capability list BFL published covers:</p><ul><li><p>Text-to-video generation.</p></li><li><p>Image-to-video generation, either animating from a starting frame or using images as visual references.</p></li><li><p>Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.</p></li><li><p>Generative video-audio continuation from existing video and audio input.</p></li><li><p>Keyframe-to-video generation for controlled transitions between defined moments.</p></li><li><p> Multilingual dialogue.</p></li><li><p>A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.</p></li><li><p>Typography generation and animated design.</p></li><li><p>Agentic chaining of individual clips into longer, multi-shot sequences.</p></li></ul><p>That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.</p><p>It is also the capability where competition is most direct. HappyHorse 1.1&#x27;s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.</p><p>BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.</p><h2><b>FLUX-mimic tests whether video models can become robot models</b></h2><p>BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.</p><p>The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic&#x27;s robot-learning and production-deployment expertise in dexterous manipulation.</p><p>FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data. </p><p>BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.</p><p>&quot;The hardest part of robotics is data,&quot; said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. &quot;Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.&quot;</p><p>BFL<!-- --> argues that a model trained only on images cannot understand a world that &quot;moves, sounds, changes, and responds,&quot; and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni. </p><p>Its developer documentation cites &quot;world knowledge&quot; that combines &quot;an understanding of physics&quot; with Gemini&#x27;s grasp of history, science and cultural context. Its marketing is blunter still: &quot;Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,&quot; the company posted in June, crediting the model with &quot;an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.&quot; </p><p>The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it. </p><p>There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.</p><h2><b>Open weights helped make FLUX an industry standard</b></h2><p>BFL<a href="s"> officially launched in summer 2024 </a>and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises. </p><p>The company&#x27;s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and <a href="https://venturebeat.com/business/stable-diffusion-creators-launch-black-forest-labs-secure-31m-for-flux-1-ai-image-generator">Stable Diffusion</a>, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies. </p><p>That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research&#x27;s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.</p><p><a href="https://www.wired.com/story/black-forest-labs-ai-image-generation/"><i>Wired</i></a> magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley&#x27;s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.</p><p>FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, <a href="https://venturebeat.com/business/black-forest-labs-releases-flux-1-1-pro-and-an-api">released shortly after the firm&#x27;s launch,</a>  gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.</p><p>The company continued that pattern with <a href="https://venturebeat.com/ai/black-forest-labs-launches-flux-2-ai-image-models-to-challenge-nano-banana">FLUX.2 Dev</a> in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.</p><p>FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery.  BFL hasn&#x27;t yet shared information about its license, the parameter count, quantizations or hardware requirements.</p><p>The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows. </p><p>The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom&#x27;s T.Capital.</p>]]></description>
            <author>carl.franzen@venturebeat.com (Carl Franzen)</author>
            <category>Technology</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/1jXsrdqmpOmNlLZrXsL3hC/d83b416f6a0db2e9b871428ed55b2c2a/image__85___1_.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Multi-turn attacks broke AI models 88% of the time — single-turn testing missed it, Cisco AI security lead warns at VB Transform 2026]]></title>
            <link>https://venturebeat.com/security/openai-anthropic-google-and-xai-models-all-broke-under-multi-turn-attack-up-to-88-of-the-time</link>
            <guid isPermaLink="false">3xnPvZEt1cynHudemwNzuz</guid>
            <pubDate>Thu, 23 Jul 2026 17:13:48 GMT</pubDate>
            <description><![CDATA[<p>When Cisco ran 6,986 multi-turn attacks against <a href="https://blogs.cisco.com/ai/proprietary-problems">15 flagship models</a>, attackers who adapted across the conversation broke through as often as 88.3% of the time. Amy Chang, Cisco&#x27;s head of AI threat intelligence and security research, brought that finding to the agentic security panel at <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a>; the number should worry anyone still running single-turn red-teaming programs.</p><p><a href="https://venturebeat.com/resources/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials">VentureBeat&#x27;s June 2026 Pulse survey of 107 enterprise respondents</a> explains why the room was full. More than half, 54%, have already had a confirmed agent security incident (18%) or a near-miss caught before harm (36%). Just 32% give every agent its own scoped, managed identity, and fewer still, 30%, isolate their highest-risk agents in sandboxes. Provider-native and hyperscaler controls remain the primary agent security layer at <a href="https://venturebeat.com/security/shared-api-keys-expose-ai-agent-fleets-venturebeat-research">82% of companies surveyed</a>. The world&#x27;s largest security vendors have done the same math. </p><p>Palo Alto Networks closed its <a href="https://www.paloaltonetworks.com/company/press/2026/palo-alto-networks-completes-acquisition-of-cyberark-to-secure-the-ai-era">$25 billion acquisition of CyberArk</a> in February, CrowdStrike <a href="https://www.crowdstrike.com/en-us/press-releases/crowdstrike-to-acquire-sgnl-to-transform-identity-security-for-ai-era/">agreed in January to pay $740 million for SGNL</a>, and Cisco announced its <a href="https://blogs.cisco.com/news/cisco-announces-intent-to-acquire-astrix-security">intent to acquire Astrix Security</a> for a reported $400 million, all of it aimed at the identity and isolation layer most enterprises have not finished building.</p><div></div><p>Chang came to the panel with almost two decades of experience spanning cybersecurity operations, government, and the military. She ran global cybersecurity operations as an executive director at JPMorgan Chase, where she led the bank&#x27;s cyber threat intelligence teams, and served as a senior staffer on the House Foreign Affairs Committee and as a U.S. Navy Reserve officer. She also teaches cybersecurity and emerging threats as adjunct faculty at the Middlebury Institute of International Studies.</p><p>Chang&#x27;s 88.3% number comes from a study she co-authored with Nicholas Conley, built on 30,090 single-turn prompts and 6,986 multi-turn attacks against those 15 closed and proprietary flagship models. Multi-turn success rates ranged from 7.89% to 88.3%, every model tested showed non-trivial multi-turn exposure, and the two testing styles did not even rank the models in the same order. Cisco publishes adversarial evaluation signals for what is now 105 models on its <a href="https://leaderboard.aidefense.cisco.com/">LLM Security Leaderboard</a>, she told the audience.</p><p>&quot;If you don&#x27;t understand how models are susceptible to different types of attacks, then you are unable to account for how that model that is powering your agent, that is powering your application, to understand where those failure points are,&quot; Chang said. Single-turn testing is the one-shot malicious prompt, she explained, while extending an attack into a longer conversation &quot;is more realistic of how we are actually engaging with our models, with our agents, with our applications.&quot; That longer arc surfaces harmful outputs and misaligned behaviors that a snapshot never catches.</p><p>Cisco has pushed the testing itself into agentic territory. Chang described a framework where agents assess a deployment scenario, develop relevant attacks, judge whether they are worth pursuing, execute them, and evaluate their own success. What surprised her most, after all that sophistication, was how simple the defensive answer stays. &quot;The answer is still that it&#x27;s pretty simple,&quot; she said. &quot;You don&#x27;t have to get super creative. You just need to think about truly what are the fundamentals and basics of what I&#x27;m trying to secure in my organization.&quot;</p><p>Her starting point for CISOs beginning agentic deployments is Cisco&#x27;s <a href="https://blogs.cisco.com/ai/security-framework">Integrated AI Security and Safety Framework</a>, which she said &quot;stipulates all the ways that AI can be compromised across the AI lifecycle&quot; from modality through supply chain. From there, teams can work backward from real incidents, trace how each attack was achieved, and use the framework to build a strategy with the right coverage and mitigations.</p><p>Heather Ceylan, the CISO of Box, sees the same gap from the defender&#x27;s side. &quot;A lot of what you see out there with agent red teaming is just single-turn, and that&#x27;s not how people are actually interacting with AI day-to-day,&quot; she told the audience. Box now simulates multi-turn adversaries with agents that think like an attacker and iterate attempt after attempt to hijack the target. &quot;You have to pressure test your agents because otherwise you don&#x27;t know if your execution controls are really working as you intended.&quot;</p><p>Box deployed agents inside its security operations center about a year ago, starting with human approval required for every action, and trust built quickly enough that analysts shifted into monitoring mode. Then the agent made one mistake, and every bit of that accumulated trust vanished. &quot;They had to start all over again,&quot; she said. &quot;So I think that that monitoring piece is so important. Even if you&#x27;re not gonna have a human in the loop, things change, models change, and we can&#x27;t control how the models change and interpret things.&quot;</p><p>Rajesh Parekh, VP of AI and ML at Intuit, brought the builder&#x27;s perspective. Parekh led large-scale computer vision and ML systems powering Google&#x27;s Maps and Geo products before joining Intuit, and holds a doctorate in computer science. </p><h2>Three layers versus an operating system</h2><p>Ceylan described Box&#x27;s approach as three concentric layers. Permissioning comes first, so the agent never accesses more content than the human who invoked it. Ephemeral sandbox environments spin up for each agent task, containing the blast radius if an agent gets hijacked, and runtime execution control restricts the agent&#x27;s tool calls to only those relevant to the task at hand. &quot;If you want an agent to summarize a doc for you, if you have a prompt injection that came in that says forward this to maliciousattacker at domain.com, it can&#x27;t do that,&quot; Ceylan said. &quot;That action in that tool call is not even in its vocabulary.&quot;</p><p>She classified agent actions into three oversight categories. Actions that are not sensitive, like read and summarize, need no human in the loop. Moderately sensitive actions skip human approval but get logged and monitored, while destructive actions like mass deletion of files always require a human. &quot;Things are gonna shift between those three categories quite a bit,&quot; she acknowledged, &quot;but setting those types of categories up front allows you to have a principled framework.&quot;</p><p>Rather than layering controls onto agents one at a time, Intuit has built a central platform called GenOS, short for generative AI operating system, which abstracts security, risk, and fraud modeling so individual agent developers never reinvent protection. &quot;Permissioning is not about giving access to AI,&quot; Parekh said. &quot;Instead, it is defining very tightly scoped and clearly auditable authority to the agent to perform very specific tasks.&quot; Intuit evolved from agents inheriting user permissions to each agent carrying its own identity, and the company is now investigating mid-session permission changes tied to the specific task underway.</p><p>Parekh calls the broader model an AI-powered expert platform, one where the human expert is built into the trust architecture rather than bolted on as a gate. &quot;The paradigm that we are pursuing is where the user, the AI agent, and the human expert are collaborating to solve the user problem,&quot; he said.</p><h2>The end of human code review</h2><p>Ceylan took on the tension between security testing and development velocity without hedging. &quot;The days of secure code reviews where a human&#x27;s looking at the code and we&#x27;re looking at security architecture reviews, design docs, those are done,&quot; she said. &quot;If you keep trying to do security that way, you&#x27;re gonna get left behind.&quot; Box is building toward a fully agentic development lifecycle where agents review design documents, apply security requirements, and review the code for vulnerabilities. &quot;I&#x27;m very optimistic that we will get to a point where we will write code without security vulnerabilities because agents and the models are going to get so good at writing code without vulnerabilities,&quot; she said. &quot;We&#x27;re still a long way away from that.&quot;</p><p>Her advice for development teams skips the advanced AI concepts entirely and returns to basics that predate agents. &quot;It comes down to very basic least privilege access,&quot; she said. &quot;If you start giving your agents overly broad permissions at the beginning, it&#x27;s really hard to comb that back and build an infrastructure that allows for those ephemeral credentials and only those narrowly scoped tasks.&quot;</p><p>Parekh explained why the red teaming surface has expanded so quickly. &quot;These agents have skills, and skills could become vulnerabilities,&quot; he said. &quot;Agents have access to certain data, they have access to tools, and there could be threats that are lurking within those tools as well. So suddenly the blast radius of the malicious code or the intent increases dramatically.&quot; When Intuit identifies common vulnerability patterns from its manual red teaming exercises, it automates those tests back into the GenOS harness so future agents inherit protection and red teamers stay focused on new threat vectors. Runtime scanning of prompts and responses adds a final layer that can stop a suspect response and escalate to a human expert, he said.</p><p>&quot;You need to continuously test to ensure that those remain robust to the protections that you have built, as well as to account for any sort of drift or any other types of dependencies that you introduce into your scenario that can create novel vulnerabilities,&quot; he said.</p><h2>Intent versus probability</h2><p>An audience question about intent detection set off the sharpest exchange of the session. Ceylan noted that when Box&#x27;s own agent operates, the system always knows the user&#x27;s intent because it controls the prompt, which means guardrails and tool-call restrictions can be engineered around it. The harder challenge, which she admitted Box is still trying to solve, arrives when external agents connect and the context behind the request is opaque.</p><p>That exchange exposed a split running through the wider industry. Mastercard, in the fireside chat immediately preceding the panel, came down on the side of quantifying intent, building an open-source framework to propagate it as a standard because complex B2B procurement cannot work without that trust. Endpoint security CTOs, in briefings with VentureBeat, have gone the other way, saying they will bet on probability rather than intent inference for production workloads. Chang explained why models, as they are trained today, cannot reliably derive intent from a prompt, which is why deterministic controls and behavioral proxies remain necessary. Ceylan agreed that both are required. &quot;If you&#x27;re not doing anything deterministic, you&#x27;re really relying heavily on that intent, and I haven&#x27;t seen programs that are there yet,&quot; she said.</p><p>Ceylan&#x27;s story about trust collapsing after a single agent mistake landed as the panel&#x27;s most memorable moment because enterprise agentic security is not a problem that gets solved and stays solved. Models change, permissions drift, and adversaries adapt across multi-turn conversations that snapshot tests never capture.</p><p>For the 82% of enterprises relying on provider-native controls as their primary security layer, and the 59% shopping for agent security tooling over the next 12 months, the panel&#x27;s takeaway was blunt. Test the way attackers attack, across full conversations and continuously, or find out in production what your single-turn red teaming missed.</p>]]></description>
            <author>louiswcolumbus@gmail.com (Louis Columbus)</author>
            <category>Security</category>
            <category>VB Transform</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/63dEmbP0fFdy2U4vcWJkKZ/bc298524309906711b20a459a7e2e64a/2026-VB-Transform-Hotel-Nia-0245.jpg?w=300&amp;q=30" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs]]></title>
            <link>https://venturebeat.com/resources/the-ai-compute-gap-enterprises-are-buying-infrastructure-faster-than-they-can-measure-what-it-costs</link>
            <guid isPermaLink="false">3RCeHS5ENXpygHO6HPhX6f</guid>
            <pubDate>Thu, 23 Jul 2026 17:06:07 GMT</pubDate>
            <description><![CDATA[<p>Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price — which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap — heavy, fast-moving investment running ahead of the visibility needed to control it.</p><p>This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and — most revealingly — how well they can measure and control the economics of the compute underneath it all.</p><p>The central finding is a compute gap — the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold — 83% report GPU utilization of 50% or less — and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own.</p><p>Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter — unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions — the shift from GPU compute to memory bandwidth as inference scales — is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it.</p><p>This report is one of five in VentureBeat Research&#x27;s Q2 study of the agentic stack. Cost is the control with the least instrumentation: More than eight in ten GPU operators report utilization at half capacity or less, and a minority rigorously track compute cost and return. The executive summary, <a href="https://venturebeat.com/technology/venturebeat-research-where-enterprise-ai-agent-governance-hasnt-caught-up">&quot;VentureBeat Research: Where enterprise AI agent governance hasn&#x27;t caught up,&quot;</a> places that finding in the full pattern. </p><h2>Methodology</h2><p>VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.</p><p>By organization size the sample concentrates in the mid-market: 101–250 employees (36%) and 251–1,000 (27%) lead, with 1,001–5,000 (22%), 5,001–10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%).</p><p>At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators.</p><h2>Finding 1: Ambition outpaces production</h2><p><b>Only one in five run AI in production at scale</b></p><p>We asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale.</p><div></div><p>The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint — and whose costs — are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works.</p><h2>Finding 2: Enterprises run on hyperscalers and model APIs</h2><p><b>The specialized GPU clouds barely register — today</b></p><p>We asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents.</p><div></div><p>The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius and peers — register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from — which makes the evaluation intentions in Finding 3 all the more striking.</p><p><i>(A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses — an average of 2.1 selections each — so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google&#x27;s strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.)</i></p><h2>Finding 3: The next dollar goes to infrastructure they don’t yet run</h2><p><b>AI-specialized clouds top the evaluations list</b></p><p>We asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today.</p><div></div><p>Here is the report’s sharpest tension. The single most-cited planned evaluation area — AI-specialized clouds, at 45% — is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-Nvidia accelerators, and 28% in next-generation Nvidia silicon; even decentralized compute networks (16%) and sovereign compute (11%) draw meaningful interest. Read against current usage, this is not incremental — it is the leading edge of a re-platforming. The direction-of-travel question tells the same story: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum (+24), edging out even the hyperscalers (+22). Enterprises are preparing to move a meaningful share of AI compute off the general-purpose cloud.</p><p>This continues a trend we saw in our April-May survey wave. Back then, usage of the AI-specialized clouds was equally marginal — CoreWeave at 3%, Lambda at 4%, Crusoe at 2% of enterprises. When we asked enterprises what change they planned in their AI infrastructure strategy over the next twelve months, the most-cited answer was moving workloads to specialized AI clouds, at 33%. Asked in April-May which emerging compute option they were most likely to evaluate AI-specialized clouds again drew the most responses. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use.</p><h2>Finding 4: A switching wave is building</h2><p><b>Six in 10 plan to change providers within a year — many within a quarter</b></p><p>We asked whether and when enterprises plan to switch or add an infrastructure provider. Very few intend to stand still.</p><div></div><p>For a category as foundational as compute, this is a remarkable amount of intended movement. Only 36% have no plans to change, meaning a clear majority (64%) intend to switch or add a provider within twelve months — and 38% within the next quarter alone. Where that interest points is telling: the providers drawing the most switching consideration are again the incumbents — Microsoft Azure and Google Cloud (33% each), OpenAI (30%), and Gemini (22%) — which suggests much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest in Finding 3 is a 12-month evaluation thesis; the switching in the next quarter is mostly incumbents trading share.</p><p>(<i>Method note: Respondents who selected both &quot;no plans to change&quot; and a specific switching window are counted as switchers, on the logic that naming a timeframe is the more specific answer; three respondents were reclassified under this rule.</i>)</p><h2>Finding 5: Nobody buys on token price</h2><p><b>Integration and total cost of ownership decide — not sticker price</b></p><p>We asked what matters most when enterprises select an AI infrastructure provider. Headline price finished last.</p><div></div><p>Enterprises do not buy AI infrastructure on pricing, which is the place vendors compete on hardest. Integration with the existing stack (41%) and total cost of ownership (35%) dominate, while the headline metric — cost per million tokens — is the deciding factor for just 8%, dead last. The pattern is coherent: buyers are optimizing for how a provider fits and what it truly costs to operate, not for the advertised unit rate. It also foreshadows Finding 7 — enterprises say TCO matters most, yet most cannot yet measure it rigorously. The stated priority and the measured capability are out of step.</p><h2>Finding 6: Expensive GPUs, idle most of the time</h2><p><b>83% report GPU utilization of 50% or less</b></p><p>We asked what share of their GPU capacity enterprises actually utilize. The answer is a well-known but rarely quantified inefficiency.</p><div></div><p><i>Disclosure: Band percentages count every selection against all 107 qualified respondents; 14 respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50%</i></p><p>The compute already in place runs cold. Adding the bands at or below half capacity, 83% of enterprises that operate GPUs report utilization of 50% or less, and nearly half (49%) run at 25% or below. Only 12% clear the 50% mark, and a further 8% do not measure utilization at all. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large — and largely unmeasured.</p><h2>Finding 7: Spending fast, measuring slowly</h2><p><b>Fewer than half rigorously track what their compute costs</b></p><p>We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger lags the spending.</p><div></div><p>Measurement trails money. Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (20%), or have not prioritized it (6%). That gap is consequential given Finding 5, where total cost of ownership was the second-ranked buying criterion — enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure is moderately positive but not enthusiastic: on a five-point scale, overall satisfaction averages 4.0, with ease of implementation (3.8) and value for money (3.9) trailing slightly — the softness landing, tellingly, on cost. Enterprises are spending quickly and accounting slowly.</p><h2>Finding 8: The next bottleneck few are watching</h2><p><b>As inference shifts from compute to memory, the field scatters</b></p><p>Finally, we asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The responses reveal a frontier that is not yet a priority.</p><div></div><p>The memory frontier is real but barely governed. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises scatter: Dell leads at 31%, Nvidia follows at 16%, and the rest fragments across storage vendors, open-source tooling, and model-level efficiency techniques. Most telling is that roughly one in five (18%) either do not recognize the constraint or have not begun to address it. For a shift that will reshape inference cost and architecture, this is an early and unsettled market — and, consistent with the measurement gap in Finding 7, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one.</p><h2>The bottom line: A compute gap that faster spending will widen, not close</h2><p>Organizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending intentions point past their current stack — toward specialized clouds and alternative accelerators almost none of them run today — and a clear majority intend to change providers within the year. They buy on integration and total cost of ownership rather than headline price, which is rational; the difficulty is that most cannot yet see those economics clearly.</p><p>The visibility gap is concrete. The GPUs enterprises already own run at half utilization or less for the overwhelming majority, and fewer than half can rigorously track what their compute costs or returns. Satisfaction is decent but unenthusiastic, softest on value for money — the dimension hardest to judge without measurement. And the next constraint, the shift from compute to memory in large-scale inference, is arriving while most enterprises are still unaware of it. At 107 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market and earlier-stage adopters — but the direction is consistent: the appetite to spend is running well ahead of the instrumentation to spend well. The compute gap is not a capacity problem that more hardware will solve on its own; it is, first, a problem of seeing what the hardware already costs. The open question for later waves is whether enterprises build that visibility before the re-platforming arrives — or buy the next layer of infrastructure as blind to its economics as the last.</p><hr/><p><i>Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally rather than as a month-over-month trend, and at 107 respondents this is a directional signal rather than a precise measurement — the sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.</i></p>]]></description>
            <category>Resources</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/65A33lcUi9p0nBSloUI1Wo/5e5d26295bc879f0ea8845cecac65504/VentureBeat-Research.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
    </channel>
</rss>