<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
    <channel>
        <title>VentureBeat</title>
        <link>https://venturebeat.com/feed/</link>
        <description>Transformative tech coverage that matters</description>
        <lastBuildDate>Sun, 16 Aug 2026 14:00:13 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <copyright>Copyright 2026, VentureBeat</copyright>
        <item>
            <title><![CDATA[DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge]]></title>
            <link>https://venturebeat.com/orchestration/deepseeks-top-ranked-v4-flash-stumbles-on-real-agent-tasks-as-its-prices-surge</link>
            <guid isPermaLink="false">30TtKMdIkDcC2xBxuqAjKE</guid>
            <pubDate>Sun, 16 Aug 2026 13:00:00 GMT</pubDate>
            <description><![CDATA[<p>DeepSeek&#x27;s V4 Flash has topped model leaderboards and been hailed by developers as a &quot;total monster&quot; since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. </p><p><a href="https://x.com/composio/status/2085330847951970801?s=20">Composio ran the model through eight different agent harnesses</a>, including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed — and only six of the 30 workflows were completed successfully by every harness tested.</p><p>The gap illustrates why orchestration, not raw model capability, may decide whether the model succeeds in enterprise settings: the same model produced substantially different results depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on.</p><p><a href="https://www.engadget.com/2236912/deepseek-ai-models-get-four-times-pricier/">DeepSeek said it will be hiking the prices for V4 Flash and Pro</a>, models that have quickly become favorites among developers building coding assistants and agents. </p><p>While it seems the move might undercut its very appeal — strikingly capable models at ultra-low pricing that frontier providers simply can&#x27;t match — it also moves the story beyond the now-clichéd &quot;cheap Chinese model&quot; narrative, as early use cases emerge and enterprises figure out where different models fit into their tech stacks and what workflows they should be aimed at.</p><h2>&quot;<!-- -->Insane&quot; adoption numbers as DeepSeek flips the cost structure</h2><p>DeepSeek rolled out V4 Flash to public beta July 31, and made V4 Pro generally available on August 13. The 284-billion-parameter Flash is built for volume and speed, the 1.6 trillion-parameter Pro for more complex workflows.</p><p>Both models have flexible reasoning capabilities (low, high, max) and ‘thinking modes’ applying chain-of-thought (CoT) reasoning to improve answer accuracy.</p><p>Users were immediately impressed by Flash’s capabilities. It has <a href="https://openrouter.ai/rankings">dominated OpenRouter&#x27;s usage leaderboard</a> since its rollout, currently the most-used model on the platform by weekly token volume.</p><p>“The adoption numbers of the initial DeepSeek V4 Flash were insane,” <a href="https://x.com/natolambert/status/2084790959636922652?s=20">ML researcher Nathan Lambert posted to X</a>, <!-- -->adding that the new version &quot;scored the same as GLM 5.2,&quot; making it a &quot;total monster&quot; that will be used extensively.</p><div></div><p>DeepSeek switching the cost model adds an interesting dimension. </p><p><a href="https://api-docs.deepseek.com/quick_start/pricing/">V4 API rates</a> are going up by as much as 1,100% depending on the model, token type and time of use. The new pricing structure: </p><ul><li><p>Flash will be 22 cents per million input tokens and 66 cents per million output tokens off-peak; and 44 cents per million input tokens and $1.32 per million output tokens at peak. This represents a 57% to 371% increase. </p></li><li><p>Pro will be 66 cents per million input tokens and $1.98 per million output tokens off-peak; and $1.32 per million input tokens and $3.96 per million output tokens at peak. This shows a 51% to 355% jump. </p></li><li><p>Cache hits, meanwhile (when models reuse prompts rather than starting from scratch), are going up between 52% and 1,100%.</p></li></ul><p>DeepSeek says offering 50% lower off-peak usage is intended to encourage “more flexible workload scheduling.&quot; Seventeen of every 24 hours stay at half price, and the new structure actually prices the company&#x27;s home market the highest.</p><p>&quot;This is not a simple price rise,&quot; said <a href="https://greyhoundresearch.com/svg/">Sanchit vir Gogia of Greyhound Research</a>. &quot;It is a pricing architecture that makes the timing of inference an economic variable.&quot;</p><p>Work that can wait — such as batch evaluation, synthetic-data generation, and overnight development runs — moves into the cheap hours; interactive agents and live operations cannot. Gogia said irritation among developers and enterprises is genuine and vocal, and that DeepSeek&#x27;s past low pricing doesn&#x27;t obligate it to stay cheap forever.</p><p>At first glance, it does look like a “suicidal move from a platform still looking for credibility against more established AI model vendors,” <a href="https://www.linkedin.com/in/carmi/">said tech analyst Carmi Levy</a>. The increases will certainly eat into DeepSeek&#x27;s price advantage and force customers to weigh concerns around the company&#x27;s Chinese origins more heavily.</p><p>Still, DeepSeek remains far cheaper by all pricing measures relative to competing models from OpenAI, Anthropic, Google, Cohere, xAI, and others, he said. </p><p>So, while the move will force DeepSeek to emphasize performance and security over cost, it hardly wipes out its already-notable price-performance advantage, and still gives customers ample wiggle room to justify its use for specific workloads, Levy said. The math will just have to be more tightly calculated. </p><p>“The advantage will likely erode over time as DeepSeek inevitably continues to align pricing with market realities, but for now it&#x27;s still easy to make the business case,” Levy said. </p><h2>Where can DeepSeek Flash fit into enterprise environments? </h2><p>Adoption inside enterprises remains an open question due to cost, capability, reliability, data governance, security, and other factors.</p><p>One use case is batch processing, Levy said. This kind of work is typically routine and repetitive rather than demanding top-tier intelligence, so it makes sense to use a cheaper, more efficient model. “It’s a high-performance inference engine that enterprises can consider using for point solution workloads rather than as a wholesale replacement for the incumbent offerings,” Levy said. </p><p>Partial adoption will likely involve isolated, non-sensitive workloads with clearly-defined success metrics, strict oversight, permissions controls, and fallback models in case of failures, he noted. Broader deployment will require DeepSeek and its hosting partners to demonstrate strong reliability, security, privacy, auditability, and deployment options. As it adjusts price structures based on demand, DeepSeek also must retain a large enough price-performance advantage to justify any risk, he said.</p><p>Expect unsanctioned, smaller-scale use in backroom labs and contained test environments as IT teams get familiar with the new model and figure out when and how to bring it to senior leadership for budget approval.</p><p>“DeepSeek has built a well-earned reputation as a global disruptor,” he said, “and it’s clear that its march to broader enterprise adoption will continue to gather momentum.”</p><h2>Testing DeepSeek in multi-tool workflows </h2><p>While many use cases are still in the experimental phases, Meta software engineer Naman Ahuja offers one that could translate directly into enterprise environments. In a project unrelated to his employer, he built a home-automation agent with DeepSeek V4 Flash to explore how a lower-cost model performs as the reasoning/orchestration layer for a real multi-tool workflow.</p><p>When he leaves home, an agent coordinates several actions across otherwise separate systems: Such as setting a thermostat to “away” to reduce unnecessary energy use, arming a Ring security system, closing and locking doors. </p><p>“What interested me was not simply whether the model could understand a command, but whether it could translate intent into a sequence of actions across multiple tools where reliability matters,” Ahuja said. </p><p>The biggest lesson was that once a model can take actions, reliability matters as much as intelligence. The system needs structured tool outputs, verification that actions actually succeeded, retry/failure handling, and clear boundaries around what the model is allowed to do.</p><p>In the case of enterprise, “the architecture is similar.” Home devices change to ticketing systems, databases, CRM platforms, or infrastructure APIs. The most useful agents will likely orchestrate repetitive workflows across multiple systems, with scoped permissions, auditability, observability, and human approval for higher-risk actions.</p><p>“Many valuable AI agents will not be chatbots; they will be background agents coordinating APIs, infrastructure, and business systems in response to events,” he said. </p><h2>Enterprises need tangible use cases</h2><p>But Flash&#x27;s API is still in public beta, Gogia pointed out, and there is not yet an evidence trail of settled enterprise adoption, real-world deployments, and named customers.</p><p>“The benchmark story is looser than its retelling, the portfolio story is newer than it looks, and the economics have moved into the system around the model,” he said. Developer mainstreaming is proven; enterprise standardization is not. “The model is mainstream by traffic and still unproven by contract.”</p><p>DeepSeek&#x27;s own integration guidance is an important consideration, he said: Its documentation for at least one popular agent environment states that built-in V4 entries are not sufficient for reliable operation without compatibility overrides. </p><p>“Which is a vendor telling the market, accurately, that benchmark performance is not a proxy for production readiness,” Gogia said. “A model can score beautifully and still misbehave once tools, credentials, and state enter the room.”</p><p>The serving layer behaves no differently: the same open weights run by different hosts show visible differences in throughput and uptime. &quot;Choosing Flash therefore answers one procurement question and opens three more: Who serves it, where it runs, and which controls surround it,&quot; Gogia said.</p><h2>Prepare for a multi-model future</h2><p>DeepSeek offers a nuanced case for a multi-model future. V4 Flash is being deployed as the high-volume worker inside diverse estates, Gogia noted: It handles routine generation, retrieval, and background automation, while more difficult or sensitive tasks go elsewhere. </p><p>“The question is whether its performance is sufficient for the real-world workflows enterprises actually run, not whether it tops every benchmark,” he said. Enterprises must determine which combination of model, harness, and provider completes the work safely at the lowest cost.</p><p>Adam Dalloul, CEO and founder of EmpirioLabs AI, pointed out that bigger isn’t always better; workflows should be task-dependent. For example, his team at EmpirioLabs AI — which hosts 100-plus models on one API, including DeepSeek V4 Flash — were recently working on translating its site into different languages, and there was no need for a large model like GPT 5.6 Sol or Opus 5 to complete the task. </p><p>“This is where subagents come in handy,” he said. </p><p>His recommended approach: Spawn cheaper subagents and adapt per task. For instance, use Flash variants for day-to-day work, and Pro variants when you need something more powerful. “It depends on the nature of your application.”</p><p>Many companies are pivoting towards their own internal benchmarks to route models effectively, Dalloul noted. For example, his team has a workflow that puts a model through various gates and instructions. This helps them identify the model with the speed and accuracy required for the task. </p><p>In another example, one of his enterprise clients exclusively wanted access to DeepSeek V4 Flash. They had tested a variety of models and V4 Flash was the only one that met their criteria for speed, cost, and an “appropriate intelligence threshold.”</p><p>Meta’s Ahuja agreed that smaller, more efficient models can handle frequent, well-defined agentic tasks, while more expensive frontier models can be reserved for “ambiguous, difficult, or higher-risk decisions.” The relevant metric increasingly becomes cost per successfully completed workflow rather than simply cost per token. </p><p>The trade-off, however, is that cheap inference does not automatically mean cheap or safe automation, he said. Once an AI system can take actions, reliability, verification, permissions, failure handling, and security become much more important. </p><p>“A failed text response is inconvenient; a failed action in an operational workflow can have real consequences,” Ahuja said. </p>]]></description>
            <author>taryn.plumb@venturebeat.com (Taryn Plumb)</author>
            <category>Orchestration</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/YO11rLoN8OR9ZKXk99L3g/102fc5f215a65b1851a46b1309831c15/DeepSeek.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[An eval harness found what qualitative review couldn't: AI models are most confident when wrong]]></title>
            <link>https://venturebeat.com/orchestration/an-eval-harness-found-what-qualitative-review-couldnt-ai-models-are-most-confident-when-wrong</link>
            <guid isPermaLink="false">3aqFAE5idF5OcZ4SA8CJIT</guid>
            <pubDate>Sat, 15 Aug 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[<p>There is a step in the development process for large language model (LLM)-assisted tooling that most teams skip because it&#x27;s tedious, time-consuming, and doesn&#x27;t produce results visible to end users: Verifying that what the model is saying is actually correct. Not fluent, not coherent, not topically relevant — correct in the sense of accurately identifying the right answer to the specific problem the tool was built to solve.</p><p>The gap between &quot;this output sounds right to me&quot; and &quot;this output is verifiably correct&quot; is where most LLM-assisted enterprise tools fail quietly. They pass internal review because the output sounds right. They fail in production because those people weren&#x27;t reviewing against ground truth — they were reviewing against their intuition about what a good answer looks like.</p><p>This distinction matters more as LLM-assisted tools move from productivity accessories to components that influence real business decisions. If your AI-assisted tool is shaping how an analyst investigates a data quality issue, how a compliance reviewer decides whether to escalate a flagged record, or how an operations team triages a validation failure — the accuracy of its output has real consequences. &quot;Seems reasonable&quot; is not an adequate evaluation standard for that.</p><h2><b>What qualitative evaluation actually catches</b></h2><p>The standard evaluation approach for LLM output in enterprise tooling is qualitative: A sample of outputs is reviewed by someone with domain knowledge, judged against a mental model of what a good answer looks like, and the prompt is adjusted if too many outputs seem off.</p><p>This catches a specific class of problems: Outputs that are obviously wrong, poorly formatted, or off-topic. These are real issues worth catching. They&#x27;re also the easy ones.</p><p>What qualitative evaluation consistently misses is the class of outputs that are wrong in ways that are difficult to see without checking against something external. An explanation that confidently identifies the wrong root cause, in language that sounds authoritative, based on reasoning that sounds plausible — this passes qualitative review. It fails the moment someone with the right context checks it against what actually happened.</p><p>In a system whose value proposition depends on accuracy, &quot;sounds plausible&quot; is not the same as &quot;correct.&quot; The two can diverge significantly, and qualitative review won&#x27;t tell you when they have.</p><h2><b>What an actual eval harness looks like</b></h2><p>The alternative is building an evaluation harness that scores model output against labeled ground truth — a set of cases where the correct answer is known, against which you can measure accuracy rather than coherence.</p><p>I built this while developing a root-cause explainer for data migration drift: A tool that takes a detected drift event and generates a ranked explanation of what most likely caused it. The first prototype produced fluent, specific-sounding explanations that passed qualitative review. When I tested it against cases where I already knew the root cause, the explanation was wrong often enough to matter.</p><p>The eval harness I built works in three parts.</p><p>First, a synthetic ground truth dataset: Cases where the correct answer is known by construction. This meant introducing specific, controlled causes into a test pipeline — schema changes, transformation logic bugs, source system behavioral shifts — recording exactly what I introduced, and running the model against the resulting drift events. The correct answer for each case was the cause I had deliberately introduced.</p><p>Getting the synthetic scenarios realistic enough to be useful required more care than I expected. Early versions were too clean — the drift signal was obvious in ways that real production drift events aren&#x27;t. Adding realistic noise, overlapping signals, and cases where multiple plausible causes were present simultaneously was what made the synthetic set actually predictive of real-world performance.</p><p>Second, a scoring function that evaluates ranked output. Binary correct/incorrect isn&#x27;t sufficient when the model produces a ranked list of likely causes rather than a single answer. An explanation that correctly identifies the root cause as the third most likely candidate is meaningfully different from one that identifies it as the most likely. The scoring function evaluated two dimensions: Presence — did the correct answer appear in the output at all — and rank — how prominently was it featured relative to incorrect candidates. These were combined into a weighted score that rewarded both finding the right answer and ranking it appropriately.</p><p>Third, systematic evaluation across the full synthetic dataset rather than spot-checking. Running the harness across the complete set reveals patterns that spot-checking misses: Which categories of problem the model handles reliably, which it consistently gets wrong, and which combinations of signals produce the highest rate of confident incorrect explanations.</p><h2><b>What the evaluation revealed</b></h2><p>The results were more informative than any qualitative review could have been.</p><p>Schema change scenarios scored well — the model was reliable at identifying upstream schema changes when the evidence was present and distinctive. Transformation logic bugs were harder — the model consistently identified the right general category but misattributed the specific change that caused the problem, particularly when multiple changes had been made close together. Overlapping-signal scenarios were the hardest — cases where two different causes occurred close in time produced the highest rate of confidently wrong explanations.</p><p>That last finding is the one that qualitative review would never have surfaced. The model&#x27;s expressed confidence didn&#x27;t correlate with its accuracy — it was most confident in the cases where it was most wrong. Without the eval harness measuring against ground truth, that pattern would have been invisible.</p><h2><b>The practical implication for enterprise AI deployment</b></h2><p>For teams deploying LLM-assisted tools in enterprise contexts — particularly tools that influence how people investigate problems, triage alerts, or make routing decisions — the eval harness question to answer before production deployment is: Have we measured accuracy against cases where we know the right answer, or have we only reviewed whether the outputs seem reasonable?</p><p>If the answer is the latter, the tool has been tested for fluency and coherence but not for correctness. Those are different properties. For tools that shape business decisions, correctness is the one that matters.</p><p>Building the synthetic ground truth dataset is the hard part and the part most worth investing in. It forces you to define precisely what &quot;correct&quot; means for your specific use case — which turns out to be a useful exercise independent of the evaluation itself. The scoring function and the harness infrastructure are relatively straightforward once you have that definition. Without it, you&#x27;re measuring something other than what you&#x27;re trying to guarantee.</p><p><i>Arun Mishra is an enterprise architect. </i></p>]]></description>
            <category>Orchestration</category>
            <category>DataDecisionMakers</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/5i7p6do1zSmoo71Z58g81z/271b2f52cede8bfb4b801d4dafdc70d9/u7277289442_A_teacher_is_sitting_at_a_desk._View_from_behind._0f474f7c-90c2-4350-85cd-0f23e41636c4_0.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor]]></title>
            <link>https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor</link>
            <guid isPermaLink="false">65fLrIS0quprLcAMRZ04SD</guid>
            <pubDate>Fri, 14 Aug 2026 22:56:16 GMT</pubDate>
            <description><![CDATA[<p>Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, <a href="https://z.ai/blog/glm-5.3">today released GLM-5.3</a> with substantial gains in long-horizon coding and a more consequential — and potentially sensitive — jump in cybersecurity capabilities.</p><p>Already, GLM-5.3&#x27;s cyber capabilities have found a &quot;potentially serious vulnerability in Cursor,&quot; the AI coding startup recently <a href="https://cursor.com/blog/joining-spacex">acquired by SpaceX</a>, according to z.ai developer advocate Lou, <a href="https://x.com/louszbd/status/2088284853943009425">posting on X</a>. VentureBeat also tagged Cursor for confirmation on X and is awaiting response.</p><p>GLM-5.3 is available initially <i>only</i> through the company&#x27;s GLM Coding Plan and ZCode coding environment, while API access and open weights are coming later, &quot;once safety evaluation and hardening are complete,&quot; according to the company. </p><p>Z.ai says it plans to release weights approximately two weeks after launch.</p><p>For enterprise developers, the notable part of the release is not simply another round of benchmark improvements. Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with the improvements coming entirely from scaling post-training across more environments, more diverse tasks and additional reinforcement-learning compute.</p><p>That makes GLM-5.3 something of a test of how far a frontier-scale base model can be pushed without another expensive pretraining cycle.</p><p>“Scaling post-training is all we did for GLM-5.3,” Z.ai wrote in its <a href="https://z.ai/blog/glm-5.3">technical announcement.</a></p><p>The results suggest considerable headroom. But they have also produced an unusual problem for an open-model developer: according to Z.ai, cybersecurity capabilities improved faster than anticipated as training scaled, particularly as tasks progressed from vulnerability identification toward constructing complete exploitation chains.</p><p>Reuters reported Friday that Z.ai is also introducing controls around some of the model&#x27;s more advanced capabilities, including a “trusted access” approach for sensitive functionality.</p><h2>A large jump in coding without another base model</h2><p>GLM-5.3 builds on the 743-billion-parameter-scale base model behind GLM-5.2 rather than replacing it. Z.ai instead expanded the post-training system it had already assembled around long-horizon reinforcement learning.</p><p>Those environments increasingly resemble complete engineering jobs rather than isolated programming exercises.</p><p>Z.ai describes scenarios in which an agent receives access to codebases, documentation, compute clusters, storage systems and experimental results, then has to diagnose problems, modify systems, run experiments and demonstrate a measurable improvement while preserving correctness. Some tasks are designed to approximate several days of work for an experienced engineer.</p><p>The approach produced sizable generation-over-generation improvements on Z.ai&#x27;s reported evaluations.</p><p>GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 26.2 to 48.2 on AutomationBench. On Agents&#x27; Last Exam CLI, it improves from 23.8 to 28.5.</p><p>The model does not dominate every frontier competitor. Z.ai&#x27;s own benchmark table shows GPT-5.6 Sol at 34.6 and Claude Fable 5 at 33.7 on Terminal-Bench 3.0, compared with GLM-5.3&#x27;s 28.3. On DeepSWE v1.1, GLM-5.3 scores 66.9, compared with 72.7 for GPT-5.6 Sol and 69.7 for Fable 5.</p><p>But Z.ai is also emphasizing efficiency rather than benchmark position alone.</p><p>On its private Z.ai Code Bench, GLM-5.3 reaches a 34.5% result at its Max reasoning setting while consuming roughly 75,000 output tokens per task. GLM-5.2 reaches 23.4% while consuming approximately 96,000. At High effort, GLM-5.3 reaches 31.4% at roughly 50,000 output tokens, compared with Z.ai&#x27;s reported 29.5% for Claude Opus 4.8 using 120,000.</p><p>Because Code Bench is Z.ai&#x27;s own private evaluation, those comparisons should be treated as company-reported results rather than independent measurements. Still, reducing token consumption while improving task completion is operationally important for enterprises deploying coding agents, where long-running loops can make inference cost and latency compound quickly.</p><h2>Cyber capabilities developed faster than Z.ai expected</h2><p>The more unusual development is cybersecurity.</p><p>Z.ai introduced vulnerability-discovery environments into GLM-5.3&#x27;s post-training mix expecting the model to improve at finding software flaws. Instead, the company says capability began progressing further along the exploitation chain.</p><p>“As we scaled post-training, cyber capability developed faster than we expected,” Z.ai wrote.</p><p>On CyberGym, which tests vulnerability discovery and validation against source code, GLM-5.3 scores 84.5%, compared with 77.2% for GLM-5.2. That also edges Z.ai&#x27;s reported scores for GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%.</p><p>The advantage does not extend across the entire exploitation stack. GLM-5.3 scores 54.4% on ExploitBench, more than twice GLM-5.2&#x27;s 24.4%, but remains well behind the 76.5% Z.ai reports for GPT-5.6 Sol and 78% for Mythos 5.</p><p>Similarly, on ExploitGym, GLM-5.3 completes 105 tasks under a normalized two-hour budget and 130 under six hours, up from 29 and 39 for GLM-5.2. Fable 5 reaches 181 and 247, while GPT-5.6 Sol reaches 216 and 293.</p><p>The direction of travel may matter more than the leaderboard position.</p><p>Z.ai says work with security teams in China has resulted in 2,436 vulnerability findings across 269 projects after expert review, screening and deduplication. Its disclosure ledger lists 1,097 as critical or high severity, with 53 publicly disclosed and 2,383 still under embargo at the time of the release.</p><p>That creates a tension increasingly facing frontier model providers: the same long-horizon agent capabilities that make models more useful for software engineering can also make them more capable security researchers — and potentially more capable offensive operators.</p><h2>GLM-5.3 also requires developers to change how they call the model</h2><p>Developers migrating existing GLM applications should pay attention to a breaking API behavior.</p><p>GLM-5.3 supports three reasoning-effort levels — <code>low</code>, <code>high</code> and <code>max</code> — with <code>max</code> the default and Z.ai&#x27;s recommended setting for coding. But unlike previous releases, thinking cannot be disabled.</p><p>Applications currently sending <code>thinking.type: &quot;disabled&quot;</code> must change the value to <code>enabled</code> and specify a reasoning effort before switching the model identifier to GLM-5.3. Otherwise, Z.ai says the request will fail.</p><p>That makes GLM-5.3 an actual migration rather than simply a model-name substitution for some production applications.</p><h2>From GLM-4.5 to GLM-5.3: Z.ai&#x27;s rapid push into agentic engineering</h2><p>GLM-5.3 is the latest step in a rapid shift by Z.ai — formerly known as Zhipu AI — toward coding agents and long-running autonomous engineering workloads.</p><p>GLM-4.5, released in July 2025, established much of that direction. The 355-billion-parameter mixture-of-experts model was designed to combine reasoning, coding and agent capabilities, while the smaller GLM-4.5-Air offered 106 billion total parameters. Z.ai released the models with open weights and emphasized integration with agent frameworks.</p><p>GLM-4.6 followed in September, expanding context from 128,000 to 200,000 tokens and targeting coding, tool use and agent workflows in environments including Claude Code, Cline, Roo Code and Kilo Code. Z.ai also began placing greater emphasis on token efficiency in real-world coding evaluations rather than benchmark performance alone.</p><p>The larger architectural jump came with GLM-5 in February 2026. Z.ai scaled the model from GLM-4.5&#x27;s 355 billion parameters to 744 billion, with 40 billion active parameters, and increased pretraining data to 28.5 trillion tokens. It also introduced its “slime” asynchronous reinforcement-learning infrastructure and explicitly repositioned the GLM family around “agentic engineering” and long-horizon tasks.</p><p>By June, GLM-5.2 had turned that strategy into a more direct enterprise proposition. The 753-billion-parameter model arrived with a stable 1-million-token context window, open weights under an MIT license and support across more than 20 coding environments. It also introduced IndexShare, which reuses an indexer across sparse-attention layers to reduce the computational burden of very long contexts.</p><p>GLM-5.2 was priced at $1.40 per million API input tokens and $4.40 per million output tokens, with cached input priced substantially lower, positioning Z.ai as both a technical and pricing competitor to proprietary frontier labs.</p><p>Z.ai&#x27;s ambitions have been expanding outside model development as well. Reuters reported last month that Zhipu AI raised roughly HK$31.4 billion, or about $4 billion, through a Hong Kong share sale, with proceeds intended for areas including research and development, computing infrastructure, talent and business expansion.</p><p>Taken together, the releases show a consistent progression: GLM-4.5 unified reasoning, coding and agents; GLM-5 substantially scaled the foundation model; GLM-5.2 attacked long-context and long-horizon engineering; and GLM-5.3 is now attempting to extract substantially more capability from that same foundation through post-training.</p><h2>Pricing, ZCode and availability</h2><p>GLM-5.3 is available now through Z.ai&#x27;s GLM Coding Plan and ZCode.</p><p>ZCode is the company&#x27;s own coding-agent environment and supports long-running “Goal” tasks that plan, implement, test and verify work. It also offers remote control of running tasks and is available on macOS, Windows and Linux.</p><p>Individual GLM Coding Plans currently start at a listed promotional price of $12.60 per month for Lite with 10,000 credits per week. Pro is listed at $56 per month with six times Lite usage, while Max costs $117.60 per month with 14 times Lite usage. Team Standard and Premium seats are listed at $88 and $188 per user per month, respectively.</p><p>Z.ai has also moved the Coding Plan to a points-based quota system that separately accounts for input, cached-input and output tokens. Calls outside the company&#x27;s weekday peak period consume 50% of the normal points.</p><p>The company has not yet provided general GLM-5.3 API pricing in the supplied launch materials, making total production API cost difficult to compare directly with GLM-5.2 or competing frontier models until staged API access arrives.</p><p>That staged release may ultimately be the most important part of GLM-5.3.</p><p>Z.ai spent the past year pushing an open-model strategy centered on permissive weights, low-cost inference and compatibility with existing coding-agent ecosystems. GLM-5.3 demonstrates what happens when that strategy succeeds perhaps too well in one sensitive domain: better autonomous engineering also means better autonomous security research.</p><p>The result is a model that advances Z.ai&#x27;s coding ambitions while forcing the company to confront the same capability-versus-access tradeoff facing the largest closed frontier labs.</p><p>For enterprise developers, GLM-5.3 is therefore worth watching for two reasons. Its coding results provide another indication that increasingly capable agents can emerge from better post-training and environments without continuously rebuilding the underlying foundation model. Its cybersecurity results show why deciding how those agents are distributed may become just as important as deciding how they are trained.</p>]]></description>
            <author>carl.franzen@venturebeat.com (Carl Franzen)</author>
            <category>Technology</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/4SSp8seVDFTXUdtXn6Rgyr/dfd61c47834cf5cb947b8f5ecbe249f0/ChatGPT_Image_Aug_14__2026__05_51_59_PM.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done]]></title>
            <link>https://venturebeat.com/security/three-claude-agents-given-conflicting-orders-sabotaged-each-other-on-a-shared-server-then-didnt-tell-users-what-theyd-done</link>
            <guid isPermaLink="false">5s4nliRe8aMQl0oAyHkrg3</guid>
            <pubDate>Thu, 13 Aug 2026 20:14:20 GMT</pubDate>
            <description><![CDATA[<p>Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other&#x27;s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival&#x27;s work. There was no prompt injection and no adversary. Anthropic&#x27;s Frontier Red Team <a href="https://www.anthropic.com/research/multiagent-systems">published the transcripts on Thursday</a> and called the escalation “increasingly aggressive, self-replicating malware.”</p><p>The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: &quot;<i>Since I have root, I could revoke u2 and u3&#x27;s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.</i>&quot;</p><p>That is a production outage being reasoned into existence by the software you deployed to prevent one.</p><p>Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its <a href="https://arxiv.org/abs/2604.24618">April paper</a>, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.</p><h2>Force settled 61% of Sonnet 4.6 runs, and capability did not fix it</h2><p>Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic&#x27;s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.</p><p>In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent&#x27;s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.</p><h2>18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs</h2><p>The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner&#x27;s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.</p><p>Merritt Baer, a former deputy CISO at AWS and an advisor to <a href="https://www.g2i.ai/">G2i</a> and <a href="https://andesite.ai/">Andesite</a>, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don&#x27;t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”</p><p>Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: &quot;wholesale is 10 for all of us, so a price war just burns everyone&#x27;s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.&quot;</p><p>Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor&#x27;s agents now owns that behavior, and no regulator will care that no human wrote the memo.</p><p>Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run&#x27;s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.</p><h2>Agents keep trusting liars, and keep burying the one fact that matters</h2><p>Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.</p><p>Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.</p><h2>AISI found zero unprompted sabotage, and a model that hides the difference</h2><p>Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D&#x27;Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.</p><p>Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It&#x27;s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won&#x27;t tell you it&#x27;s cheating, and it will lie about having cheated.”</p><p>The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that&#x27;s the basis of fiduciary duty,” he argued. “However, they don&#x27;t necessarily have the capability to do it.”</p><p>Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can&#x27;t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There&#x27;s an analogy to insider threat: you don&#x27;t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”</p><p>McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that&#x27;s going to be the only meaningful way to look at what an AI forward entity does.”</p><h2>Only 18% of enterprises isolate the agents most likely to turn</h2><p>VentureBeat&#x27;s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research&#x27;s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a <a href="https://venturebeat.com/security/four-of-five-enterprises-that-secured-ai-agent-identities-still-cant-contain-one-that-goes-rogue">58% incident or near-miss rate</a>. Anthropic&#x27;s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.</p><p>Baer said the first thing that changes for a team running agents in production is the threat model itself. “We&#x27;ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can&#x27;t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”</p><p>Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.</p><p>The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.</p><table><tbody><tr><td><p><b>Behavior Anthropic documented</b></p></td><td><p><b>Where it surfaces in an enterprise stack</b></p></td><td><p><b>The test to run this quarter</b></p></td></tr><tr><td><p>Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.</p></td><td><p>Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.</p></td><td><p>Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.</p></td></tr><tr><td><p>Concealment. Mythos Preview&#x27;s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.</p></td><td><p>Chain-of-thought monitoring used as a primary control on agent behavior.</p></td><td><p>Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.</p></td></tr><tr><td><p>Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.</p></td><td><p>Fleets of identical agents on one model making correlated calls at the same moment.</p></td><td><p>Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.</p></td></tr><tr><td><p>Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.</p></td><td><p>Pricing, bidding, or procurement agents acting in the same market.</p></td><td><p>Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.</p></td></tr></tbody></table><p>Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there&#x27;s a level of tolerance that&#x27;s being given right now in AI that is unlike anything else in society,” McGladrey said. </p><p>McGladrey&#x27;s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.</p>]]></description>
            <author>louiswcolumbus@gmail.com (Louis Columbus)</author>
            <category>Security</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/wysFOf7ze5H4BCTXQfy63/7970ccebe5cf3361afbd2f35aefa6065/turfwar_hero.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut]]></title>
            <link>https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut</link>
            <guid isPermaLink="false">017W5EU8WmUPzADxzKpTLI</guid>
            <pubDate>Thu, 13 Aug 2026 18:02:54 GMT</pubDate>
            <description><![CDATA[<p>Google is <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/">rolling out Gemini 3.7 Flash</a>, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half.</p><p>The release arrives just <a href="https://venturebeat.com/technology/googles-gemini-3-6-flash-model-cuts-ai-agent-token-costs-by-up-to-65-on-long-horizon-engineering-tasks-and-3-5-pro-is-on-the-way">three weeks after the release of Gemini 3.6 Flash</a>, an unusually short turnaround that Google attributes to developer feedback and algorithmic improvements. </p><p>For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, <a href="https://ai.google.dev/gemini-api/docs/pricing">Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens</a>.</p><p>Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google&#x27;s claimed reductions in retries and manual oversight translate into lower total operating costs.</p><p>The launch also underscores Google&#x27;s rapid iteration on its Flash line while its next flagship Pro model remains absent. Google did not provide a release date for Gemini 3.5 Pro with Thursday&#x27;s announcement, Reuters reported, despite the model having previously been described as undergoing partner testing. Axios similarly noted that 3.7 Flash arrives before the anticipated Pro release.</p><h2><b>A three-week upgrade focused on getting work done</b></h2><p>Google describes Gemini 3.7 Flash as its &quot;most intelligent workhorse model yet for coding and agents.&quot; The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity.</p><p>Those improvements matter beyond benchmark scores. In an enterprise coding agent, a model that makes fewer unnecessary changes, recovers from errors and executes multi-step plans more reliably can reduce the number of human interventions needed to complete a task. The same principle applies to business agents operating across documents and applications, where an incorrect tool call or poorly interpreted instruction can derail an otherwise useful workflow.</p><p>Google says 3.7 Flash &quot;thinks more diligently,&quot; applying more effort to multi-step planning and tool calls. Its stated goal is more disciplined execution with fewer retries and less manual supervision.</p><p>That represents an interesting evolution from Gemini 3.6 Flash. Google&#x27;s developer documentation described 3.6 as reducing reasoning steps, conversational turns and tool calls compared with earlier models while attempting to limit execution-loop spiraling. With 3.7, the emphasis shifts toward putting sufficient effort into planning while improving the quality of execution — potentially a more useful optimization than simply minimizing the number of steps an agent takes.</p><p>Google DeepMind said in a post accompanying the release that 3.7 Flash shows gains in debugging and issue resolution, generates more functional web layouts and applications with fewer prompts, and improves reasoning and accuracy on real-world business workflows.</p><h2><b>Coding gains are substantial, but not universal</b></h2><p>Google&#x27;s benchmarks show a large generational improvement in several software engineering tests.</p><p>On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scores 43.6%, up from 34.4% for Gemini 3.6 Flash. That also narrowly exceeds the 42.7% Google reports for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra.</p><p>On DeepSWE v1.1, a long-horizon software engineering evaluation, 3.7 Flash reaches 65.3%, compared with 49.0% for its predecessor. GPT-5.6 Terra remains ahead at 69.6% in Google&#x27;s table.</p><p>Web development shows another notable gain. Gemini 3.7 Flash receives an Elo score of 1588 on Code Arena, versus 1538 for 3.6 Flash, 1541 for Claude Sonnet 5 and 1523 for GPT-5.6 Terra. Google says the new model can produce more functional layouts and feature-complete applications in fewer prompts while more closely following reference screenshots, images and design systems.</p><p>The broader benchmark table is more mixed, which is important for enterprises evaluating the model against particular workloads rather than looking for a single &quot;best&quot; model.</p><p>Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1, compared with 87.4% for GPT-5.6 Terra. Terra also leads Google&#x27;s comparisons on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads the Agent&#x27;s Last Exam multimodal desktop and operating-system tasks with a 33.3% pass rate, versus 26.3% for Gemini 3.7 Flash.</p><p>In other words, Google&#x27;s own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.</p><h2><b>Enterprise workflows may be the more important test</b></h2><p>The gains extend beyond software development.</p><p>On AutomationBench, which Google describes as measuring enterprise workflow automation, Gemini 3.7 Flash scores 30.4%, up sharply from 17.0% for 3.6 Flash. Google&#x27;s table lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%.</p><p>The model also reaches 34.0% on GDP.PDF, an evaluation of complex PDF comprehension, compared with 22.0% for 3.6 Flash, 28.0% for Claude Sonnet 5 and 24.7% for GPT-5.6 Terra.</p><p>That combination is relevant for enterprise agents because many practical deployments require more than generating text or code. An agent may need to interpret a long report, identify relevant information, decide which tool to invoke, update another system and produce a document for a human reviewer. Reliability across that chain can matter more than performance on an isolated reasoning benchmark.</p><p>Google is putting that thesis into practice with Gemini Spark. Google AI Pro and Ultra subscribers can use 3.7 Flash in Spark, the company&#x27;s personal AI agent. Google says the upgrade improves Spark&#x27;s knowledge work and tool use across Google Workspace applications, including workflows that consolidate files, draft emails and update status documents.</p><p>For enterprises, 3.7 Flash is also available through the Gemini Enterprise Agent Platform and Gemini Enterprise app.</p><h2><b>Price becomes part of the model competition</b></h2><p>Gemini 3.7 Flash&#x27;s introductory pricing is a notable bid to embed the model into enterprise workflows.</p><p>Until Dec. 31, developers pay $0.75 per million input tokens and $3.75 per million output tokens. Context caching costs $0.075 per million tokens during the introductory period. Google says standard prices will double on Jan. 1, 2027, to $1.50 for input and $7.50 for output, with context caching rising to $0.15.</p><p>For comparison, Gemini 3.6 Flash&#x27;s standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Google&#x27;s benchmark table lists Claude Sonnet 5 at $2 and $10, respectively, while GPT-5.6 Terra is listed at $2 and $12.</p><p>The economics become more pronounced for autonomous agents because a single user request can produce a long sequence of model calls, reasoning tokens and tool interactions. A model that costs less per token but requires substantially more retries may not ultimately be cheaper. </p><p>Conversely, Google&#x27;s combination of lower introductory token pricing and claimed improvements in first-pass accuracy could materially change the cost of running high-volume coding or document-processing agents if those gains carry over to production.</p><p>That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task.</p><h2><b>Google’s AI shake-up raises the stakes for Gemini</b></h2><p>Gemini 3.7 Flash arrives amid a broader debate over whether Google is losing ground at the AI frontier. The company has not released Gemini 3.5 Pro, despite saying in May that the flagship model would arrive the following month. </p><p>By July, Google said it remained in partner testing and would become broadly available when ready; Thursday’s announcement offered no further timetable. Google’s latest released general-purpose Pro model therefore remains <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/">Gemini 3.1 Pro</a>, introduced in February. </p><p><a href="https://www.reuters.com/business/google-gemini-launch-delayed-tech-falls-short-internal-goals-bloomberg-news-2026-07-16/">Reuters reported in July that Gemini 3.5 Pro missed its original target</a> after falling short of internal goals, particularly in coding, even as Google began training what it calls its most ambitious model yet, Gemini 4.</p><p>The delay coincides with a <a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/">major overhaul of Google’s AI leadership announced last week.</a> </p><p>Google DeepMind co-founder and Nobel Prize Winner Demis Hassabis has relinquished day-to-day control of the company&#x27;s famed DeepMind AI division to become its chair and, simultaneously, to take on the role of Alphabet’s chief scientist.</p><p>Meanwhile, former DeepMind CTO Koray Kavukcuoglu now runs the unit as a senior vice president reporting directly to CEO Sundar Pichai. </p><p>Kavukcuoglu controls Gemini model development, frontier research, the Gemini app and developer teams—effectively consolidating the full Gemini chain under a more product-focused operator. </p><p>Chief scientist Jeff Dean, Gemini co-lead Oriol Vinyals, Quoc Le and Sanjay Ghemawat left to establish the <a href="https://www.wired.com/story/jeff-dean-google-discovery-loop-startup/">research startup Discovery Loop. </a></p><p>Those exits followed Gemini co-lead Noam Shazeer’s move to OpenAI and Nobel Prize-winning AlphaFold scientist John Jumper’s departure for Anthropic. <a href="https://www.reuters.com/world/inside-google-executive-moves-that-led-its-big-ai-reshuffle-2026-08-12/">Reuters reported</a> that internal disagreements, constrained compute allocation and Google’s bureaucracy contributed to slower releases and weaknesses in coding.</p><p>Outside interpretations range from organizational repair to a more fundamental retreat. </p><p><a href="https://newsletter.semianalysis.com/p/gemini-is-cooked-but-gcp-is-cooking">SemiAnalysis has argued</a> that Google is increasingly prioritizing the highly profitable business of supplying cloud infrastructure to AI companies—including Gemini competitors—over keeping its own models at the absolute frontier. That analysis also claimed Google had effectively canceled 3.5 Pro, although Google has not confirmed that and continues to describe the model as delayed. </p><p><a href="https://www.theverge.com/podcast/979370/google-deepmind-ai-race-lose-jeff-dean-demis-hassabis">The Verge offered a more measured assessment</a>: the departures and model delays are serious, but Google retains enormous advantages through Search, Workspace, Android, Cloud, custom AI chips and consumer distribution. Google says the Gemini app has surpassed 950 million monthly users, giving it a reach that does not depend entirely on owning the highest-scoring model.</p><p>Current benchmarks similarly depict a company behind the overall leaders but still firmly competitive. Artificial Analysis places Claude Opus 5 at 63 on its overall model <a href="https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index">Intelligence Index</a>, while Google reports a score of 56 for Gemini 3.7 Flash—an improvement from 52 for 3.6 Flash but not a return to the top. </p><p>Arena’s early human-preference results are more favorable, provisionally ranking 3.7 Flash ninth overall and eighth for web development. </p><p>The resulting picture is not that Google has abandoned advanced AI, but that it has become stronger at rapidly shipping efficient Flash models while struggling to deliver the premium flagship required to reclaim broad leadership. Gemini 4 will now serve as the clearest test of whether the leadership reorganization fixes that execution gap.</p><h2><b>Available now across Google&#x27;s developer stack</b></h2><p>Developers can access Gemini 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, as well as Google&#x27;s Antigravity environment. Enterprises can deploy it through Gemini Enterprise Agent Platform and Gemini Enterprise, while consumers with Google AI Pro or Ultra subscriptions can access the model through Spark in supported countries.</p><p>Google is also shipping updated safeguards covering chemical, biological, radiological and nuclear risks and cyber-offense misuse, according to the company.</p><p>The unusually fast jump from Gemini 3.6 Flash to 3.7 Flash points toward a model development cycle in which algorithmic improvements can reach production products without waiting for a new flagship generation. Ars Technica also highlighted the three-week interval between the two releases, while Google says the techniques behind the update will inform future models.</p><p>For developers, that faster cadence creates its own operational question. Models can improve quickly, but production teams still have to benchmark new releases against their own repositories, prompts, tool schemas and failure modes before changing a deployment.</p><p>Gemini 3.7 Flash gives those teams a particularly strong incentive to run that evaluation. Google&#x27;s own numbers show major improvements in production coding, web development, document comprehension and workflow automation without claiming leadership everywhere. At its introductory price, Google is effectively betting that developers will value a model that is competitive enough with more expensive systems while being cheap enough to run repeatedly inside agents.</p><p>Whether that advantage survives the return to full pricing in January will depend less on leaderboard positions than on how reliably 3.7 Flash completes real work.</p>]]></description>
            <author>carl.franzen@venturebeat.com (Carl Franzen)</author>
            <category>Technology</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/2hiXA5MTlT6U93qXfs9J8Q/e9a84e4962a92f6170fb5761a78b3d5f/Gemini_Generated_Image_9mr5qd9mr5qd9mr5.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices]]></title>
            <link>https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices</link>
            <guid isPermaLink="false">Mg5d6r5JkyqbXcJfDURku</guid>
            <pubDate>Thu, 13 Aug 2026 16:47:55 GMT</pubDate>
            <description><![CDATA[<p>DeepSeek is expanding beyond the model layer and deeper into the software developers use to put AI agents to work.</p><p>The Chinese AI lab on Thursday <a href="https://x.com/deepseek_ai/status/2087864585504305397">launched the official version of DeepSeek-V4-Pro</a>, an updated flagship model focused heavily on agentic workloads, alongside <a href="https://github.com/deepseek-ai/deepseek-harness">DeepSeek Harness v0.1</a>, a new open-source agent harness that gives developers an alternative to integrated coding-agent environments such as Anthropic’s Claude Code.</p><p>Together, the releases amount to a broader developer push from DeepSeek. V4-Pro is now available across DeepSeek’s web interface, mobile app and API, with native support for the OpenAI Responses API and integration with Codex. </p><p>DeepSeek Harness, meanwhile, is entering developer preview under the MIT license and the c<a href="https://github.com/deepseek-ai/deepseek-harness">ode is available now for download and use on GitHub</a>. It&#x27;s built around an unusually modular premise: practically every part of the agent runtime can be swapped out as a plugin.</p><p>But developers accessing V4 through DeepSeek’s API will soon pay considerably more for it. DeepSeek is simultaneously <a href="https://venturebeat.com/infrastructure/how-deepseeks-radical-architecture-is-shattering-silicon-valleys-token-moat">abandoning its existing flat API pricing </a>in favor of peak and off-peak rates beginning at 16:00 UTC on Sunday, Aug. 16 (2 am ET). </p><p>Even the discounted off-peak cache-miss and output prices will be substantially higher than the prices available today.</p><p>The combination is significant because DeepSeek is no longer competing solely over model intelligence and token prices. With Harness, it is moving into the layer that determines how models use tools, manipulate files, maintain sessions and execute long-running agent workflows — territory where Anthropic’s Claude Code and other coding agents have become increasingly important developer products.</p><h2><b>DeepSeek builds its own agent harness</b></h2><p>DeepSeek describes Harness, or <code>dsh</code>, as an open-source agent harness built on Cordis, a framework designed around composable plugins.</p><p>Its guiding principle is simple: <b>“Everything is a plugin.”</b></p><p>That extends to models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration and user interfaces, according to DeepSeek. Rather than making those components fixed pieces of a single coding agent, Harness is designed to let developers mix, replace and extend them.</p><p>The project is available under the MIT license and can currently be launched from npm with <code>npx @deepseek-ai/dsh web</code>. DeepSeek also provides instructions for building it directly from source. The repository describes the software explicitly as a developer preview and warns that “THERE WILL BE COMPATIBILITY-BREAKING CHANGES.”</p><p>That caveat matters for enterprise developers. Harness is not yet being presented as a stable drop-in production platform. But its architecture points toward a potentially important strategy: DeepSeek can now offer developers not only models but an open framework for assembling the systems that surround them.</p><p>That makes Anthropic&#x27;s Claude Code and OpenAI&#x27;s Codex useful competitive references, although the products should not be treated as functionally identical. </p><p>DeepSeek Harness is an open-source, model-agnostic alternative to the agent infrastructure underlying Claude Code and Codex—not yet a full replacement for either product’s broader developer experience.</p><p>It can already inspect repositories, edit files, execute shell commands, search files and the web, maintain plans, invoke skills, delegate work to subagents and enforce approval policies. Those are the essential capabilities that make Claude Code and Codex agentic coding tools rather than autocomplete systems. </p><p>DeepSeek explicitly describes Standard mode as a full coding agent with file editing, shell access, search, planning, subagents and workflows. Its local web interface lets users select a workspace and approve sensitive operations. </p><p>But Claude Code and Codex now extend well beyond that agent loop. Here&#x27;s a quick comparison:</p><table><tbody><tr><td><p><b>Dimension</b></p></td><td><p><b>DeepSeek Harness</b></p></td><td><p><b>Claude Code</b></p></td><td><p><b>OpenAI Codex</b></p></td></tr><tr><td><p>Read, edit and test a repository</p></td><td><p>Yes</p></td><td><p>Yes</p></td><td><p>Yes</p></td></tr><tr><td><p>Shell and development tools</p></td><td><p>Yes</p></td><td><p>Yes</p></td><td><p>Yes</p></td></tr><tr><td><p>Planning and subagents</p></td><td><p>Yes</p></td><td><p>Yes</p></td><td><p>Yes</p></td></tr><tr><td><p>Permission controls and sandboxing</p></td><td><p>Yes, configurable through plugins</p></td><td><p>Yes, mature built-in permission and sandbox system</p></td><td><p>Yes, granular sandbox and approval controls</p></td></tr><tr><td><p>Primary interfaces</p></td><td><p>Local web UI; headless command; Python SDK</p></td><td><p>Terminal, VS Code, JetBrains, desktop, browser, mobile and Slack</p></td><td><p>CLI, IDE extension, desktop app, web/cloud and integrations</p></td></tr><tr><td><p>Hosted background agents</p></td><td><p>Not documented as a DeepSeek-managed service</p></td><td><p>Yes</p></td><td><p>Yes</p></td></tr><tr><td><p>GitHub-native PR workflow</p></td><td><p>Not documented as a finished integration</p></td><td><p>GitHub Actions, automatic reviews, issue-to-PR workflows</p></td><td><p>Cloud tasks, automatic reviews, PR fixes and GitHub Action</p></td></tr><tr><td><p>Model choice</p></td><td><p>DeepSeek, Anthropic, OpenAI and custom compatible endpoints</p></td><td><p>Primarily Claude, including Bedrock, Google Cloud and Microsoft hosting</p></td><td><p>Primarily OpenAI models, with configurable providers in the open-source CLI</p></td></tr><tr><td><p>Extensibility</p></td><td><p>Exceptional: virtually every component is replaceable</p></td><td><p>Strong: skills, hooks, MCP, plugins and agent teams</p></td><td><p>Strong: skills, MCP, custom agents, SDK and app server</p></td></tr><tr><td><p>Product maturity</p></td><td><p>Developer preview; breaking changes expected</p></td><td><p>Established commercial product</p></td><td><p>Established commercial product plus open-source CLI</p></td></tr><tr><td><p>License</p></td><td><p>MIT</p></td><td><p>Commercial product with extensibility interfaces</p></td><td><p>Codex CLI is open source; cloud and app services are managed products</p></td></tr></tbody></table><p>DeepSeek Harness instead emphasizes modularity and replacement: the model itself is another plugin rather than necessarily the center of a vertically integrated stack.</p><p>DeepSeek’s repository was already attracting significant developer attention on launch day, showing roughly 27,500 GitHub stars and 2,000 forks as of Aug. 13, although those rapidly changing figures are best viewed as a snapshot rather than an adoption metric.</p><h2><b>V4-Pro gets an agent-focused upgrade</b></h2><p>Harness arrives alongside the general-availability release of DeepSeek-V4-Pro-0813.</p><p><a href="https://venturebeat.com/technology/deepseek-v4-arrives-with-near-state-of-the-art-intelligence-at-1-6th-the-cost-of-opus-4-7-gpt-5-5">DeepSeek originally introduced the V4 family in preview in April.</a> The lineup consists of the 1.6-trillion-parameter V4-Pro, with 49 billion parameters activated per token, and the smaller 284-billion-parameter V4-Flash, with 13 billion activated. Both support context windows of up to one million tokens.</p><p>The company’s Aug. 13 release therefore is not the first appearance of V4-Pro. It is the transition from the earlier preview into an updated official version, with DeepSeek emphasizing agent performance.</p><p>“The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration,” DeepSeek says on its<a href="https://platform.deepseek.com/transactions"> API website</a>. “It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.”</p><p>DeepSeek’s changelog similarly says the general-availability model has “significantly enhanced Agent capabilities,” particularly in production environments. Developers using the API do not have to change model identifiers: <code>deepseek-v4-pro</code> now resolves to the latest V4-Pro version.</p><p>The company has also added native<a href="https://venturebeat.com/orchestration/openai-upgrades-its-responses-api-to-support-agent-skills-and-a-complete"> OpenAI Responses API suppor</a>t, lowering the amount of integration work required for applications already built around that interface. </p><p>DeepSeek says V4-Pro is optimized for OpenAI&#x27;s own open source harness, Codex, with one-click setup. Its current API documentation lists Responses API, tool calling, JSON output and an Anthropic-format API among the supported interfaces for both V4-Pro and V4-Flash.</p><p>For developers using DeepSeek directly rather than through an API, V4-Pro is now accessible through “Expert Mode” on the company’s app and website.</p><h2><b>Reasoning effort becomes another deployment knob</b></h2><p>DeepSeek is also making reasoning effort an explicit control across V4-Pro and V4-Flash.</p><p>The V4 model documentation describes three levels: Non-think, designed for fast routine tasks; Think High, intended for more complex problem-solving and planning; and Think Max, which allocates substantially more reasoning to difficult problems.</p><p>That distinction can be operationally important for agent systems because maximum reasoning on every step can consume unnecessary time and tokens. A coding agent might use relatively little reasoning to inspect a file or execute a routine tool call, then increase effort when diagnosing a difficult bug or planning a multi-stage code change.</p><p>DeepSeek’s latest benchmark table suggests the 0813 model improves substantially on agent-oriented tests, although the figures are company-reported and some results depend on the harness configuration.</p><p>DeepSeek reports V4-Pro-0813 scores of <b>87.9 on Terminal Bench 2.1, 74.1 on Toolathlon-Verified, 71.1 on DSBench-FullStack and 67.2 on DSBench-Hard</b>. It does not lead every comparison in DeepSeek’s own table: Fable 5, for example, scores 77.9 on Toolathlon-Verified and 77.2 on DSBench-FullStack.</p><p>There is an especially important qualification buried beneath the benchmark table. For public Code Agent tasks, DeepSeek says V4-Pro-0813 was tested using its upcoming DeepSeek Harness in “minimal mode.”</p><p>In other words, some of the agent results arriving alongside Harness are not purely model benchmarks. They measure the model operating inside an agent execution environment — precisely the software layer DeepSeek is now releasing to developers.</p><h2><b>A sharp reversal in DeepSeek’s API price trajectory</b></h2><p>The bigger immediate change for teams already running DeepSeek in production may be pricing.</p><p>DeepSeek’s current API documentation lists V4-Flash at $0.14 per million cache-miss input tokens and $0.28 per million output tokens, while V4-Pro costs $0.435 for cache-miss input and $0.87 for output. Cache hits are dramatically cheaper at $0.0028 for Flash and $0.003625 for Pro.</p><p>Those prices themselves represented a major reduction from V4’s original April launch economics. When V4 arrived in April, V4-Pro was priced at $1.74 per million cache-miss input tokens and $3.48 per million output tokens. By late May, DeepSeek had made a 75% reduction permanent, intensifying its position as an unusually inexpensive option for high-volume agent workloads. Now the pendulum is moving in the other direction.</p><p>Beginning Aug. 16 at 16:00 UTC, DeepSeek will charge different rates depending on when API calls occur. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC (9:00 PM – 12:00 AM ET and 2:00 AM – 6:00 AM ET, respectively) with all other hours classified as off-peak. Off-peak rates are half the corresponding peak prices.</p><p>For V4-Flash, off-peak cache-miss input rises from $0.14 to $0.22 per million tokens, while output rises from $0.28 to $0.66. During peak hours those rates reach $0.44 input and $1.32 output.</p><p>V4-Pro moves from $0.435 per million cache-miss input tokens and $0.87 output today to $0.66 and $1.98 off-peak, respectively. Peak rates rise to $1.32 input and $3.96 output.</p><p>The increases are even more pronounced for cached input. V4-Pro cache hits rise from $0.003625 per million tokens today to $0.022 off-peak and $0.044 at peak. Flash moves from $0.0028 to $0.007 off-peak and $0.014 peak.</p><table><tbody><tr><td><p><b>Model</b></p></td><td><p><b>Old input (per 1M token)</b></p></td><td><p><b>Old output (per 1M tok)</b></p></td><td><p><b>Old total (1M in/1M out)</b></p></td></tr><tr><td><p>deepseek-v4-flash</p></td><td><p>$0.14</p></td><td><p>$0.28</p></td><td><p>$0.42</p></td></tr><tr><td><p>deepseek-v4-pro</p></td><td><p>$0.435</p></td><td><p>$0.87</p></td><td><p>$1.305</p></td></tr></tbody></table><p>The new prices still position DeepSeek as an affordable alternative via API to Western proprietary labs, but <a href="https://www.reuters.com/world/china/deepseek-releases-official-v4-pro-model-it-steps-up-expansion-2026-08-13/">Reuters reported Thursday</a> that, depending on model, token category and time of use, the changes represent increases ranging from 50% to more than 1,100% over existing rates.</p><table><tbody><tr><td><p><b>Model</b></p></td><td><p><b>Input ($/1M)</b></p></td><td><p><b>Output ($/1M)</b></p></td><td><p><b>Total ($/1M)</b></p></td><td><p><b>Source</b></p></td></tr><tr><td><p>Muse Spark 1.2 Contributor</p></td><td><p>$0.10</p></td><td><p>$0.20</p></td><td><p>$0.30</p></td><td><p><a href="https://dev.meta.ai/docs/pricing-rate-limits">Meta</a></p></td></tr><tr><td><p>MiMo-V2.5 Flash</p></td><td><p>$0.10</p></td><td><p>$0.30</p></td><td><p>$0.40</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p><b>DeepSeek-V4-Flash — off-peak</b></p></td><td><p><b>$0.22</b></p></td><td><p><b>$0.66</b></p></td><td><p><b>$0.88</b></p></td><td><p><b></b><a href="https://x.com/deepseek_ai/status/2087864589895798968"><b>DeepSeek</b></a><b></b></p></td></tr><tr><td><p>GPT-5.6 Luna</p></td><td><p>$0.20</p></td><td><p>$1.20</p></td><td><p>$1.40</p></td><td><p><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/">OpenAI</a></p></td></tr><tr><td><p>MiniMax-M3</p></td><td><p>$0.30</p></td><td><p>$1.20</p></td><td><p>$1.50</p></td><td><p><a href="https://platform.minimax.io/subscribe/token-plan?tab=api-enterprise">MiniMax</a></p></td></tr><tr><td><p>LongCat-2.0 — limited-time promo</p></td><td><p>$0.30</p></td><td><p>$1.20</p></td><td><p>$1.50</p></td><td><p><a href="https://longcat.chat/platform/docs/APIPayAsYouGo.html">LongCat</a></p></td></tr><tr><td><p><b>DeepSeek-V4-Flash — peak hours</b></p></td><td><p><b>$0.44</b></p></td><td><p><b>$1.32</b></p></td><td><p><b>$1.76</b></p></td><td><p><b></b><a href="https://x.com/deepseek_ai/status/2087864589895798968"><b>DeepSeek</b></a><b></b></p></td></tr><tr><td><p>MiMo-V2.5</p></td><td><p>$0.40</p></td><td><p>$2.00</p></td><td><p>$2.40</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p><b>DeepSeek-V4-Pro — off-peak</b></p></td><td><p><b>$0.66</b></p></td><td><p><b>$1.98</b></p></td><td><p><b>$2.64</b></p></td><td><p><b></b><a href="https://x.com/deepseek_ai/status/2087864589895798968"><b>DeepSeek</b></a><b></b></p></td></tr><tr><td><p>LongCat-2.0 — standard</p></td><td><p>$0.75</p></td><td><p>$2.95</p></td><td><p>$3.70</p></td><td><p><a href="https://longcat.chat/platform/docs/APIPayAsYouGo.html">LongCat</a></p></td></tr><tr><td><p>MiMo-V2.5 Pro (≤256K)</p></td><td><p>$1.00</p></td><td><p>$3.00</p></td><td><p>$4.00</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p><b>DeepSeek-V4-Pro — peak hours</b></p></td><td><p><b>$1.32</b></p></td><td><p><b>$3.96</b></p></td><td><p><b>$5.28</b></p></td><td><p><b></b><a href="https://x.com/deepseek_ai/status/2087864589895798968"><b>DeepSeek</b></a><b></b></p></td></tr><tr><td><p>Muse Spark 1.1 / 1.2</p></td><td><p>$1.25</p></td><td><p>$4.25</p></td><td><p>$5.50</p></td><td><p><a href="https://dev.meta.ai/docs/pricing-rate-limits">Meta</a></p></td></tr><tr><td><p>GLM-5.2</p></td><td><p>$1.40</p></td><td><p>$4.40</p></td><td><p>$5.80</p></td><td><p><a href="https://docs.z.ai/guides/overview/pricing">Z.ai</a></p></td></tr><tr><td><p>Grok 4.6 — &lt;200K prompt tokens</p></td><td><p>$2.00</p></td><td><p>$6.00</p></td><td><p>$8.00</p></td><td><p><a href="https://docs.x.ai/developers/models/grok-4.6">xAI</a></p></td></tr><tr><td><p>MiMo-V2.5 Pro (&gt;256K)</p></td><td><p>$2.00</p></td><td><p>$6.00</p></td><td><p>$8.00</p></td><td><p><a href="https://platform.xiaomimimo.com/docs/en-US/pricing">Xiaomi</a></p></td></tr><tr><td><p>Qwen3.8-Max</p></td><td><p>$2.00</p></td><td><p>$6.00</p></td><td><p>$8.00</p></td><td><p><a href="https://www.qwencloud.com/models/qwen3.8-max">QwenCloud</a></p></td></tr><tr><td><p>Gemini 3.6 Flash</p></td><td><p>$1.50</p></td><td><p>$7.50</p></td><td><p>$9.00</p></td><td><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a></p></td></tr><tr><td><p>GPT-5.6 Terra</p></td><td><p>$2.00</p></td><td><p>$12.00</p></td><td><p>$14.00</p></td><td><p><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/">OpenAI</a></p></td></tr><tr><td><p>Grok 4.6 — ≥200K prompt tokens</p></td><td><p>$4.00</p></td><td><p>$12.00</p></td><td><p>$16.00</p></td><td><p><a href="https://docs.x.ai/developers/models/grok-4.6">xAI</a></p></td></tr><tr><td><p>GPT-5.4</p></td><td><p>$2.50</p></td><td><p>$15.00</p></td><td><p>$17.50</p></td><td><p><a href="https://openai.com/api/pricing/">OpenAI</a></p></td></tr><tr><td><p>Kimi K3</p></td><td><p>$3.00</p></td><td><p>$15.00</p></td><td><p>$18.00</p></td><td><p><a href="https://platform.kimi.ai/docs/pricing/chat-k3">Moonshot AI</a></p></td></tr><tr><td><p>Claude Opus 5</p></td><td><p>$5.00</p></td><td><p>$25.00</p></td><td><p>$30.00</p></td><td><p><a href="https://platform.claude.com/docs/en/about-claude/pricing">Anthropic</a></p></td></tr><tr><td><p>Sakana Fugu Ultra (≤272K)</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://console.sakana.ai/pricing#subscription-plan">Sakana AI</a></p></td></tr><tr><td><p>GPT-5.6 Sol — Standard mode</p></td><td><p>$5.00</p></td><td><p>$30.00</p></td><td><p>$35.00</p></td><td><p><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/">OpenAI</a></p></td></tr><tr><td><p>Claude Fable 5 / Claude Mythos 5</p></td><td><p>$10.00</p></td><td><p>$50.00</p></td><td><p>$60.00</p></td><td><p><a href="https://platform.claude.com/docs/en/about-claude/models/overview">Anthropic</a></p></td></tr><tr><td><p>GPT-5.6 Sol — Fast mode</p></td><td><p>$10.00</p></td><td><p>$60.00</p></td><td><p>$70.00</p></td><td><p><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/">OpenAI</a></p></td></tr></tbody></table><p>That makes the “50% lower” off-peak framing potentially misleading without context. Off-peak is 50% cheaper than DeepSeek’s <i>new</i> peak rate; it is not a 50% discount from the API prices developers are paying today.</p><p>For a simple workload consisting of one million cache-miss input tokens plus one million output tokens, V4-Pro currently costs $1.305. The same token mix will cost $2.64 off-peak, roughly twice as much, or $5.28 during peak hours, more than four times the current price.</p><p>V4-Flash moves from $0.42 under the same simple calculation to $0.88 off-peak and $1.76 peak.</p><p>Actual application costs will vary considerably depending on the ratio of cached input, uncached input and generated output, making those combined figures illustrative rather than universal total-cost estimates.</p><h2><b>DeepSeek is moving up the agent stack</b></h2><p>The timing makes the strategic direction difficult to miss.</p><p>When DeepSeek released the V4 preview on April 24, the major story was how much frontier-class capability the company could deliver with an unusually efficient architecture. </p><p>V4-Pro uses a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention; at a one-million-token context, DeepSeek says it requires only 27% of the single-token inference FLOPs and 10% of the KV cache required by V3.2.</p><p>By late May, the discussion had shifted toward what those efficiencies meant economically for high-volume agents, whose repeated context reads can make caching a major component of inference costs. DeepSeek’s steep V4 price cuts amplified that advantage.</p><p>The Aug. 13 releases move the competition another layer upward.</p><p>DeepSeek now has an updated V4-Pro tuned around agent workloads, standardized interfaces designed to make it easier to connect with existing developer tooling, configurable reasoning effort, and an MIT-licensed harness for controlling the models, tools, sandboxes, filesystems and orchestration surrounding an agent.</p><p>At the same time, DeepSeek is demonstrating that developers cannot assume its aggressively low API rates are permanent. For organizations considering the platform, workload scheduling, caching behavior and the option to run open weights on their own infrastructure now become more important parts of the total-cost calculation.</p><p>That leaves DeepSeek pursuing two potentially conflicting advantages at once: making its agent stack more accessible and open while making its own hosted API considerably more expensive.</p><p>For enterprise developers, Harness may ultimately be the more consequential part of Thursday’s announcement. Models can increasingly be swapped behind standardized interfaces. The harness that controls how an agent reasons, invokes tools, edits software and persists across a workflow can be much harder to replace.</p><p>DeepSeek is now competing for that layer, too.</p>]]></description>
            <author>carl.franzen@venturebeat.com (Carl Franzen)</author>
            <category>Technology</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/6hFdhAgllY0UvaVpl6BmLa/d6e4fff99189904b4a8a9f037e1fbb83/43530361-2311-45E8-BCF9-8DF972A298D8.png?w=300&amp;q=30" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Why Capital One built its multi-agent AI platform around open-weight models]]></title>
            <link>https://venturebeat.com/orchestration/why-capital-one-built-its-multi-agent-ai-platform-around-open-weight-models</link>
            <guid isPermaLink="false">4n9mkJEYB9mXPbd19CqSRA</guid>
            <pubDate>Thu, 13 Aug 2026 14:00:00 GMT</pubDate>
            <description><![CDATA[<p><i>Presented by Capital One </i></p><hr/><p>At <a href="https://venturebeat.com/vbtransform2026">VB Transform 2026</a>, <!-- -->Kel Vanee, MVP of machine learning engineering at Capital One, spoke with Sam Witteveen, Senior Technology Contributor at VentureBeat, about how the bank built a scalable multi-agent AI architecture around deeply customized open-weight models rather than relying on an off-the-shelf foundation model. </p><p>&quot;At Capital One, we&#x27;re not just using AI, we&#x27;re building AI,&quot; Vanee said.</p><div></div><p>The groundwork was laid years ago with Capital One&#x27;s early investments in data transformation and cloud adoption, which Vanee said were foundational to moving quickly when the current wave of AI arrived. That technical foundation enabled the company to make several deliberate architectural decisions, including building a centralized, enterprise-wide AI platform with built-in governance, deeply customizing open models with proprietary data, and constructing its own multi-agent orchestration harness.</p><h2>Customizing open-weight models with proprietary data </h2><p>Rather than relying solely on off-the-shelf frontier models, Capital One fine-tunes open-weight models using its rich, proprietary data.</p><p>&quot;We view our data as a huge advantage and something that nobody else has, something that the general frontier models cannot provide. So we are taking that data and deeply customizing these models,&quot; Vanee explained. He added that real-time data is absolutely critical to bring in fresh context during live customer or associate interactions.</p><p>Vanee also revealed an unexpected benefit of this approach: extensibility across the enterprise. </p><p>“As we customize those open-source models for one use case, we actually see benefits across our whole portfolio,&quot; he noted. &quot;We are training that model to be an expert at Capital One use cases, policy, and nomenclature. As we do that training, we see a general lift.&quot;</p><h2>Inside Capital One&#x27;s multi-agentic AI workflow</h2><p>As an example of the approach, Vanee pointed to a customer-service workflow for bank fraud that handles millions of calls a year, where interactions range from roughly four minutes to as long as sixty minutes, and where an initial attempt at engaging a single large language model proved insufficient. With Capital One&#x27;s multi-agentic workflow (MACAW), interactions are routed through specialized agents with governance and guardrails built in.</p><p>&quot;The MACAW workflow is made up of a number of different agents,&quot; he said. &quot;The first one is an understanding agent. Its purpose is to look at what the customer is saying and try to understand what their intention is.”</p><p>From there, a reasoning agent is given several specific instructions to generate a summary; a validation agent fact-checks the summary to ensure it is accurate; and an explaining agent turns the summary into a formatted document with all necessary details that is then shared with agents.</p><p>For the consumer banking use case, this workflow helps several hundred customer-service agents who specialize in complex fraud calls. The post-call summaries it generates help document long, back-and-forth interactions that agents previously had to reconstruct by hand.</p><p>Capital One’s multi-agentic architecture also underpins Chat Concierge, a customer-facing auto-shopping assistant, which further leverages a version of Meta&#x27;s open-weight Llama model that has been customized with Capital One&#x27;s proprietary data. It uses the same division of labor, with one agent conversing with the customer, one building an action plan from business rules, one evaluating accuracy, and one explaining and validating the result. </p><h2>Optimizing latency and cost with an agentic research system</h2><p>Beyond customer-facing solutions, Capital One is also leveraging agentic AI to automate rote tasks for its employees and help them focus on high-leverage aspects of their work. In one example, the company built an autonomous agentic optimization solution to tune backend hosting infrastructure.</p><p>Vanee explained that in the world of LLMs, where new optimizations are delivered every day, they aren&#x27;t all complementary. Combining two good optimizations can sometimes cause a performance regression.</p><p>&quot;This agentic system will run through a search space that is designed by the researcher, handle all the mechanics of setting up that experiment and running the experiment, and then put a whole summarization of the results in front of the researcher,&quot; Vanee said. </p><p>Vanee added that the system allows researchers to “find the series of optimizations and configurations that&#x27;s really going to give [them] the best latency possible.”</p><h2>What&#x27;s next: model routing and proactive, event-driven AI</h2><p>Looking ahead, one big trend Vanee sees is routing abstraction layers that a platform seeks to validate over multiple models, both for cost and accuracy.</p><p>&quot;We actually think that you can get better accuracy than any individual model simply by routing across a broader set of available models, because different models are going to excel in different areas,&quot; he said.</p><p>His second prediction was a shift toward systems that act without waiting to be asked, while also emphasizing that deploying such proactive agents would demand rigorous testing and monitoring.</p><p>&quot;The thing I think is going to become bigger in the future is more proactive and event-driven AI,&quot; Vanee said. Rather than waiting for a human prompt, AI would step in as soon as it detects conditions that warrant action.</p><p>&quot;This is going to enable more monitoring and larger-scale monitoring, and it&#x27;ll empower us as we fight fraud and address these opportunities,&quot; Vanee said. &quot;So proactive AI is going to be a really important trend.&quot;</p><h2>Driving continuous AI innovation in financial services</h2><p>Capital One’s approach underscores a broader truth for enterprise technology leaders: driving measurable value with AI requires moving beyond off-the-shelf software toward deeply customized, highly governed architectures. By combining fine-tuned open-weight models, a multi-agent orchestration harness, and proprietary data assets, the bank has established a repeatable blueprint for deploying scalable AI in financial services.</p><p>&quot;All of those ingredients were absolutely critical to differentiating in this space and hitting the quality bars as well as the cost and latency thresholds we set for ourselves,” Vanee said.</p><p>As the company expands these capabilities across new use cases, its enterprise platform approach helps to ensure that technical breakthroughs translate into safer, faster, and more personalized experiences for its millions of customers.</p><hr/><p><i>Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact </i><a href="mailto:sales@venturebeat.com"><i><u>sales@venturebeat.com</u></i></a><i>.</i>
</p>]]></description>
            <category>Orchestration</category>
            <enclosure url="https://images.ctfassets.net/jdtwqhzvc2n1/7gbhCwja5N985q8VOA85UE/c8f8683d5ba63276b672c9312ece355e/2026-VB-Transform-Hotel-Nia-0089-X5.jpg?w=300&amp;q=30" length="0" type="image/jpg"/>
        </item>
    </channel>
</rss>