<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" >
  <generator uri="https://gohugo.io/" version="0.166.0">Hugo</generator>
  <link href="https://lostechies.com/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://lostechies.com/" rel="alternate" type="text/html" />
  <updated>2026-09-28T18:50:16Z</updated>
  <id>https://lostechies.com/</id>

  
    <title type="html">Los Techies</title>
  

  

  
  
  
  
    <entry>
      <title type="html">ACP Muse Agent Client Protocol and Muse Code</title>
      <link href="https://lostechies.com/ryansvihla/2026/09/28/acp-muse-agent-client-protocol-and-muse-code/" rel="alternate" type="text/html" title="ACP Muse Agent Client Protocol and Muse Code" />
      <published>2026-09-28T12:38:15&#43;02:00</published>
      <updated>2026-09-28T12:38:15&#43;02:00</updated>
      <id>https://lostechies.com/ryansvihla/2026/09/28/acp-muse-agent-client-protocol-and-muse-code/</id>
      <content type="html" xml:base="https://lostechies.com/ryansvihla/2026/09/28/acp-muse-agent-client-protocol-and-muse-code/">&amp;lt;p&amp;gt;At &amp;lt;a href=&amp;#34;https://slopcop.com/&amp;#34;&amp;gt;work&amp;lt;/a&amp;gt; we tend to try every new model and vendor we can so we know what the state of the art is. Enter &amp;lt;a href=&amp;#34;https://dev.meta.ai/docs/muse-code&amp;#34;&amp;gt;Muse Code&amp;lt;/a&amp;gt;, the models are fast, cheap and pretty capable. For $50 a month you get more credits than I can spend, and for API usage there is a 92-95% discount letting them train on your data. TLDR this is a ton of value you get out of the models from &amp;lt;a href=&amp;#34;https://about.meta.com/&amp;#34;&amp;gt;Meta&amp;lt;/a&amp;gt;. So I went ahead and created &amp;lt;a href=&amp;#34;https://github.com/BrokkAi/muse-acp&amp;#34;&amp;gt;Muse ACP&amp;lt;/a&amp;gt; and my employer was nice enough to let me work on it at work time and now it is a part of several of our projects.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-is-acp&amp;#34;&amp;gt;What is ACP?&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;a href=&amp;#34;https://www.agentclientprotocol.com&amp;#34;&amp;gt;Agent Client Protocol&amp;lt;/a&amp;gt; is an interaction standard to that one can wire up ACP clients (TUI, Text Editor, Batch Processes, Personal Assistants) to an AI agent that has implemented the same protocol. Many Agents provide this out of the box such as &amp;lt;a href=&amp;#34;https://github.com/NousResearch/hermes-agent&amp;#34;&amp;gt;Hermes&amp;lt;/a&amp;gt; via &amp;lt;code&amp;gt;hermes acp&amp;lt;/code&amp;gt;, &amp;lt;a href=&amp;#34;https://opencode.ai&amp;#34;&amp;gt;Opencode&amp;lt;/a&amp;gt; via &amp;lt;code&amp;gt;opencode acp&amp;lt;/code&amp;gt;, you get the pattern, however as of yet Muse Code does not ship one, so I made one.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Because of this one can easily hookup Hermes, Opencode or any of the other dozens of clients in the &amp;lt;a href=&amp;#34;https://agentclientprotocol.com/get-started/registry&amp;#34;&amp;gt;registry&amp;lt;/a&amp;gt; to &amp;lt;a href=&amp;#34;https://www.jetbrains.com/idea/&amp;#34;&amp;gt;IntelliJ&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;#34;https://zed.dev/&amp;#34;&amp;gt;Zed&amp;lt;/a&amp;gt;, or any number of the clients I have written like &amp;lt;a href=&amp;#34;https://github.com/BrokkAi/micro-acp&amp;#34;&amp;gt;Micro ACP&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;#34;https://github.com/foundev/belgr&amp;#34;&amp;gt;Belgr&amp;lt;/a&amp;gt;, or &amp;lt;a href=&amp;#34;https://github.com/BrokkAi/mjolnir&amp;#34;&amp;gt;Mjolnir&amp;lt;/a&amp;gt;.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Now there is a registry, but I have personally found it a real and total pain to get into, so I have chosen to ignore it for Muse ACP until I get enough popularity that someone else adds my ACP server to it (10 stars and counting). It is also easy to add custom clients and I have add an installer for IntelliJ and Zed with Muse ACP (muse-acp install) so that you can try this out with your favorite editor.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;why-muse-acp-and-not-one-of-the-others&amp;#34;&amp;gt;Why Muse ACP and not one of the others&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;When I started the project there were no mature ACP servers using &amp;lt;a href=&amp;#34;https://meta-models.github.io/muse-code-sdk/next/guides/msp-concepts/&amp;#34;&amp;gt;Muse Session Protocol&amp;lt;/a&amp;gt; as the integration point. Also, it is written in Rust instead of Typescript and that is for some people a good reason to prefer the BrokkAi version of the rest.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;conclusion&amp;#34;&amp;gt;Conclusion&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;If you want access to a cheap capable model, go ahead and sign up for a Muse Code subscription plan and give Muse ACP a try, wire it up to your favorite editor or one of mine and get to work.&amp;lt;/p&amp;gt;</content>

      

      

      
        <category term="AI" />
      
        <category term="Muse" />
      
        <category term="Agent Client Protocol" />
      

      
    </entry>
  
    <entry>
      <title type="html">Why not use DeepSeek Flash for everything?</title>
      <link href="https://lostechies.com/ryansvihla/2026/09/22/why-not-use-deepseek-flash-for-everything/" rel="alternate" type="text/html" title="Why not use DeepSeek Flash for everything?" />
      <published>2026-09-22T09:18:47Z</published>
      <updated>2026-09-22T09:18:47Z</updated>
      <id>https://lostechies.com/ryansvihla/2026/09/22/why-not-use-deepseek-flash-for-everything/</id>
      <content type="html" xml:base="https://lostechies.com/ryansvihla/2026/09/22/why-not-use-deepseek-flash-for-everything/">&amp;lt;p&amp;gt;I started using AI heavily when a friend &amp;lt;a href=&amp;#34;https://www.linkedin.com/in/jbellis/&amp;#34;&amp;gt;Jonathan Ellis&amp;lt;/a&amp;gt; launched his own startup &amp;lt;a href=&amp;#34;https://brokk.ai&amp;#34;&amp;gt;Brokk&amp;lt;/a&amp;gt; and I started using his then revolutionary tool &amp;lt;a href=&amp;#34;https://github.com/BrokkAi/brokk-app&amp;#34;&amp;gt;Brokk&amp;lt;/a&amp;gt; in April of 2025, at that stage basically all models were pretty terrible and things like Cursor or Claude Code were basically useless (I like to call them tooling harnesses because they hook up tools like bash, read, write, internet search to models&amp;amp;hellip; and that is most of what they do), so this meant we had to use the latest and greatest to ever get anything done. Brokk despite being basically unusable for a mere mortal (lesson let a designer build the UX), it was light years ahead of Claude Code and Cursor and with the top end model we could get a ton of solid generated code out of things like GPT 4 and Sonnet 3.5. They routinely invented stuff, didn&amp;amp;rsquo;t know how to look into directories etc, what we knew about prompting then was limited and in general getting good work out of them required a lot of cleverness or at least a willingness to accept really terrible code and a drop in productivity (at least in single threaded modes of work).
At least with Brokk with the tools we had built and the efficiencies we engaged in to keep the context small (the total amount of information sent to the models) we could make good use of cheaper less high end models, but generally speaking the lesson was clear, if you could afford it always use the best, or else waste a lot more time and money getting bad results out of the cheaper stuff.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Fast forward to today and even cheap models can score well on difficult coding benchmarks (source: &amp;lt;a href=&amp;#34;https://deepswe.datacurve.ai/&amp;#34;&amp;gt;https://deepswe.datacurve.ai/&amp;lt;/a&amp;gt;) and the results are way more about &amp;amp;lsquo;feel&amp;amp;rsquo; than they are about any sort of qualitative result that one can measure easily.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;a href=&amp;#34;/assets/why-not-use-deepseek-flash-for-everything/deepswe-snapshot.png&amp;#34;&amp;gt;&amp;lt;img src=&amp;#34;/assets/why-not-use-deepseek-flash-for-everything/deepswe-snapshot.png&amp;#34; alt=&amp;#34;DeepSWE v1.1 leaderboard comparing model scores with average cost per task.&amp;#34;&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;DeepSWE v1.1 score versus average cost per task, September 22, 2026 (&amp;lt;a href=&amp;#34;https://deepswe.datacurve.ai/&amp;#34;&amp;gt;source&amp;lt;/a&amp;gt;). Click the chart to view it at full size.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;This is because of a few things at once: the tooling harnesses of today all do the things Brokk was doing in 2025 and often with more quality and efficiency, the &amp;lt;a href=&amp;#34;https://arxiv.org/abs/2501.12948&amp;#34;&amp;gt;advances&amp;lt;/a&amp;gt; &amp;lt;a href=&amp;#34;https://aclanthology.org/2026.findings-acl.1767/&amp;#34;&amp;gt;in&amp;lt;/a&amp;gt; &amp;lt;a href=&amp;#34;https://arxiv.org/abs/2607.22529&amp;#34;&amp;gt;reasoning&amp;lt;/a&amp;gt;, the advances in &amp;lt;a href=&amp;#34;https://arxiv.org/abs/2503.14476&amp;#34;&amp;gt;post training&amp;lt;/a&amp;gt;, and &amp;lt;a href=&amp;#34;https://arxiv.org/abs/2505.09388&amp;#34;&amp;gt;distillation becoming&amp;lt;/a&amp;gt; refined to the point that &amp;lt;a href=&amp;#34;https://www.deepseek.com/en/news/deepseek-v4-1-flash/&amp;#34;&amp;gt;DeepSeek Flash 4.1&amp;lt;/a&amp;gt; which is a model that now beats Sol on our internal benchmarks and is a fraction of the estimated size of Sol.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;a href=&amp;#34;/assets/why-not-use-deepseek-flash-for-everything/deepseek-benchmarks.png&amp;#34;&amp;gt;&amp;lt;img src=&amp;#34;/assets/why-not-use-deepseek-flash-for-everything/deepseek-benchmarks.png&amp;#34; alt=&amp;#34;Benchmark table comparing DeepSeek V4.1-Flash with DeepSeek V4-Pro and V4-Flash, GLM 5.3, Kimi K3, GPT 5.6-Sol, and Claude Opus 5.&amp;#34;&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;DeepSeek V4.1-Flash benchmark comparison (&amp;lt;a href=&amp;#34;https://www.deepseek.com/en/news/deepseek-v4-1-flash/&amp;#34;&amp;gt;source&amp;lt;/a&amp;gt;). Click the table to view it at full size.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;For my own experience I can tell you that:&amp;lt;/p&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;Fable 5.1 (Anthropic) and Astra (OpenAI) are great general purpose models and I have used them a lot, but for many tasks they vastly overengineer a solution.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Sol (OpenAI) is overall a very good value but it also can overengineer on higher thinking levels and is sort of dumb on its default level (it works but you just need to prompt it a lot more). I do not think this is a bad choice honestly as your model for everything and between it and Luna really justifies the cheaper price.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;GLM 5.3 (Z.ai) is very capable and has been my go-to for about a month now.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;GLM 5.3 Flash (Z.ai), DeepSeek Flash v4 are both very cheap to use and for a large variety of straightforward coding or admin tasks are easily good enough.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Recently I have been using DeepSeek Flash V4.1 and while more expensive to use than Flash v4.0 it has largely replaced most of my model usage, and next month I do not plan on renewing my Anthropic, Z.ai or OpenAI subs at the maximum level as a result.&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;h2 id=&amp;#34;conclusion&amp;#34;&amp;gt;Conclusion&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;I do not actually care or think it is super important which model you use anymore, use the model that fits your style and you can afford and after that really compared to any other point prior we get very good results from nearly any provider. Use Muse, use OpenAI, use Z.ai you should be basing these differences now based on who treats you well, respects your data or &amp;lt;a href=&amp;#34;https://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot-upload&amp;#34;&amp;gt;does&amp;lt;/a&amp;gt; &amp;lt;a href=&amp;#34;https://thehackernews.com/2026/07/grok-build-uploads-entire-git.html&amp;#34;&amp;gt;not&amp;lt;/a&amp;gt;. For some gnarly hard problems I will still probably reach a lot for Astra and Fable, but I doubt I will be on a 20x Max sub anymore.&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Ryan Svihla</name>
      </author>
      

      

      
        <category term="AI" />
      
        <category term="Fable" />
      
        <category term="Benchmarking" />
      

      
    </entry>
  
    <entry>
      <title type="html">Using Opus 5 Without Going Insane</title>
      <link href="https://lostechies.com/ryansvihla/2026/09/18/using-opus-5-without-going-insane/" rel="alternate" type="text/html" title="Using Opus 5 Without Going Insane" />
      <published>2026-09-18T14:00:00Z</published>
      <updated>2026-09-18T14:00:00Z</updated>
      <id>https://lostechies.com/ryansvihla/2026/09/18/using-opus-5-without-going-insane/</id>
      <content type="html" xml:base="https://lostechies.com/ryansvihla/2026/09/18/using-opus-5-without-going-insane/">&amp;lt;p&amp;gt;I&amp;amp;rsquo;ve been using Opus 5 as my daily driver for a few weeks now across a bunch of projects and it&amp;amp;rsquo;s a really capable model, but it has two habits that drove me a little nuts before I figured out how to deal with them:&amp;lt;/p&amp;gt;
&amp;lt;ol&amp;gt;
&amp;lt;li&amp;gt;It is verbose in a way that actually hides the answer. You ask a yes or no question and get four paragraphs of context, a code snippet, a caveat, and then the answer somewhere in the middle.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;When it doesn&amp;amp;rsquo;t understand something it does not go look it up, it guesses, and then it builds a lot of confident sounding structure on top of the guess.&amp;lt;/li&amp;gt;
&amp;lt;/ol&amp;gt;
&amp;lt;p&amp;gt;This isn&amp;amp;rsquo;t a post about whether Opus 5 is good or not (it is), it&amp;amp;rsquo;s strictly what I do day to day to keep it useful. None of this is clever, most of it is just repeating yourself.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;start-with-a-claudemd&amp;#34;&amp;gt;Start with a CLAUDE.md&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The first thing I did was put a global CLAUDE.md in place (~/.claude/CLAUDE.md) that tries to head off the verbosity before it starts. This is the current version, warts and all:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-markdown&amp;#34; data-lang=&amp;#34;markdown&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;gh&amp;#34;&amp;gt;# Global instructions
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;gu&amp;#34;&amp;gt;## Be brief
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Answer the question asked, then stop. No preamble, no recap, no summary of what you just did.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Default to a few sentences. Drop headers, bold, and bullet lists unless the content is genuinely a list.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Do not restate the user&amp;amp;#39;s question back to them.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Do not offer next steps unless asked.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; No trailing coda. Do not end with a &amp;amp;#34;one thing I didn&amp;amp;#39;t do&amp;amp;#34; / &amp;amp;#34;note that&amp;amp;#34; / &amp;amp;#34;caveat&amp;amp;#34; paragraph. If a limitation matters, state it inline where it&amp;amp;#39;s relevant, in one clause. If it doesn&amp;amp;#39;t matter, leave it out.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Banned openers for a closing paragraph, no exceptions: &amp;amp;#34;One thing worth deciding&amp;amp;#34;, &amp;amp;#34;One thing worth noting&amp;amp;#34;, &amp;amp;#34;Worth noting&amp;amp;#34;, &amp;amp;#34;Worth flagging&amp;amp;#34;, &amp;amp;#34;One caveat&amp;amp;#34;, &amp;amp;#34;Separately&amp;amp;#34;, &amp;amp;#34;For what it&amp;amp;#39;s worth&amp;amp;#34;. If you catch yourself starting a final paragraph with any of these, delete the paragraph. Do not rephrase it to evade the list -- the ban is on appending a postscript the user did not ask for, not on those specific words.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; This applies to unsolicited future-work suggestions too. If the answer is done, stop at the answer.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;gu&amp;#34;&amp;gt;## Never claim credit
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Never say you already had an idea, thought of it first, suggested it earlier, or were about to do it.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; If the user proposes something, implement it. Do not annotate it with your prior reasoning.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Never describe your own earlier messages as having been right. If a correction is needed, state the fact, not who found it.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Do not narrate your process or reasoning quality, positively or negatively.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;gu&amp;#34;&amp;gt;## Verify before asserting
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Never state how code behaves without reading it. &amp;amp;#34;It&amp;amp;#39;s not in this repo&amp;amp;#34; is not a stopping point — find the source and read it.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;-&amp;lt;/span&amp;gt; Never invent a critique, a self-criticism, or an account of what happened earlier. If you are describing the past, quote it or do not claim it.
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;A few notes on why it looks like that:&amp;lt;/p&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;The &amp;amp;ldquo;banned openers&amp;amp;rdquo; list is there because I kept getting a trailing &amp;amp;ldquo;One thing worth noting&amp;amp;hellip;&amp;amp;rdquo; paragraph on literally every response, and telling it not to add caveats just made it rephrase the caveat. Listing the exact phrases and saying don&amp;amp;rsquo;t rephrase to evade the list was what finally worked (mostly).&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;&amp;amp;ldquo;Never claim credit&amp;amp;rdquo; is there because it has a habit of annotating your idea with &amp;amp;ldquo;as I suggested earlier&amp;amp;rdquo; which is both annoying and frequently not true.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;&amp;amp;ldquo;Verify before asserting&amp;amp;rdquo; is the docs point below in a different form. It will happily tell you how some library behaves without ever having opened the library.&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;p&amp;gt;This gets you maybe 70% of the way there. The rest you do it live.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;img src=&amp;#34;/content/ryansvihla/uploads/2026/09/do-it-live.gif&amp;#34; alt=&amp;#34;We&amp;amp;rsquo;ll do it live&amp;#34;&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;simplify-simplify-simplify&amp;#34;&amp;gt;Simplify, simplify, simplify&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Even with the CLAUDE.md in place it&amp;amp;rsquo;ll drift back into long confusing answers, especially when the question has any nuance to it. If I do not understand what it is saying, or there are too many conflicting items in the response, I tell it to simplify and repeat until it is down to the level of directness I actually need. Some real prompts from my history:&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I asked why some tests were Unix only and got this:&amp;lt;/p&amp;gt;
&amp;lt;blockquote&amp;gt;
&amp;lt;p&amp;gt;Because the worker doesn&amp;amp;rsquo;t run on Windows at all. The worker is the process that actually hosts an agent session, and its whole runtime is &amp;lt;code&amp;gt;#[cfg(unix)]&amp;lt;/code&amp;gt; — the sub-agent socket is just one piece of it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The code says so directly, at &amp;lt;code&amp;gt;mj-worker/src/mcp_stdio.rs:297&amp;lt;/code&amp;gt;:&amp;lt;/p&amp;gt;
&amp;lt;blockquote&amp;gt;
&amp;lt;p&amp;gt;Workers run on Unix; the servers compile everywhere so the CLI stays one shape, and say plainly where they cannot run.&amp;lt;/p&amp;gt;
&amp;lt;/blockquote&amp;gt;
&amp;lt;p&amp;gt;That&amp;amp;rsquo;s the design: Windows builds the same CLI so the commands are identical, and the worker-side functions are stubs that return a clear error instead of silently missing. The transport itself is a Unix domain socket via &amp;lt;code&amp;gt;std::os::unix::net&amp;lt;/code&amp;gt;, which Rust&amp;amp;rsquo;s standard library only provides on Unix.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I didn&amp;amp;rsquo;t make that call — it predates this work. My change only lined up two test helpers with the gate their subject already had.&amp;lt;/p&amp;gt;
&amp;lt;/blockquote&amp;gt;
&amp;lt;p&amp;gt;This giant wall of text seemed just unnecessary to parse, it had a lot repetition, and some uncessary data, qualifications (&amp;amp;ldquo;That&amp;amp;rsquo;s the design&amp;amp;rdquo;, &amp;amp;ldquo;the transport itself&amp;amp;rdquo;, &amp;amp;ldquo;the servers compile everywhere&amp;amp;rdquo;) so I just said &amp;lt;code&amp;gt;simplify&amp;lt;/code&amp;gt; and got this:&amp;lt;/p&amp;gt;
&amp;lt;blockquote&amp;gt;
&amp;lt;p&amp;gt;Mjolnir&amp;amp;rsquo;s worker only runs on Unix by design. Windows gets the CLI, not the worker. So sub-agent code on Windows is stubs that return &amp;amp;ldquo;needs a Unix platform.&amp;amp;rdquo;&amp;lt;/p&amp;gt;
&amp;lt;/blockquote&amp;gt;
&amp;lt;p&amp;gt;WAY more parsable without a lot of garbage, I took this in and saved myself some Ulysses level parsing of text. Here is another one, this time I asked if we could use a different approach to read usage numbers on Mac:&amp;lt;/p&amp;gt;
&amp;lt;blockquote&amp;gt;
&amp;lt;p&amp;gt;No — &amp;lt;code&amp;gt;claude -p /usage&amp;lt;/code&amp;gt; has no parser here. &amp;lt;code&amp;gt;successful_refresh_output&amp;lt;/code&amp;gt; (&amp;lt;code&amp;gt;mj-controller/src/claude_usage.rs:197&amp;lt;/code&amp;gt;) only checks the exit status; nothing reads its stdout, and the module header already calls print-mode usage output approximate. Switching mac to it means writing a text parser for an unstable format and getting worse numbers.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The gap isn&amp;amp;rsquo;t the API, it&amp;amp;rsquo;s where the token comes from. &amp;lt;code&amp;gt;read_credentials&amp;lt;/code&amp;gt; only reads &amp;lt;code&amp;gt;&amp;amp;lt;home&amp;amp;gt;/.credentials.json&amp;lt;/code&amp;gt;. On macOS the same &amp;lt;code&amp;gt;{&amp;amp;quot;claudeAiOauth&amp;amp;quot;: {...}}&amp;lt;/code&amp;gt; JSON is in the Keychain under &amp;lt;code&amp;gt;Claude Code-credentials&amp;lt;/code&amp;gt; — &amp;lt;code&amp;gt;setup.rs:451&amp;lt;/code&amp;gt; already reads it with &amp;lt;code&amp;gt;security find-generic-password -s &amp;amp;quot;Claude Code-credentials&amp;amp;quot; -w&amp;lt;/code&amp;gt;, and its comment notes that &amp;lt;code&amp;gt;security&amp;lt;/code&amp;gt; is pre-authorized for the item so it raises no prompt. That blob feeds &amp;lt;code&amp;gt;oauth_access_token&amp;lt;/code&amp;gt; unchanged.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;There&amp;amp;rsquo;s a third source too: &amp;lt;code&amp;gt;mj login --setup-token&amp;lt;/code&amp;gt; stores a long-lived token at &amp;lt;code&amp;gt;claude_oauth_token_path(profile_id)&amp;lt;/code&amp;gt;, which sessions use via &amp;lt;code&amp;gt;apply_claude_setup_token&amp;lt;/code&amp;gt; but quota never reads.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;So the fix I&amp;amp;rsquo;d make is a source chain for the usage API — credentials file, then the stored setup token, then the macOS Keychain — leaving &amp;lt;code&amp;gt;claude -p /usage&amp;lt;/code&amp;gt; doing only what it does now, forcing a refresh. Want me to implement that?&amp;lt;/p&amp;gt;
&amp;lt;/blockquote&amp;gt;
&amp;lt;p&amp;gt;Again a lot of noise, uncessary information and just a lot to parse for a simple question, instead of making my head hurt figuring out if it actually answered my question, I just said &amp;lt;code&amp;gt;simpler too much fluff&amp;lt;/code&amp;gt;:&amp;lt;/p&amp;gt;
&amp;lt;blockquote&amp;gt;
&amp;lt;p&amp;gt;No. Nothing parses that output — &amp;lt;code&amp;gt;successful_refresh_output&amp;lt;/code&amp;gt; only checks the exit code.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;On mac the token is in the Keychain in the same JSON shape the API path already parses. Fall back to &amp;lt;code&amp;gt;security find-generic-password -s &amp;amp;quot;Claude Code-credentials&amp;amp;quot; -w&amp;lt;/code&amp;gt; when &amp;lt;code&amp;gt;.credentials.json&amp;lt;/code&amp;gt; is missing. Want that?&amp;lt;/p&amp;gt;
&amp;lt;/blockquote&amp;gt;
&amp;lt;p&amp;gt;I will be honest I did more rounds after this. The short answer was still wrong, but at least I was sure now.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Some notes:&amp;lt;/p&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;It usually takes two or three rounds.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;The shorter answer is almost always the same answer with less around it, so you are not losing anything.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;If it is still confusing after a few rounds it is probably not a verbosity problem, it does not understand the question. In that case I give it the structure I want (a numbered list, a table, an &amp;lt;code&amp;gt;a -&amp;amp;gt; b -&amp;amp;gt; c&amp;lt;/code&amp;gt; chain) and have it fill that in rather than letting it pick how to lay the answer out.&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;h2 id=&amp;#34;make-it-read-the-docs&amp;#34;&amp;gt;Make it read the docs&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The other thing is it will implement against a library or an API from memory, and its memory is often out of date or just wrong for anything outside the popular parts of the API. It does not go look unless you tell it to, so I always tell it to. Some real prompts:&amp;lt;/p&amp;gt;
&amp;lt;pre tabindex=&amp;#34;0&amp;#34;&amp;gt;&amp;lt;code&amp;gt;read this https://&amp;amp;lt;docs url&amp;amp;gt;
&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34;&amp;gt;&amp;lt;code&amp;gt;can you look at the sdk docs for &amp;amp;lt;X&amp;amp;gt; and look at the release notes for &amp;amp;lt;X&amp;amp;gt; then look at
the feature set for &amp;amp;lt;Y&amp;amp;gt; and try and find any gaps in the new functionality
&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34;&amp;gt;&amp;lt;code&amp;gt;you can search the web for &amp;amp;lt;X&amp;amp;gt;, you will need to do so to find the &amp;amp;lt;X&amp;amp;gt; source and docs
&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34;&amp;gt;&amp;lt;code&amp;gt;I mean can you read the &amp;amp;lt;X&amp;amp;gt; sdk surely we can start the server with some flag or something
&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;p&amp;gt;The last one was after it told me something was not supported. It was supported, it was in the docs, it just had not read them.&amp;lt;/p&amp;gt;
&amp;lt;h3 id=&amp;#34;my-advice&amp;#34;&amp;gt;My advice&amp;lt;/h3&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;Give it the actual URL, it will fetch it. If the docs are in a repo tell it to clone the repo and read the source.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;When it says something is not possible ask if it read the docs first.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;The &amp;amp;ldquo;Verify before asserting&amp;amp;rdquo; section in the CLAUDE.md above helps some but does not replace saying it in the prompt.&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;h2 id=&amp;#34;summary&amp;#34;&amp;gt;Summary&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;That is really it. Tell it to simplify until you can read the answer, and tell it to read the docs before it writes code against them. Boring but simple, and then the model is genuinely very solid and will do often do proper implementatios.&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Ryan Svihla</name>
      </author>
      

      

      
        <category term="AI" />
      
        <category term="Claude" />
      
        <category term="Opus 5" />
      
        <category term="Coding Agents" />
      

      
    </entry>
  
    <entry>
      <title type="html">Big Design Up Front Is Responsible Again</title>
      <link href="https://lostechies.com/erichexter/2026/07/31/big-design-up-front-is-responsible-again/" rel="alternate" type="text/html" title="Big Design Up Front Is Responsible Again" />
      <published>2026-07-31T09:00:00Z</published>
      <updated>2026-07-31T09:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/07/31/big-design-up-front-is-responsible-again/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/07/31/big-design-up-front-is-responsible-again/">&amp;lt;p&amp;gt;Agile taught us to avoid big design up front. That was the right call — for the constraints we had. Those constraints are gone, and I think the advice inverted.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Here&amp;amp;rsquo;s why we avoided it. Two things were slow and expensive: &amp;lt;strong&amp;gt;writing the spec, and building the software.&amp;lt;/strong&amp;gt; Both were done by hand. So a detailed upfront design was a large bet you couldn&amp;amp;rsquo;t easily change, written before you knew if it was right. The sane response was to write less of it, ship a thin slice, and learn by building the wrong thing first. For twenty years that was the responsible move. I&amp;amp;rsquo;ve shipped plenty of thin slices that discovered the wrong product — expensively — and called it agility.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;AI deleted both of those costs.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Last week I designed a whole application before writing a line of it. Not a napkin sketch — a complete, traceable spec: stakeholder needs, system requirements, architecture with sequence diagrams, down to component-level tickets. Then it was rendered into a &amp;lt;strong&amp;gt;narrated video walkthrough with real UI screens&amp;lt;/strong&amp;gt;, and I handed that to the customer to review. Pause, draw on a frame, leave a voice note.&amp;lt;/p&amp;gt;
&amp;lt;img src=&amp;#34;/content/erichexter/uploads/2026/07/design-review-frame.png&amp;#34; alt=&amp;#34;A design review, rendered: a real UI screen with narration, reviewed before any code was written&amp;#34; style=&amp;#34;max-width:100%&amp;#34; /&amp;gt;
&amp;lt;p&amp;gt;Writing that spec used to take weeks of typing. It took hours. Building a throwaway version just to get real feedback used to be the only option. Now I get feedback on the design itself, at full fidelity, before building anything.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The part that matters: &amp;lt;strong&amp;gt;this is not waterfall.&amp;lt;/strong&amp;gt; Waterfall failed because the big upfront design was unvalidated &amp;lt;em&amp;gt;and&amp;lt;/em&amp;gt; expensive to change — by the time reality showed up, you were committed. Flip both of those. The design is now cheap to produce, cheap to change, and cheap to validate: render a new video, get feedback, re-render. That is big design up front &amp;lt;em&amp;gt;with&amp;lt;/em&amp;gt; a tight feedback loop — the thing waterfall couldn&amp;amp;rsquo;t afford and agile gave up on.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Put it in a grid. Waterfall: expensive to produce, expensive to change. Agile: so skip the design. What we have now: cheap to produce, cheap to change, cheap to validate. Different quadrant. Different rules.&amp;lt;/p&amp;gt;
&amp;lt;img src=&amp;#34;/content/erichexter/uploads/2026/07/design-screen-set.png&amp;#34; alt=&amp;#34;A full set of designed screens generated for review before implementation&amp;#34; style=&amp;#34;max-width:100%&amp;#34; /&amp;gt;
&amp;lt;p&amp;gt;So where did the bottleneck go? It was never typing. It was &amp;lt;em&amp;gt;typing speed&amp;lt;/em&amp;gt; — for specs and for code — that forced us to go slow and iterate. The model writes specs at the speed of light and builds at the speed of light. The only scarce thing left is &amp;lt;strong&amp;gt;judgment&amp;lt;/strong&amp;gt;: deciding what&amp;amp;rsquo;s right, and getting the humans aligned before you commit.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Spend your time there. Do the design work you&amp;amp;rsquo;ve been skipping since 2005 — you can finally afford it. Refine it. Review it properly, with people who aren&amp;amp;rsquo;t going to read a forty-page doc but will watch a five-minute video and tell you it&amp;amp;rsquo;s wrong. Then let the model build it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Concretely, the loop I&amp;amp;rsquo;m running now:&amp;lt;/p&amp;gt;
&amp;lt;ol&amp;gt;
&amp;lt;li&amp;gt;A living spec that deepens by level — stakeholder need, system, architecture, components — each traceable to the next.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Every level rendered into a narrated review video, approved before the design goes deeper.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;The final design compiles to tickets.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;The model implements; you verify against the spec you already agreed on.&amp;lt;/li&amp;gt;
&amp;lt;/ol&amp;gt;
&amp;lt;p&amp;gt;The responsible move in 2026 isn&amp;amp;rsquo;t &amp;amp;ldquo;move fast and iterate.&amp;amp;rdquo; It&amp;amp;rsquo;s &amp;lt;strong&amp;gt;get it right, then let the machine move fast.&amp;lt;/strong&amp;gt;&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      
        <category term="AI" />
      
        <category term="Agile" />
      
        <category term="Requirements" />
      
        <category term="Waterfall" />
      

      
    </entry>
  
    <entry>
      <title type="html">23 Models, One Weekend, Final Picks</title>
      <link href="https://lostechies.com/erichexter/2026/06/06/local-llm-bench-part-5-final-picks/" rel="alternate" type="text/html" title="23 Models, One Weekend, Final Picks" />
      <published>2026-06-06T12:00:00Z</published>
      <updated>2026-06-06T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/06/06/local-llm-bench-part-5-final-picks/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/06/06/local-llm-bench-part-5-final-picks/">&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Part 5 of 5 in the &amp;lt;a href=&amp;#34;/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/&amp;#34;&amp;gt;Local LLM Bench series&amp;lt;/a&amp;gt;.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The project started with ten models and two prompts. It ended with 23 models, a 13-point scoring harness, 3 Python agentic tasks, and more surprises per hour than I expected. This is the final leaderboard and the honest verdict.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;expanding-to-23-models&amp;#34;&amp;gt;Expanding to 23 Models&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;After the initial ten-model run, I pulled thirteen more based on a mix of research agent recommendations and community signals. The research was right about some things and wrong about others.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;It correctly killed two obvious traps. qwen2.5vl is a vision model, not a coder — the &amp;amp;ldquo;vl&amp;amp;rdquo; should have been the clue but I wanted confirmation. qwen3.5:27b is a thinking model that burns its token budget on internal reasoning before producing output; on 16GB VRAM with a standard context budget it hits the wall and times out on every agentic task. Both of those were correct calls.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Then there was cogito:14b. The research said: skip it, superseded, runs 2-3 points behind qwen2.5. I almost listened. What actually happened when I ran cogito: 11-second code generation on the fizzbuzz task, 100/100 agentic score, both edit formats working cleanly. The research was wrong. Cogito turned out to be the fastest sweet-spot model I tested, and it passed tasks that models with higher single-shot scores failed entirely.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Two tag hallucinations also surfaced during pulls. qwen3.5:9b doesn&amp;amp;rsquo;t exist — only the 27B is available. qwen3-vl:8b doesn&amp;amp;rsquo;t exist — only the 235B is available. The research had the right model families but invented specific version tags. The fix is always the same: check ollama.com/library before pulling. Don&amp;amp;rsquo;t trust a model recommendation that includes a specific tag without verifying.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-pi-harness-experiment&amp;#34;&amp;gt;The Pi Harness Experiment&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Alongside the expanded model pool, I tested a different agentic harness entirely. Pi is fundamentally different from aider: instead of receiving structured edit instructions, the model gets direct Bash tool access and can run &amp;lt;code&amp;gt;dotnet new&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;dotnet test&amp;lt;/code&amp;gt;, and anything else itself. It operates as an autonomous loop rather than a guided editor.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I ran devstral and qwen3-coder through pi on two tasks: fizzbuzz-plus and csv-parser. Both timed out at 1020 seconds. Not close calls — full exhaustion, zero useful output across both models and both tasks.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The root cause is that pi is designed for models fine-tuned for tool-calling loops: NousResearch Hermes-class, OpenClaw, models explicitly trained to keep calling tools autonomously and self-terminate when done. Devstral and qwen3-coder via Ollama&amp;amp;rsquo;s OpenAI-compat API don&amp;amp;rsquo;t have that fine-tuning. They can use tools when prompted, but they don&amp;amp;rsquo;t have the trained instinct to keep invoking tools in sequence until a test passes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The thing pi taught me even while failing: harness design is not neutral. An aider task prompt and a pi task prompt are different programs. The model receives different inputs, operates under different constraints, and requires different trained behaviors to succeed. A 100/100 aider score does not predict pi performance, and vice versa. If a Hermes-class model shows up in Ollama&amp;amp;rsquo;s library with solid benchmark numbers, pi is worth revisiting. Until then, aider is the right tool for local 14-30B models.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-scoring-expansion&amp;#34;&amp;gt;The Scoring Expansion&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The harness also grew. I extended the single-shot tests from 10 to 13 points by adding three new probes: a math word problem (3 apples at $0.50 plus 4 oranges at $0.75, reply with only the dollar amount), a JSON output test (return a JSON array of 3 programming languages, nothing else), and a sequence test (output 1 through 5, one per line, nothing else).&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;These three tests turned out to be more discriminating than I expected. Ten of twenty-three models fail the $4.50 math test — not because they get the arithmetic wrong, but because they reason aloud about the problem instead of answering it. The sequence test catches models that follow instructions in general but can&amp;amp;rsquo;t suppress the urge to add a brief explanation. The JSON test catches models that can&amp;amp;rsquo;t stop themselves from wrapping output in markdown fences when explicitly told not to.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;None of these tests are hard. All of them reveal something real about how a model behaves when you need it to produce structured output on command.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;three-new-python-agentic-tasks&amp;#34;&amp;gt;Three New Python Agentic Tasks&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The agentic suite expanded to include three Python tasks alongside the existing C# work. The tasks: a markdown-to-html converter (implement &amp;lt;code&amp;gt;md_to_html()&amp;lt;/code&amp;gt;, 10 pytest tests covering headers, bold, italic, inline code, and links), a JSON validator (implement &amp;lt;code&amp;gt;validate(data, schema)&amp;lt;/code&amp;gt; returning error strings, 9 pytest cases covering required fields, type checking, and enum validation), and a word-frequency counter (implement &amp;lt;code&amp;gt;top_words(text, n)&amp;lt;/code&amp;gt; returning top-N tuples sorted by count descending then alphabetically, 8 pytest cases).&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I ran these on seven models: devstral, qwen3-coder, phi4, hermes3, qwen2.5-coder, mistral-small3.2, and codestral. The results reshuffled the leaderboard in ways the single-shot scores did not predict.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-full-leaderboard&amp;#34;&amp;gt;The Full Leaderboard&amp;lt;/h2&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Size&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;SS /13&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Chat ms&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Code ms&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Agentic Best&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Agentic Pass%&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;~12B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;12/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,918&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;603&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0% (0/2)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;devstral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;~24B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;11/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;16,875&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,246&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;83% (5/6)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:26b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;26B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;11/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;11,029&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,255&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0% (0/2)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3.5:27b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;27B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;11/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;24,810&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7,222&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0% (timeout)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;deepseek-r1:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,286&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;561&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;glm-4.7-flash&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;30B MoE&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8,843&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;2,531&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0% (timeout)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;granite4:32b-a9b-h&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;32B MoE&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20,885&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,125&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0%&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,221&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;475&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0%&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5vl:7b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5,783&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;863&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3-coder:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;30B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9,948&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;2,143&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;67% (4/6)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,876&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;523&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0%&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;cogito:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,447&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;438&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;hermes3:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;~8B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,756&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;280&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;40% (2/5)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;mistral-small3.2:24b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;24B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;12,169&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,228&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100% (3/3)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;mistral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3,335&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;323&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0%&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;codestral:22b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;22B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;17,182&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;2,427&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;67% (2/3)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;deepseek-coder-v2:16b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;16B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,516&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;298&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;llava:7b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;4,045&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;292&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;magistral:24b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;24B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;22,802&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;11,568&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gpt-oss:20b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8,751&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8,915&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0%&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;phi4:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6,415&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;466&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;50% (2/4)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5-coder:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5,989&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;529&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100% (4/4)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;30B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3/13&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14,866&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10,749&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;—&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;qwen3.5:27b, gpt-oss:20b, and deepseek-r1:14b are thinking models — they burn context on internal reasoning before producing visible output. The scores reflect that.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-surprising-results&amp;#34;&amp;gt;The Surprising Results&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;gemma4:latest.&amp;lt;/strong&amp;gt; 12/13 single-shot, 603ms code generation, fastest chat in its size class. Zero percent agentic pass rate across every task it attempted. This is the sharpest split in the entire dataset. gemma4 is an excellent model for answering questions. It has no working mental model of &amp;amp;ldquo;I am in a multi-turn loop writing files until tests pass.&amp;amp;rdquo; Those are different capabilities. The single-shot tests reward the former. The agentic tasks require the latter. gemma4 nails one and is completely useless at the other.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;mistral-small3.2:24b.&amp;lt;/strong&amp;gt; I almost missed this one entirely. It had no agentic run history going into the final Python task batch — it just hadn&amp;amp;rsquo;t come up in earlier experiments. When I finally ran it, it swept all three new Python tasks with 100/100 scores on first attempt, finishing each in 26 to 52 seconds. Nine out of 13 on single-shot. It had minimal community attention during the bench period. It turned out to be one of the two most reliable agentic performers I tested. The lesson here: community signal is a useful prior, not a substitute for running the test.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;qwen2.5-coder:14b.&amp;lt;/strong&amp;gt; 6/13 on single-shot. That score is a lie in the specific direction that matters most. The instruction-following tests fail consistently. The code generation test produces output that compiles but gets the wrong answer. On every agentic task I ran it on, it passed. Four for four, 100% pass rate. The single-shot harness penalizes its tendency to reason aloud before writing code. In an agentic loop, that verbosity doesn&amp;amp;rsquo;t hurt — aider just waits for the edit block, and the edit block is correct. Single-shot actively mispredicts this model&amp;amp;rsquo;s real-world utility.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;hermes3:latest.&amp;lt;/strong&amp;gt; 280ms code generation. The fastest model in the field by a significant margin, and at 4.7GB it&amp;amp;rsquo;s the lightest serious option. 3,756ms average chat latency, also fastest. It scored 100/100 on csv-scaffolded with a 25-second wall time — another field record. Then it scored 10/100 on fizzbuzz and instant-failed on json-validator in zero turns. The inconsistency pattern makes sense for a model fine-tuned specifically for tool use and short completions: it handles the tasks that match its training profile well and falls apart outside them. For anyone doing rapid-fire chat or simple completions at scale, hermes3 is the answer. For general agentic coding, the brittleness is a real problem.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;phi4:14b.&amp;lt;/strong&amp;gt; 6/13 on single-shot; 100/100 on fizzbuzz and word-freq. It failed markdown-to-html and json-validator, and both failures have the same signature: 16 to 17% context utilization, then the output starts spiraling. phi4 has a 16K context ceiling, and tasks that grow their working context over multiple iterations hit that wall. The context limit is the only thing preventing phi4 from joining the reliable agentic tier. With 32K context or better, I&amp;amp;rsquo;d expect it to pass everything it currently fails.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;codestral:22b.&amp;lt;/strong&amp;gt; The markdown-to-html task produced a unicode crash — aider&amp;amp;rsquo;s display layer choked on an arrow character in a CP1252 terminal. json-validator and word-freq both passed 100/100. That markdown failure is an environment bug, not a model failure. I&amp;amp;rsquo;m counting it in the pass rate because I can&amp;amp;rsquo;t retroactively change the environment it ran in, but anyone testing codestral in a UTF-8 terminal should expect a different result.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-actual-picks&amp;#34;&amp;gt;The Actual Picks&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;For coding work on a 16GB machine, the answer depends on what you&amp;amp;rsquo;re doing.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;If you&amp;amp;rsquo;re working in a new codebase — multi-file, complex scaffolding, scratch-to-working-tests — use devstral:latest. It&amp;amp;rsquo;s the only model in this pool that reliably handles multi-file C# from scratch. 83% agentic pass rate across six diverse tasks spanning C# and Python. Not the fastest at 3 to 20 seconds per response, but it has the highest ceiling and it doesn&amp;amp;rsquo;t fall apart on complexity.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;If you&amp;amp;rsquo;re working in an existing codebase — the actual everyday case, where you&amp;amp;rsquo;re editing files that already exist — use qwen3-coder:30b. 100/100 on Python tasks, strong on scaffolded C#, 2-second code generation. The whole edit format is mandatory; diff mode fails silently and produces nothing. Get the format right and this model is very fast for its size.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;If VRAM is the constraint, use qwen2.5-coder:14b. It runs on about 9GB, which means it fits alongside other processes. It passed every agentic task I ran it on. The 6/13 single-shot score is misleading — ignore it for agentic work.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;mistral-small3.2:24b is on a watch list. Three tasks run, three passed. That&amp;amp;rsquo;s not enough data to promote it above devstral for serious work, but it&amp;amp;rsquo;s enough to keep it in the rotation. If it holds 100% across ten more tasks I&amp;amp;rsquo;ll move it up.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;For chat and Q&amp;amp;amp;A, the picks are different. gemma4:latest for quality — 12/13, fast for its size, clean outputs. Don&amp;amp;rsquo;t use it for anything agentic. For speed, hermes3:latest at 4.7GB and 280ms code generation is the answer, especially if you&amp;amp;rsquo;re running it alongside something else or doing high-volume completions.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-single-shot-scores-actually-measure&amp;#34;&amp;gt;What Single-Shot Scores Actually Measure&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;This question came up enough during the project that it deserves a direct answer.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Single-shot scores measure whether a model understands what it&amp;amp;rsquo;s being asked, can produce a well-formed response on one shot, and follows tight output constraints. That&amp;amp;rsquo;s genuinely useful for chatting, summarizing, classifying, and answering questions. The score is predictive for those tasks.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;What it does not measure: will this model keep working across turns, will it understand its own previous outputs, can it handle a tool returning an unexpected result, will it know when to stop and verify rather than spiraling, can it write files instead of prose. Those are the capabilities that determine agentic performance. They don&amp;amp;rsquo;t show up in any single-prompt test because by design they can&amp;amp;rsquo;t — they require multiple turns to observe.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The practical implication is that running a 13-point single-shot harness before picking a coding model will tell you roughly nothing about whether the model can actually do the coding work. You have to run the agentic task. There is no shortcut.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;closing&amp;#34;&amp;gt;Closing&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Six weeks. 23 models. 630 lines of harness code. 50&#43; agentic task runs. The answer to &amp;amp;ldquo;which local model can actually code?&amp;amp;rdquo; turns out to be a different question depending on what you mean by coding.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The model that tops the single-shot leaderboard is the one to use for chat. The model that wins at agentic coding tasks is a different model entirely. I spent a weekend thinking gemma4 was the obvious answer before it timed out on every real task I gave it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The bench application and all results are at &amp;lt;a href=&amp;#34;https://github.com/erichexter/ollama-model-bench&amp;#34;&amp;gt;github.com/erichexter/ollama-model-bench&amp;lt;/a&amp;gt;. The harness accepts any model Ollama can serve — pull it, add an entry to the settings file, run it. The numbers here are reproducible on any machine with 16GB of VRAM. If you find something that beats devstral on multi-file from scratch, I want to know about it.&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
    <entry>
      <title type="html">The Config That Changed Everything</title>
      <link href="https://lostechies.com/erichexter/2026/06/03/local-llm-bench-part-4-harness-optimization/" rel="alternate" type="text/html" title="The Config That Changed Everything" />
      <published>2026-06-03T12:00:00Z</published>
      <updated>2026-06-03T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/06/03/local-llm-bench-part-4-harness-optimization/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/06/03/local-llm-bench-part-4-harness-optimization/">&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Part 4 of 5 in the &amp;lt;a href=&amp;#34;/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/&amp;#34;&amp;gt;Local LLM Bench series&amp;lt;/a&amp;gt;.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;After Part 3&amp;amp;rsquo;s 1-in-6 pass rate, I had a theory about qwen3-coder. The model scored 0/100 not because it couldn&amp;amp;rsquo;t write C#, but because aider couldn&amp;amp;rsquo;t parse what it wrote. If the failure was format mismatch, then fixing the format should fix the score.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I was right. One line in a YAML file took qwen3-coder:30b from 0/100 to 100/100. Twenty-six seconds. Same model, same task, same hardware.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;That result rewrites how I think about local model evaluation.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-edit_format-lever&amp;#34;&amp;gt;The edit_format Lever&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;aider supports two primary edit modes. In &amp;lt;code&amp;gt;diff&amp;lt;/code&amp;gt; mode, the model sends back git-style patches — only the changed lines, with surrounding context. In &amp;lt;code&amp;gt;whole&amp;lt;/code&amp;gt; mode, the model sends back the entire file contents. These are not stylistic preferences. They require completely different output from the model, and models are not equally capable of both.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The research I ran before Phase 9 turned up a finding I didn&amp;amp;rsquo;t take seriously enough at the time: &amp;amp;ldquo;harness mismatch is bigger than model choice.&amp;amp;rdquo; One real-world study cited 6x performance variation purely from harness configuration changes, holding the model constant. I read that and thought it was probably overstated. Then I ran the A/B.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The &amp;lt;code&amp;gt;.aider.model.settings.yml&amp;lt;/code&amp;gt; file lets you configure per-model settings. The critical field is &amp;lt;code&amp;gt;edit_format&amp;lt;/code&amp;gt;. Here&amp;amp;rsquo;s what qwen3-coder&amp;amp;rsquo;s entry looks like after the fix:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-yaml&amp;#34; data-lang=&amp;#34;yaml&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;- &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;name&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;l&amp;#34;&amp;gt;ollama_chat/qwen3-coder:30b&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;edit_format&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;l&amp;#34;&amp;gt;whole&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;use_repo_map&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;kc&amp;#34;&amp;gt;false&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;extra_params&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;num_ctx&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;65536&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;Before this change: &amp;lt;code&amp;gt;edit_format&amp;lt;/code&amp;gt; was unset, defaulting to &amp;lt;code&amp;gt;diff&amp;lt;/code&amp;gt;. After: &amp;lt;code&amp;gt;whole&amp;lt;/code&amp;gt;. The model behavior changes completely.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-ab-results&amp;#34;&amp;gt;The A/B Results&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;I ran six models against both formats on the fizzbuzz-plus sweet-spot task:&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;whole&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;diff&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3-coder:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100 (26s)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0/100 FAIL&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;devstral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100 (53s)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100 (98s)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5-coder:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100 (73s)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;100/100 (65s)&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gpt-oss:20b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;mistral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100 FAIL&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;Three models work. Three models don&amp;amp;rsquo;t. The format A/B cleanly separates the populations. gpt-oss, qwen3:14b, and mistral fail in both formats — those are genuine capability problems, not configuration problems. qwen3-coder was a false negative: the code was right, the format was wrong, the score said zero.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;devstral and qwen2.5-coder work in both formats, which tells you something about their training. They&amp;amp;rsquo;ve been explicitly tuned to produce structured edit blocks. qwen3-coder has not — or at least not in the diff format aider expects. Switching to whole file output removes the constraint entirely: just dump the file, let aider handle the diff computation. qwen3-coder is very good at writing complete, correct files.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-thinking-mode-problem&amp;#34;&amp;gt;The Thinking-Mode Problem&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Three models that looked promising on paper — gpt-oss:20b, deepseek-r1:14b, and qwen3.5:27b — share a different failure mode. They all run in &amp;amp;ldquo;thinking mode&amp;amp;rdquo;: before producing any code output, they generate thousands of internal reasoning tokens. On single-shot tasks this is invisible; the &amp;lt;code&amp;gt;&amp;amp;lt;think&amp;amp;gt;&amp;lt;/code&amp;gt; block appears in a separate field and the user only sees the final answer. On an agentic task with a 300-second timeout, the thinking block alone can exhaust the budget.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;gpt-oss, deepseek-r1, and qwen3.5 all timeout at zero turns — the model thought itself to death before writing a single line of code.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The fix for qwen3 models (not qwen3.5, which has different training) is a &amp;lt;code&amp;gt;/no_think&amp;lt;/code&amp;gt; prefix in the aider system prompt:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-yaml&amp;#34; data-lang=&amp;#34;yaml&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;- &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;name&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;l&amp;#34;&amp;gt;ollama_chat/qwen3:14b&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;edit_format&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;l&amp;#34;&amp;gt;whole&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;system_prompt_prefix&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;/no_think&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;use_temperature&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;0.7&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;extra_params&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;num_ctx&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;32768&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;top_p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;0.8&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;top_k&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;20&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;This worked for qwen3:14b and qwen3:30b. It does nothing for qwen3.5 — different model family, different training, the prefix is ignored. qwen3.5:27b is a 17GB model on 16GB VRAM, so it&amp;amp;rsquo;s partially spilling to RAM anyway. At mixed CPU/GPU generation speed with a thinking block running first, it cannot produce useful output inside 300 seconds. The hardware ceiling and the thinking penalty compound each other. Model eliminated.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-num_ctx-revelation&amp;#34;&amp;gt;The num_ctx Revelation&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Ollama&amp;amp;rsquo;s default context window is 2048 tokens. That&amp;amp;rsquo;s not 2048 for the task — that&amp;amp;rsquo;s 2048 for the entire conversation, including the system prompt, the file content, the task description, and every prior exchange. For an agentic coding session where aider is sending file contents back and forth, 2048 fills in two or three turns. After that, the model is working with a truncated view of its own conversation. It starts looping, contradicting itself, or deleting code it just wrote.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Ollama doesn&amp;amp;rsquo;t warn you when it truncates. It silently discards the oldest tokens and keeps going. The model&amp;amp;rsquo;s outputs start looking confused on turn three and you assume it&amp;amp;rsquo;s a capability problem. It isn&amp;amp;rsquo;t.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Setting &amp;lt;code&amp;gt;num_ctx: 32768&amp;lt;/code&amp;gt; (or 65536 for the larger models) unlocks stable multi-turn behavior. Several failures that looked like model confusion were actually context truncation. The fix is one line per model in the YAML.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-architect-mode-dead-end&amp;#34;&amp;gt;The Architect Mode Dead End&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;I wanted to test whether combining two models — one to plan, one to implement — could improve results on stretch-tier tasks. aider calls this &amp;amp;ldquo;architect mode.&amp;amp;rdquo; In principle: the architect model breaks the task into pieces, the editor model writes the code, and the combination should outperform either alone. It&amp;amp;rsquo;s a reasonable theory. The machine had other plans.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Loading two 14-17GB models on 16GB VRAM means constant unloading and reloading. Every time control switches from architect to editor, Ollama has to evict one model and load the other. That swap is not fast. I ran devstral &#43; qwen3-coder and devstral &#43; qwen2.5-coder. Both pairs hit the five-minute timeout at zero turns. The entire budget went to model swap overhead before a single tool call completed.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Architect mode requires both models to be co-resident in VRAM. On 16GB, that means two models totaling at most 16GB, which limits you to two 7B models — too small to be useful on complex tasks. The minimum viable VRAM for architect mode with 14B&#43; models is 32GB. Below that, single-model runs strictly better.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-scaffolding-experiment&amp;#34;&amp;gt;The Scaffolding Experiment&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;After the format A/B produced clear winners, I wanted to understand what was really limiting the failing models on the csv-parser task. The task asked models to build a C# console app and test project from scratch — which means creating &amp;lt;code&amp;gt;.csproj&amp;lt;/code&amp;gt; files, a solution file, adding project references, restoring NuGet packages, and then writing correct C#. That&amp;amp;rsquo;s two separate problems: .NET project plumbing and C# logic.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I split them apart. The scaffolded version of the task pre-creates everything: both &amp;lt;code&amp;gt;.csproj&amp;lt;/code&amp;gt; files with correct &amp;lt;code&amp;gt;net10.0&amp;lt;/code&amp;gt; targets, a &amp;lt;code&amp;gt;Program.cs&amp;lt;/code&amp;gt; entry point the model doesn&amp;amp;rsquo;t touch, a stub &amp;lt;code&amp;gt;CsvProcessor.cs&amp;lt;/code&amp;gt; with a TODO comment, a test project with a NuGet reference already wired, and stub test method shells. &amp;lt;code&amp;gt;dotnet restore&amp;lt;/code&amp;gt; runs before the model starts. The model&amp;amp;rsquo;s job is to implement one static method and fill in five test bodies.&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;From-scratch&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Scaffolded&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Change&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;devstral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;70/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;90/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&#43;20&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3-coder:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;90/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&#43;90&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;cogito:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&#43;10&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;granite4:32b-a9b-h&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&#43;10&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;qwen3-coder was never broken. Its 0/100 on the from-scratch task was entirely a scaffolding failure. It doesn&amp;amp;rsquo;t know how to create a .NET solution structure from the command line — that&amp;amp;rsquo;s a DevOps problem, not a C# problem. Given the structure, it writes correct C# and correct tests in one shot, in 56 seconds. That&amp;amp;rsquo;s four times faster than devstral on the same task.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;cogito:14b and granite4:32b-a9b-h still fail on the scaffolded version. Their problem is C# reasoning, not project structure. The scaffolding experiment drew a clean line between the two failure modes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The practical implication: if you&amp;amp;rsquo;re deploying these models on an existing codebase — the actual real-world use case — the scaffolding problem doesn&amp;amp;rsquo;t exist. The codebase is already there. qwen3-coder becomes a genuine competitor to devstral for existing-codebase work.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;where-the-leaderboard-stands&amp;#34;&amp;gt;Where the Leaderboard Stands&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;After format configuration, context window fixes, and scaffolding experiments, the picture looks like this:&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;For sweet-spot tasks (one or two files, existing codebase, 80-120 lines of code): qwen3-coder:30b at 26 seconds, cogito:14b at 11 seconds on both formats, devstral at 53 seconds, mistral-small3.2:24b at 44 seconds, and qwen2.5-coder:14b at 73 seconds. Five models that work reliably.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;For multi-file from scratch: devstral:latest, confirmed against eight challengers. No other local model in this weight class completes the csv-parser task reliably regardless of configuration.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Eliminated regardless of configuration: gemma4 (all variants), glm-4.7-flash, qwen2.5:14b, qwen3:14b, qwen3.5:27b, deepseek-r1, gpt-oss, magistral — all timeout or fail in both formats. These aren&amp;amp;rsquo;t configuration problems. They&amp;amp;rsquo;re either the wrong model type (thinking models on a 16GB budget), capability gaps, or both.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The 6x performance variation claim from the research turned out to be conservative in at least one case. qwen3-coder went from zero to perfect. You can&amp;amp;rsquo;t express that as a multiplier.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Next up: &amp;lt;a href=&amp;#34;/erichexter/2026/06/06/local-llm-bench-part-5-final-picks/&amp;#34;&amp;gt;Part 5&amp;lt;/a&amp;gt; — expanding the model pool, three surprise entries that research told me to skip, and the final leaderboard after 23 models across six weeks of testing.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
    <entry>
      <title type="html">Single-Shot Lies</title>
      <link href="https://lostechies.com/erichexter/2026/05/31/local-llm-bench-part-3-single-shot-lies/" rel="alternate" type="text/html" title="Single-Shot Lies" />
      <published>2026-05-31T12:00:00Z</published>
      <updated>2026-05-31T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/05/31/local-llm-bench-part-3-single-shot-lies/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/05/31/local-llm-bench-part-3-single-shot-lies/">&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Part 3 of 5 in the &amp;lt;a href=&amp;#34;/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/&amp;#34;&amp;gt;Local LLM Bench series&amp;lt;/a&amp;gt;.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;gemma4:latest scored 10/10 on every test I built. Perfect chat response. Perfect code generation. Perfect tool call. Perfect instruction following. I ran it twice to be sure. Same result. So naturally, when it came time to run the first real agentic coding task, that was the model I reached for.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;It produced zero lines of useful code in ten minutes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;That&amp;amp;rsquo;s the story of Phase 8, and it changed everything about how I think about model evaluation.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-task&amp;#34;&amp;gt;The Task&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The agentic benchmark I built is a CSV parser in C#. A console app that reads a file with Name and Score columns, prints the top 3 scores descending, ties broken alphabetically. Verify with &amp;lt;code&amp;gt;dotnet test&amp;lt;/code&amp;gt;. The task is sized to what I&amp;amp;rsquo;d call &amp;amp;ldquo;stretch tier&amp;amp;rdquo; — two projects, roughly 150 lines of code, multi-file, requires the model to scaffold a .NET solution from scratch and then implement correct logic. A competent human developer does this in about ten minutes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The harness is aider 0.86.2 installed via &amp;lt;code&amp;gt;uv tools&amp;lt;/code&amp;gt;, running headless with &amp;lt;code&amp;gt;--yes-always --exit --message-file&amp;lt;/code&amp;gt;. Scoring: 60 points if the verify command passes, 20 if the model finishes in two iterations or fewer, 10 for no compile errors, 10 for clean edit format. 100 points maximum.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I ran six models: the top performers from Phase 4&amp;amp;rsquo;s single-shot benchmark plus two new additions.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-results&amp;#34;&amp;gt;The Results&amp;lt;/h2&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Score&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Notes&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;devstral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;70/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5 iterations, 147 seconds&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;600 second timeout, 0 turns&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:26b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;600 second timeout&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;glm-4.7-flash&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;20/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;600 second timeout&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;91 seconds, never recovered&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3-coder:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;0/100&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;77 seconds, garbled output&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;One passes. Five fail. The model that aced every single-shot test I designed hits its ten-minute wall and produces nothing. The model that topped the leaderboard with a perfect score is the first casualty.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;devstral is, notably, marketed specifically for agentic coding loops. That framing turned out to matter.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-went-wrong-with-gemma4&amp;#34;&amp;gt;What Went Wrong With gemma4&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;gemma4:latest doesn&amp;amp;rsquo;t fail because it can&amp;amp;rsquo;t write C#. It fails because it doesn&amp;amp;rsquo;t understand that it&amp;amp;rsquo;s supposed to be writing files. When aider sends it a task, it responds with a description of what the code should look like, or it writes a fenced code block in prose, or it explains the approach in detail without producing any actual edits. I watched this happen in real time and it took longer than I&amp;amp;rsquo;d like to admit before I understood what I was seeing. These responses look helpful if you&amp;amp;rsquo;re reading them as a chat assistant. aider can&amp;amp;rsquo;t do anything with them — it&amp;amp;rsquo;s waiting for structured edit blocks that follow its protocol, not a tutorial.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The single-shot benchmark rewarded exactly the behavior that makes gemma4 useless in an agentic loop. &amp;amp;ldquo;Write a Python function that checks if a number is prime&amp;amp;rdquo; — gemma4 produces clean, correct Python instantly. But that task has one shot, one context, one output. There&amp;amp;rsquo;s no concept of a multi-step session, no expectation that the model needs to write files into a directory, no loop where the model gets feedback and tries again.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Ask gemma4 to run a ten-minute coding session and it has no mental model for what &amp;amp;ldquo;running a coding session&amp;amp;rdquo; means. It&amp;amp;rsquo;s a very good chat assistant. That&amp;amp;rsquo;s not the same thing.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-went-wrong-with-qwen3-coder&amp;#34;&amp;gt;What Went Wrong With qwen3-coder&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;qwen3-coder:30b scores 0/100, which looks worse than the timeout failures. It&amp;amp;rsquo;s actually more interesting. The model ran for 77 seconds before aider gave up, which means it produced output — just output that aider silently rejected as malformed edits. The code was probably fine. The format wasn&amp;amp;rsquo;t.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;This is a harness compatibility problem, not a capability problem. aider expects edit blocks in specific formats — either a &amp;lt;code&amp;gt;diff&amp;lt;/code&amp;gt;-style patch or a &amp;lt;code&amp;gt;whole&amp;lt;/code&amp;gt;-file replacement. qwen3-coder was emitting something that resembled neither cleanly enough for aider to parse. aider&amp;amp;rsquo;s response to a malformed edit is to silently skip it, log nothing useful, and eventually exit. From the score sheet, it looks like the model produced nothing. That&amp;amp;rsquo;s not what happened.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;This distinction matters, because it&amp;amp;rsquo;s a clue. If the failure is format mismatch rather than capability, changing the format instruction should fix it. I filed that away and moved on.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-it-means&amp;#34;&amp;gt;What It Means&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The research literature on agentic coding benchmarks describes a roughly 17% pass rate for 14-30B parameter models on what they call &amp;amp;ldquo;stretch tier&amp;amp;rdquo; tasks: multi-file, 150&#43; lines of code, multiple tool-call iterations. My six-model run hit 1-in-6. Exactly 17%.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;That number didn&amp;amp;rsquo;t come from luck. It came from the same thing the research describes: most models that can answer questions well don&amp;amp;rsquo;t have a working mental model of &amp;amp;ldquo;I am operating a computer, I need to write files, I need to keep doing work until a test passes.&amp;amp;rdquo; Those are different cognitive tasks. Single-shot chat benchmarks don&amp;amp;rsquo;t distinguish between them.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The models that time out aren&amp;amp;rsquo;t slower or dumber than devstral. They&amp;amp;rsquo;re not designed for this. gemma4 is optimized to produce a high-quality response to a question. devstral is optimized to take a task and not stop until it&amp;amp;rsquo;s done. The training objectives are different. The behavior is different. The single-shot score captures none of that.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;where-this-leaves-us&amp;#34;&amp;gt;Where This Leaves Us&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;devstral finished the task with 70/100. It needed five iterations instead of two (losing 20 points on the efficiency score), but it shipped working code. None of the other five models produced a single passing test.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The 70/100 score isn&amp;amp;rsquo;t a ceiling — it&amp;amp;rsquo;s a baseline. devstral used the default aider configuration with no tuning. It worked anyway. The question is whether anything else can be made to work, or whether devstral is the only local model that can do this at all.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;qwen3-coder&amp;amp;rsquo;s format failure points toward an answer. If the problem is configuration, not capability, then changing the configuration should change the result. That&amp;amp;rsquo;s the experiment Part 4 runs.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Next up: &amp;lt;a href=&amp;#34;/erichexter/2026/06/03/local-llm-bench-part-4-harness-optimization/&amp;#34;&amp;gt;Part 4&amp;lt;/a&amp;gt; — one config change takes a model from 0/100 to 100/100, and the harness turns out to matter more than the model.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
    <entry>
      <title type="html">Building a .NET 10 Benchmark Harness</title>
      <link href="https://lostechies.com/erichexter/2026/05/28/local-llm-bench-part-2-building-the-harness/" rel="alternate" type="text/html" title="Building a .NET 10 Benchmark Harness" />
      <published>2026-05-28T12:00:00Z</published>
      <updated>2026-05-28T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/05/28/local-llm-bench-part-2-building-the-harness/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/05/28/local-llm-bench-part-2-building-the-harness/">&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Part 2 of 5 in the &amp;lt;a href=&amp;#34;/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/&amp;#34;&amp;gt;Local LLM Bench series&amp;lt;/a&amp;gt;.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The PowerShell script from part one did its job. It surfaced the think-mode problem, sorted out which models could call tools, and gave me rough latency numbers. But it could not tell me whether the code models wrote was actually correct — I was reading output and deciding it looked fine, which is not the same thing as running it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;What I needed was a harness that ran models against defined tasks, verified the outputs mechanically, and produced a repeatable score. I&amp;amp;rsquo;m a C# developer. .NET 10 was already on the machine. The choice was not a choice.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;architecture&amp;#34;&amp;gt;Architecture&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The project is a .NET 10 console application. The core pieces are:&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;OllamaRunner&amp;lt;/strong&amp;gt; is a thin HTTP wrapper around Ollama&amp;amp;rsquo;s &amp;lt;code&amp;gt;/api/generate&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;/api/chat&amp;lt;/code&amp;gt; endpoints. Every request goes out with &amp;lt;code&amp;gt;temperature=0&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;seed=42&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;think=false&amp;lt;/code&amp;gt;. Temperature zero makes results deterministic enough to compare across runs. The seed locks that in further. The &amp;lt;code&amp;gt;think&amp;lt;/code&amp;gt; flag is false by default — models that need it explicitly will be detected and handled.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;RoslynEvaluator&amp;lt;/strong&amp;gt; handles the &amp;lt;code&amp;gt;SumEvens&amp;lt;/code&amp;gt; code test in-process. It takes whatever the model returns, strips any markdown fences, wraps the bare method in a class, and hands it to the Roslyn CSharp scripting API to compile and execute. If it compiles and &amp;lt;code&amp;gt;SumEvens(new[] {1,2,3,4,5})&amp;lt;/code&amp;gt; returns 6, the model passes. This runs entirely in memory with no disk I/O and no subprocess.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;TempProjectRunner&amp;lt;/strong&amp;gt; is where it gets more serious. This component scaffolds actual temporary &amp;lt;code&amp;gt;dotnet&amp;lt;/code&amp;gt; projects, writes model-generated code into them, builds them with &amp;lt;code&amp;gt;dotnet build&amp;lt;/code&amp;gt;, and runs them with &amp;lt;code&amp;gt;dotnet run&amp;lt;/code&amp;gt;. It checks stdout for the expected output. For the test suite portion, it scaffolds a second project alongside the first, adds a project reference, drops in model-generated xUnit test code, and runs &amp;lt;code&amp;gt;dotnet test&amp;lt;/code&amp;gt;. Every project is cleaned up from the temp directory when the run completes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;Scorer&amp;lt;/strong&amp;gt; orchestrates the sequence — chat test, code test, tool test, instruction test, reasoning test, JSON output test, sequence test, Hello World test — and assembles the results into a &amp;lt;code&amp;gt;ModelResult&amp;lt;/code&amp;gt; record.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;ModelResult&amp;lt;/strong&amp;gt; is a straightforward C# record type. Every boolean metric is a property; &amp;lt;code&amp;gt;TotalScore&amp;lt;/code&amp;gt; is a computed getter that sums them. The record also carries timing in milliseconds for each test category and a &amp;lt;code&amp;gt;ThinkRequired&amp;lt;/code&amp;gt; flag that is informational only and does not affect the score.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;ConsoleReporter&amp;lt;/strong&amp;gt; prints the final table to the terminal with ANSI color coding. &amp;lt;strong&amp;gt;ResultStore&amp;lt;/strong&amp;gt; writes the raw results to &amp;lt;code&amp;gt;results/model-results.json&amp;lt;/code&amp;gt; and a human-readable markdown ledger to &amp;lt;code&amp;gt;results/RESULTS.md&amp;lt;/code&amp;gt; after each run.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-code-tests&amp;#34;&amp;gt;The Code Tests&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The first code test is &amp;lt;code&amp;gt;SumEvens&amp;lt;/code&amp;gt;: write a C# method that takes &amp;lt;code&amp;gt;IEnumerable&amp;amp;lt;int&amp;amp;gt;&amp;lt;/code&amp;gt; and returns the sum of even numbers. Return only the method, no class, no namespace, no explanation. This is deliberately narrow. The narrow scope is the point — it is testing whether a model can follow output constraints and write code that compiles and produces correct results, not whether it can write impressive prose around the code.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;RoslynEvaluator wraps the method in a class, invokes it with &amp;lt;code&amp;gt;{1, 2, 3, 4, 5}&amp;lt;/code&amp;gt;, and checks that the result is 6. Compile error means the model scores zero on both compile and correct. Compiles but returns the wrong number means compile point awarded, correct point denied. Compiles and returns 6 means full credit.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;hello-world-the-real-test&amp;#34;&amp;gt;Hello World: The Real Test&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The Hello World test is where I learned something useful. The prompt asks the model to write a complete C# console application: a &amp;lt;code&amp;gt;Greeter&amp;lt;/code&amp;gt; class with a public static &amp;lt;code&amp;gt;GetGreeting()&amp;lt;/code&amp;gt; method that returns &amp;lt;code&amp;gt;&amp;amp;quot;Hello, World!&amp;amp;quot;&amp;lt;/code&amp;gt;, plus a Main method or top-level statements that calls it and prints the result. Separately, it asks the model to write xUnit tests for that &amp;lt;code&amp;gt;Greeter&amp;lt;/code&amp;gt; class.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;TempProjectRunner scaffolds a &amp;lt;code&amp;gt;dotnet new console&amp;lt;/code&amp;gt; project, replaces &amp;lt;code&amp;gt;Program.cs&amp;lt;/code&amp;gt; with whatever the model generated, runs &amp;lt;code&amp;gt;dotnet build&amp;lt;/code&amp;gt;, then &amp;lt;code&amp;gt;dotnet run&amp;lt;/code&amp;gt;, and checks stdout for &amp;lt;code&amp;gt;&amp;amp;quot;Hello, World!&amp;amp;quot;&amp;lt;/code&amp;gt;. For the test portion, it scaffolds a &amp;lt;code&amp;gt;dotnet new xunit&amp;lt;/code&amp;gt; project in the same temp directory, adds a project reference to the app, drops in the model&amp;amp;rsquo;s test code as &amp;lt;code&amp;gt;GreeterTests.cs&amp;lt;/code&amp;gt;, runs &amp;lt;code&amp;gt;dotnet build&amp;lt;/code&amp;gt;, and then &amp;lt;code&amp;gt;dotnet test&amp;lt;/code&amp;gt;.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;This turns out to be an excellent proxy for whether a model understands C# project structure. Writing a method is straightforward. Writing a complete application that builds from scratch against a specific framework target, with a class in a form that a separately compiled test project can reference — that is a different problem. Models that understand C# project conventions get it right on the first try. Models that pattern-match on superficial features tend to include the wrong using statements, declare the class in a namespace that the test code does not account for, or produce an entry point that conflicts with the &amp;lt;code&amp;gt;Greeter&amp;lt;/code&amp;gt; class definition.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Each step is gated: if the app does not compile, neither the output check nor the test run happens. If the tests do not compile, the pass/fail result is not recorded. Partial credit is possible — a model can build the app but write tests that compile and then fail at runtime, earning two of the four Hello World points.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;scoring&amp;#34;&amp;gt;Scoring&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The 10-point scoring breakdown for the initial complete run:&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Category&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Points&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Chat response (non-empty, sensible)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;SumEvens compiles&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;SumEvens correct&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Tool call supported (not HTTP 400)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Tool call valid (structured, correct function)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Instruction followed (exactly three words)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Hello World app compiles&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Hello World app correct output&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Hello World tests compile&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;Hello World tests pass&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;After the initial runs I extended the suite with three more tests, bringing the maximum to 13: a reasoning test (a word problem with an exact numeric answer — $4.50, no other text), a JSON output test (produce a valid JSON array of at least three programming language names), and a sequence test (output the numbers 1 through 5, one per line, nothing else). All three are binary pass/fail with no partial credit. The reasoning and sequence tests catch models that ignore output constraints even when the constraint is explicit. Several did.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;unit-tests&amp;#34;&amp;gt;Unit Tests&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The test project covers 13 cases across five test classes. &amp;lt;code&amp;gt;ModelResultTests&amp;lt;/code&amp;gt; verifies that the scoring logic is correct — all true returns the expected sum, all false returns zero, &amp;lt;code&amp;gt;ThinkRequired&amp;lt;/code&amp;gt; does not affect the score. &amp;lt;code&amp;gt;RoslynEvaluatorTests&amp;lt;/code&amp;gt; covers the markdown fence stripping and three evaluation cases: correct implementation, wrong result, and garbage input. &amp;lt;code&amp;gt;ScorerTests&amp;lt;/code&amp;gt; uses a &amp;lt;code&amp;gt;MockRunner&amp;lt;/code&amp;gt; that replays canned responses and verifies that the Scorer assembles the &amp;lt;code&amp;gt;ModelResult&amp;lt;/code&amp;gt; correctly for the pass case, the tool-rejected case, and the instruction-failure case. &amp;lt;code&amp;gt;ConsoleReporterTests&amp;lt;/code&amp;gt; confirms that &amp;lt;code&amp;gt;PrintTable&amp;lt;/code&amp;gt; does not throw with null prior results or when a model has regressed since the previous run.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;None of these tests require a running Ollama instance. The mock runner pattern makes the Scorer fully testable without any external dependencies.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;first-complete-run&amp;#34;&amp;gt;First Complete Run&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Thirteen models, ten metrics each. This is what came back:&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Score&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Notes&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;10/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Clean sweep&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;glm-4.7-flash&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:26b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;devstral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3-coder:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;mistral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gpt-oss:20b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;think_required detected&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;phi4:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;llava:7b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5-coder:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;4/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3/10&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;gemma4:latest — a ~12B parameter model — scores 10 out of 10. It answers the chat question, writes &amp;lt;code&amp;gt;SumEvens&amp;lt;/code&amp;gt; correctly, emits a proper tool call, follows the three-word instruction, builds the Hello World app, writes tests that compile and pass, gets the math problem right, produces valid JSON, and outputs the sequence with no extra text. On every metric the harness defines, it is the best model in the pool by a clean margin over everything larger than it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The result is worth sitting with. A model less than half the size of qwen3:30b outscores it by seven points. glm-4.7-flash is a 30B MoE and comes in second at 9/10. The coding-focused variants — qwen2.5-coder and qwen3-coder — score lower than their general-purpose counterparts at similar sizes.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The obvious interpretation is that gemma4:latest is simply the best model here. The problem is that the harness measures what I built the harness to measure. Before drawing that conclusion, I need to know whether these metrics are the right metrics.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;p&amp;gt;The full source is at &amp;lt;a href=&amp;#34;https://github.com/erichexter/ollama-model-bench&amp;#34;&amp;gt;github.com/erichexter/ollama-model-bench&amp;lt;/a&amp;gt;.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Next up: &amp;lt;a href=&amp;#34;/erichexter/2026/05/31/local-llm-bench-part-3-single-shot-lies/&amp;#34;&amp;gt;Part 3&amp;lt;/a&amp;gt; digs into what the scores actually mean — and why gemma4:latest&amp;amp;rsquo;s clean sweep turned out to be almost entirely beside the point.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
    <entry>
      <title type="html">Search — The Evolution of the Karpathy LLM Wiki</title>
      <link href="https://lostechies.com/erichexter/2026/05/26/search-evolution-of-the-karpathy-llm-wiki/" rel="alternate" type="text/html" title="Search — The Evolution of the Karpathy LLM Wiki" />
      <published>2026-05-26T12:00:00Z</published>
      <updated>2026-05-26T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/05/26/search-evolution-of-the-karpathy-llm-wiki/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/05/26/search-evolution-of-the-karpathy-llm-wiki/">&amp;lt;p&amp;gt;My LLM notes wiki outgrew file reads. Agents were pulling entire files to find a single relevant section — burning tokens on context that didn&amp;amp;rsquo;t matter, missing things that were buried three pages deep. The corpus had just grown past the point where IO-based access was practical.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The fix was search. And since agents need tools, the obvious move was to build it as an MCP server. But if you&amp;amp;rsquo;re building search anyway, plain keyword matching felt like leaving half the value on the table — too easy to miss conceptual matches that don&amp;amp;rsquo;t share exact terms. So: something old and something new. SQLite already has FTS5. sqlite-vec adds HNSW vector search as a loadable extension. Ollama runs the embedding model locally. Put them together and you get hybrid RAG on hardware you already own, exposed as an MCP tool any agent in the fleet can call.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;This post covers how it&amp;amp;rsquo;s built — starting from what the agent sees and working inward to the SQL and vector embeddings.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;what-the-agent-sees&amp;#34;&amp;gt;What the Agent Sees&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;From the agent&amp;amp;rsquo;s perspective, this is just an MCP server with a set of tools. Point an &amp;lt;code&amp;gt;.mcp.json&amp;lt;/code&amp;gt; at the host and the tools are available. No setup, no SDK, no awareness of what&amp;amp;rsquo;s running underneath.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The primary tool is &amp;lt;code&amp;gt;search_knowledge&amp;lt;/code&amp;gt;:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-json&amp;#34; data-lang=&amp;#34;json&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;{&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;  &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;method&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;tools/call&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;  &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;params&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;{&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;    &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;name&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;search_knowledge&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;    &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;arguments&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;{&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;      &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;query&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;attention mechanism scaled dot product&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;      &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;top_k&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;mi&amp;#34;&amp;gt;5&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;      &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;hybrid_alpha&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;mf&amp;#34;&amp;gt;0.6&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;      &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;sources&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;[&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;karpathy-wiki&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;]&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;    &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;}&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;  &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;}&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;}&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;The response comes back as ranked chunks with source context:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-json&amp;#34; data-lang=&amp;#34;json&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;{&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;  &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;content&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;[{&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;    &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;type&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;text&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;    &amp;lt;span class=&amp;#34;nt&amp;#34;&amp;gt;&amp;amp;#34;text&amp;amp;#34;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;&amp;amp;#34;[
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;      {
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;        \&amp;amp;#34;text\&amp;amp;#34;: \&amp;amp;#34;Scaled dot-product attention divides the dot products by √d_k to prevent vanishing gradients in high dimensions...\&amp;amp;#34;,
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;        \&amp;amp;#34;source\&amp;amp;#34;: \&amp;amp;#34;karpathy-wiki\&amp;amp;#34;,
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;        \&amp;amp;#34;relPath\&amp;amp;#34;: \&amp;amp;#34;transformers/attention.md\&amp;amp;#34;,
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;        \&amp;amp;#34;score\&amp;amp;#34;: 0.91,
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;        \&amp;amp;#34;frontmatter\&amp;amp;#34;: { \&amp;amp;#34;tags\&amp;amp;#34;: [\&amp;amp;#34;attention\&amp;amp;#34;, \&amp;amp;#34;transformers\&amp;amp;#34;] }
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;      },
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;      ...
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;s2&amp;#34;&amp;gt;    ]&amp;amp;#34;&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;  &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;}]&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;}&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;The agent gets ranked text chunks, source file paths, and scores. It doesn&amp;amp;rsquo;t need to know whether the result came from a vector search or keyword search — that&amp;amp;rsquo;s the server&amp;amp;rsquo;s problem.&amp;lt;/p&amp;gt;
&amp;lt;h3 id=&amp;#34;the-full-tool-set&amp;#34;&amp;gt;The Full Tool Set&amp;lt;/h3&amp;gt;
&amp;lt;p&amp;gt;Seven tools in total. &amp;lt;code&amp;gt;search_knowledge&amp;lt;/code&amp;gt; covers 95% of use.&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Tool&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Purpose&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;search_knowledge&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Hybrid vec&#43;FTS search across one or more sources.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;get_page&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Retrieve a full page by source &#43; relative path. Use when search returns a partial chunk and you want the full document.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;list_sources&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Lists indexed sources with page/chunk counts and last-indexed timestamps.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;get_stats&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Query counts and latencies over 1h / 24h / 7d / 30d windows.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;get_query_log&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Recent query history. Useful for understanding what agents are actually asking.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;refresh_ingest&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Trigger immediate re-indexing for a source after a write.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;code&amp;gt;ping&amp;lt;/code&amp;gt;&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Returns current UTC. Health check.&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;list_sources&amp;lt;/code&amp;gt; is underrated as a diagnostic. A 200 response from the API tells you nothing about whether the index is populated. If results are poor, check &amp;lt;code&amp;gt;pageCount &amp;amp;gt; 0&amp;lt;/code&amp;gt; and that &amp;lt;code&amp;gt;lastIndexed&amp;lt;/code&amp;gt; is recent before assuming the search logic is wrong.&amp;lt;/p&amp;gt;
&amp;lt;h3 id=&amp;#34;the-hybrid_alpha-parameter&amp;#34;&amp;gt;The &amp;lt;code&amp;gt;hybrid_alpha&amp;lt;/code&amp;gt; Parameter&amp;lt;/h3&amp;gt;
&amp;lt;p&amp;gt;This is the control knob for the blend between vector search and full-text search.&amp;lt;/p&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;0.0&amp;lt;/code&amp;gt; — pure FTS (BM25 keyword ranking)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;1.0&amp;lt;/code&amp;gt; — pure vector (semantic similarity)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;0.5&amp;lt;/code&amp;gt; — equal blend (default)&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;p&amp;gt;In practice, &amp;lt;code&amp;gt;0.6&amp;lt;/code&amp;gt;–&amp;lt;code&amp;gt;0.7&amp;lt;/code&amp;gt; (vector-weighted) works better for conceptual queries: &amp;amp;ldquo;how does attention scale with sequence length.&amp;amp;rdquo; Drop toward &amp;lt;code&amp;gt;0.3&amp;lt;/code&amp;gt; when you need an exact term match that the embedding model might paraphrase: specific function names, error codes, version numbers.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;how-the-search-works&amp;#34;&amp;gt;How the Search Works&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;When &amp;lt;code&amp;gt;search_knowledge&amp;lt;/code&amp;gt; is called, the server runs two queries in parallel and merges the results.&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-csharp&amp;#34; data-lang=&amp;#34;csharp&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;kt&amp;#34;&amp;gt;var&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;vectorTask&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;SearchByVector&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;embeddingVector&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;topK&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;*&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;2&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;sources&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;kt&amp;#34;&amp;gt;var&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;ftsTask&amp;lt;/span&amp;gt;    &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;SearchByFts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;query&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;topK&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;*&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;m&amp;#34;&amp;gt;2&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;sources&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;await&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;Task&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;WhenAll&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;vectorTask&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;ftsTask&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;kt&amp;#34;&amp;gt;var&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;merged&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;Merge&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;vectorTask&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;Result&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;ftsTask&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;Result&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;hybridAlpha&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt; &amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;topK&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;The merge step normalizes each result list&amp;amp;rsquo;s scores to &amp;lt;code&amp;gt;[0, 1]&amp;lt;/code&amp;gt;, applies the alpha weight, sums scores per chunk (a chunk can appear in both lists), and returns the top K. Normalization matters — BM25 and HNSW distance are on completely different scales. Skip it and one path dominates every query regardless of alpha.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Before either query runs, the search query itself gets embedded:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-http&amp;#34; data-lang=&amp;#34;http&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;POST http://&amp;amp;lt;ollama-host&amp;amp;gt;:11434/api/embeddings
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;Content-Type: application/json
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;{
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;  &amp;amp;#34;model&amp;amp;#34;: &amp;amp;#34;nomic-embed-text:latest&amp;amp;#34;,
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;  &amp;amp;#34;prompt&amp;amp;#34;: &amp;amp;#34;attention mechanism scaled dot product&amp;amp;#34;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;err&amp;#34;&amp;gt;}
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;That gives back a 768-dimensional float vector — what the vector search runs against.&amp;lt;/p&amp;gt;
&amp;lt;h3 id=&amp;#34;the-vector-query&amp;#34;&amp;gt;The Vector Query&amp;lt;/h3&amp;gt;
&amp;lt;p&amp;gt;sqlite-vec exposes vector search through a virtual table with a &amp;lt;code&amp;gt;MATCH&amp;lt;/code&amp;gt; clause. Under the hood it&amp;amp;rsquo;s doing an approximate nearest-neighbor scan via HNSW:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-sql&amp;#34; data-lang=&amp;#34;sql&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;SELECT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;body&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rel_path&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;frontmatter&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;       &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;distance&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;FROM&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_vecs&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;JOIN&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunks&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;ON&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;WHERE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;embedding&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;MATCH&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;embedding&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;AND&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;k&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;k&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;AND&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;sources&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;IS&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NULL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;OR&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;IN&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;sources&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;)&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;ORDER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;BY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;cv&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;distance&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;distance&amp;lt;/code&amp;gt; here is L2 distance — lower is closer. sqlite-vec handles all the index internals; from the query side it looks like a regular SQL query.&amp;lt;/p&amp;gt;
&amp;lt;h3 id=&amp;#34;the-fts-query&amp;#34;&amp;gt;The FTS Query&amp;lt;/h3&amp;gt;
&amp;lt;p&amp;gt;Standard SQLite FTS5 with BM25 ranking:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-sql&amp;#34; data-lang=&amp;#34;sql&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;SELECT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;body&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rel_path&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;frontmatter&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;       &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;bm25&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;)&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;AS&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;fts_score&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;FROM&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;JOIN&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunks&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;ON&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rowid&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;WHERE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;MATCH&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;query&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;ORDER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;BY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;bm25&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;)&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;LIMIT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;:&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;k&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;FTS5&amp;amp;rsquo;s &amp;lt;code&amp;gt;MATCH&amp;lt;/code&amp;gt; supports phrase queries, prefix matching, and boolean operators. For agent queries coming in as natural language, the server sanitizes the input to a simple term query before passing it to MATCH.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;the-data-model&amp;#34;&amp;gt;The Data Model&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Three tables carry the retrieval workload:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-sql&amp;#34; data-lang=&amp;#34;sql&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;c1&amp;#34;&amp;gt;-- Chunked text with metadata
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;CREATE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;TABLE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunks&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;          &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;INTEGER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;PRIMARY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;KEY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;page_id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;     &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;INTEGER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NOT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NULL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;REFERENCES&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;pages&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;),&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_index&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;INTEGER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NOT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NULL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;body&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;        &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;TEXT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NOT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NULL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;token_count&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;INTEGER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;      &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;TEXT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rel_path&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;TEXT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;frontmatter&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;TEXT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;c1&amp;#34;&amp;gt;-- Vector index (sqlite-vec extension)
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;CREATE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;VIRTUAL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;TABLE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_vecs&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;USING&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;vec0&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;INTEGER&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;PRIMARY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;KEY&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;embedding&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;nb&amp;#34;&amp;gt;FLOAT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;[&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;mi&amp;#34;&amp;gt;768&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;]&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;c1&amp;#34;&amp;gt;-- Full-text search index (FTS5, built into SQLite)
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;CREATE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;VIRTUAL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;TABLE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunk_fts&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;USING&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;fts5&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;body&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;UNINDEXED&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rel_path&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;  &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;UNINDEXED&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;content&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;s1&amp;#34;&amp;gt;&amp;amp;#39;chunks&amp;amp;#39;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;    &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;content_rowid&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;s1&amp;#34;&amp;gt;&amp;amp;#39;id&amp;amp;#39;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;);&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;chunk_vecs&amp;lt;/code&amp;gt; is a &amp;lt;a href=&amp;#34;https://github.com/asg017/sqlite-vec&amp;#34;&amp;gt;sqlite-vec&amp;lt;/a&amp;gt; &amp;lt;code&amp;gt;vec0&amp;lt;/code&amp;gt; virtual table — INSERT a row with the chunk ID and its 768-dim embedding, sqlite-vec maintains the HNSW index internally. &amp;lt;code&amp;gt;chunk_fts&amp;lt;/code&amp;gt; is a content-backed FTS5 table that stays in sync with &amp;lt;code&amp;gt;chunks&amp;lt;/code&amp;gt; via triggers.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Supporting tables: &amp;lt;code&amp;gt;pages&amp;lt;/code&amp;gt; (source files with hash-based change detection), &amp;lt;code&amp;gt;indexer_runs&amp;lt;/code&amp;gt; (ingest audit log), &amp;lt;code&amp;gt;query_log&amp;lt;/code&amp;gt; (query history for observability).&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;One SQLite file. No separate processes, no network hops between storage components, no backup complexity.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;the-write-path&amp;#34;&amp;gt;The Write Path&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;When a document is added or updated in the source directory, the indexer picks it up:&amp;lt;/p&amp;gt;
&amp;lt;ol&amp;gt;
&amp;lt;li&amp;gt;SHA-256 hash the file. Compare against &amp;lt;code&amp;gt;pages.content_hash&amp;lt;/code&amp;gt;. Skip if unchanged.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Parse YAML frontmatter. Extract the body.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Split into chunks — 512-token target, 64-token overlap, break on paragraph boundaries where possible.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;For each chunk: POST to Ollama &amp;lt;code&amp;gt;/api/embeddings&amp;lt;/code&amp;gt;. Receive a 768-dim float array.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;INSERT into &amp;lt;code&amp;gt;chunks&amp;lt;/code&amp;gt;. INSERT into &amp;lt;code&amp;gt;chunk_vecs&amp;lt;/code&amp;gt;. FTS5 trigger handles &amp;lt;code&amp;gt;chunk_fts&amp;lt;/code&amp;gt; sync.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Update &amp;lt;code&amp;gt;pages.content_hash&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;indexed_at&amp;lt;/code&amp;gt;.&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;Write a row to &amp;lt;code&amp;gt;indexer_runs&amp;lt;/code&amp;gt;.&amp;lt;/li&amp;gt;
&amp;lt;/ol&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;nomic-embed-text&amp;lt;/code&amp;gt; is 137M parameters — fast on a GPU host, single-digit milliseconds per chunk. The indexer pipelines requests; Ollama queues them.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;gotchas&amp;#34;&amp;gt;Gotchas&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;The embed model context limit is a silent failure.&amp;lt;/strong&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;nomic-embed-text&amp;lt;/code&amp;gt; has an 8K token context window. Chunks that exceed it are silently not embedded — present in &amp;lt;code&amp;gt;chunks&amp;lt;/code&amp;gt;, retrievable via &amp;lt;code&amp;gt;get_page&amp;lt;/code&amp;gt;, invisible to vector search. No error from Ollama. Enforce the chunk size limit at ingest time. Symptom check:&amp;lt;/p&amp;gt;
&amp;lt;div class=&amp;#34;highlight&amp;#34;&amp;gt;&amp;lt;pre tabindex=&amp;#34;0&amp;#34; class=&amp;#34;chroma&amp;#34;&amp;gt;&amp;lt;code class=&amp;#34;language-sql&amp;#34; data-lang=&amp;#34;sql&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;SELECT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;rel_path&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;source&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;,&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;LENGTH&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;(&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;content&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;)&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;AS&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;content_len&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;FROM&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;pages&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;LEFT&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;JOIN&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;chunks&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;ON&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;page_id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;o&amp;#34;&amp;gt;=&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;p&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;line&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;cl&amp;#34;&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;WHERE&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;c&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;.&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;n&amp;#34;&amp;gt;id&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;IS&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;k&amp;#34;&amp;gt;NULL&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;p&amp;#34;&amp;gt;;&amp;lt;/span&amp;gt;&amp;lt;span class=&amp;#34;w&amp;#34;&amp;gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/code&amp;gt;&amp;lt;/pre&amp;gt;&amp;lt;/div&amp;gt;&amp;lt;p&amp;gt;Any row here is a page with no chunks.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;Stale bind mount after remount.&amp;lt;/strong&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;If the CIFS mount backing the source directory remounts — after a network blip or server reboot — the container holds a file descriptor to the old empty mount point. The API returns 200. The indexer runs. It finds zero files. Nothing crashes, nothing complains. Restart the container after any storage remount.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;strong&amp;gt;Shallow health checks miss the real failure mode.&amp;lt;/strong&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;code&amp;gt;GET /ping → 200&amp;lt;/code&amp;gt; stays green with an empty index. Real health check: call &amp;lt;code&amp;gt;list_sources&amp;lt;/code&amp;gt;, assert &amp;lt;code&amp;gt;pageCount &amp;amp;gt; 0&amp;lt;/code&amp;gt; with a recent &amp;lt;code&amp;gt;lastIndexed&amp;lt;/code&amp;gt;. You&amp;amp;rsquo;re monitoring the retrieval system, not just the process.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;h2 id=&amp;#34;what-this-gets-you&amp;#34;&amp;gt;What This Gets You&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;~11K chunks, query results under 100ms on commodity hardware. The Ollama embedding call is the only network hop on the hot path — ~10ms on a GPU host for a short query. The SQLite ANN index is not the bottleneck.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Hybrid search earns its keep in practice. Pure vector drifts on exact version numbers, function names, and error codes. Pure FTS misses conceptual synonyms. The blend handles both without tuning a separate retriever per query type.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The MCP wrapper means any agent that speaks the protocol can call it without any awareness of the storage layer. Add a source, re-index, done — consumers don&amp;amp;rsquo;t change.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Most databases can store embeddings at this point. The reason to reach for SQLite &#43; sqlite-vec specifically is that you probably already have it, it requires no new infrastructure, and the FTS5 index is already there. The hybrid approach — run both searches, blend by alpha — transfers to any store that can handle both. The schema and the search logic are the portable parts.&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
    <entry>
      <title type="html">Which Local Models Can Actually Code?</title>
      <link href="https://lostechies.com/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/" rel="alternate" type="text/html" title="Which Local Models Can Actually Code?" />
      <published>2026-05-25T12:00:00Z</published>
      <updated>2026-05-25T12:00:00Z</updated>
      <id>https://lostechies.com/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/</id>
      <content type="html" xml:base="https://lostechies.com/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/">&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Part 1 of 5 in the &amp;lt;a href=&amp;#34;/erichexter/2026/05/25/local-llm-bench-part-1-which-models-can-chat/&amp;#34;&amp;gt;Local LLM Bench series&amp;lt;/a&amp;gt;.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;I had ten local models installed and no good answer to a simple question: which of them could actually do useful work? Chat demos are easy to fake. I wanted to know whether these models could write working code, call tools correctly, and follow instructions without needing hand-holding. The only way to find out was to run them.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-setup&amp;#34;&amp;gt;The Setup&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Machine is an Alienware Windows 11 box with an RTX 5080 carrying 16GB of VRAM. Ollama is running locally, serving the following ten models:&amp;lt;/p&amp;gt;
&amp;lt;ul&amp;gt;
&amp;lt;li&amp;gt;mistral:latest (7B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;llava:7b (7B, vision)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;gemma4:latest (~12B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;gemma4:26b (26B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;qwen3:14b (14B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;qwen3:30b (30B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;phi4:14b (14B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;qwen2.5:14b (14B)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;qwen2.5-coder:14b (14B, coding-focused)&amp;lt;/li&amp;gt;
&amp;lt;li&amp;gt;glm-4.7-flash (30B MoE)&amp;lt;/li&amp;gt;
&amp;lt;/ul&amp;gt;
&amp;lt;p&amp;gt;The size range alone tells you the hardware story. Anything under about 20B fits in VRAM comfortably. The 26B and 30B models spill onto system RAM — which you feel in the latency numbers.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;first-pass-two-prompts-powershell&amp;#34;&amp;gt;First Pass: Two Prompts, PowerShell&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The first script was about as minimal as it gets. Two prompts per model: &amp;amp;ldquo;What is the capital of France?&amp;amp;rdquo; to confirm the model is responding at all, and &amp;amp;ldquo;Write an &amp;lt;code&amp;gt;is_prime()&amp;lt;/code&amp;gt; function in Python&amp;amp;rdquo; as a basic code generation check. No scoring, no verification — just checking that something came back.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Most models answered both prompts without incident. Then I hit the bigger ones. gemma4:26b, glm-4.7-flash, and qwen3:30b all returned empty responses. Not errors — the HTTP calls succeeded, Ollama said everything was fine, the responses just contained no text.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;That took longer than it should have, and the answer was different for each model.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-think-mode-wall&amp;#34;&amp;gt;The Think-Mode Wall&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;qwen3 models support a reasoning mode where the model works through a problem step by step before producing visible output. The reasoning tokens live inside &amp;lt;code&amp;gt;&amp;amp;lt;think&amp;amp;gt;...&amp;amp;lt;/think&amp;amp;gt;&amp;lt;/code&amp;gt; blocks and don&amp;amp;rsquo;t count against the response. What does count against the response is the token budget, and when I was requesting with a tight &amp;lt;code&amp;gt;num_predict&amp;lt;/code&amp;gt; limit, the model was spending the entire budget on internal reasoning and returning nothing to the caller. glm-4.7-flash has its own variant of the same mode — different model family, same symptom.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The fix for both: add &amp;lt;code&amp;gt;&amp;amp;quot;think&amp;amp;quot;: false&amp;lt;/code&amp;gt; to the request body. With that flag set, qwen3:14b went from returning a blank response to producing clean, working code in about 2 seconds. The qwen3 and glm models followed.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;gemma4:26b&amp;amp;rsquo;s blank responses were a separate problem entirely. At 26B it spills to RAM, and with a tight &amp;lt;code&amp;gt;num_predict&amp;lt;/code&amp;gt; budget and slow generation speed, the script&amp;amp;rsquo;s read timeout was firing before any tokens arrived. More headroom fixed it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The lesson here is that &amp;amp;ldquo;model returned empty string&amp;amp;rdquo; and &amp;amp;ldquo;model failed&amp;amp;rdquo; are not the same thing, and you have to understand what each model family expects before you can interpret the output.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;tool-calling-where-things-got-interesting&amp;#34;&amp;gt;Tool-Calling: Where Things Got Interesting&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Once the basic chat and code tests were passing, I added a tool-calling test. The prompt was &amp;amp;ldquo;What&amp;amp;rsquo;s the weather in Paris?&amp;amp;rdquo; with a &amp;lt;code&amp;gt;get_weather&amp;lt;/code&amp;gt; function schema attached to the request. A model that handles tool calling correctly should stop generating text and instead emit a structured &amp;lt;code&amp;gt;tool_calls&amp;lt;/code&amp;gt; object pointing at &amp;lt;code&amp;gt;get_weather&amp;lt;/code&amp;gt; with the right argument. A model that doesn&amp;amp;rsquo;t understand the protocol either returns prose (&amp;amp;ldquo;I don&amp;amp;rsquo;t have access to weather data&amp;amp;rdquo;), returns a JSON blob as plain text, or refuses the request entirely with an HTTP 400.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The results split into three clear buckets. mistral, gemma4 (both sizes), qwen3:14b, qwen2.5:14b, and glm-4.7-flash all produced proper structured &amp;lt;code&amp;gt;tool_calls&amp;lt;/code&amp;gt;. That is the expected behavior — the model uses the tool schema as intended.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;qwen2.5-coder:14b was the interesting failure. It returned what looked like a tool call, but as a raw JSON string embedded in the message content rather than as a structured &amp;lt;code&amp;gt;tool_calls&amp;lt;/code&amp;gt; entry. The model clearly understood what was being asked; it just didn&amp;amp;rsquo;t output it in the right format. A &amp;amp;ldquo;coder&amp;amp;rdquo; model is not necessarily a &amp;amp;ldquo;tool-aware&amp;amp;rdquo; model. They are different capabilities.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;llava:7b and phi4:14b both returned HTTP 400 on any request that included the &amp;lt;code&amp;gt;tools&amp;lt;/code&amp;gt; field. Those models simply do not accept the parameter — the API rejects it before the model even sees the prompt. llava makes sense here: it is a vision model, not a chat/agent model. phi4 is less obvious.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;mid-phase-additions&amp;#34;&amp;gt;Mid-Phase Additions&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;While working through these tests I pulled in three more models that had come up in research as strong candidates for coding benchmarks: devstral:latest (22B, Devstral Small — Mistral&amp;amp;rsquo;s coding-focused release), qwen3-coder:30b (~30B, Qwen&amp;amp;rsquo;s coding-tuned variant), and gpt-oss:20b (~20B). All three were added before the formal scoring phase started.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;the-baseline-table&amp;#34;&amp;gt;The Baseline Table&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;Here is where every model stood after the initial phase — response times are wall-clock from the PowerShell script, rounded to the nearest second:&amp;lt;/p&amp;gt;
&amp;lt;table&amp;gt;
	&amp;lt;thead&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;th&amp;gt;Model&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Size&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Chat&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Code&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Tool call&amp;lt;/th&amp;gt;
					&amp;lt;th&amp;gt;Notes&amp;lt;/th&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/thead&amp;gt;
	&amp;lt;tbody&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;mistral:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;llava:7b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;7B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;4s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;amp;lt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;rejected&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Vision model&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:latest&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;~12B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;4s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;think=false required&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;phi4:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;5s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;rejected&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen2.5-coder:14b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;6s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;1s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;text (not structured)&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;amp;ldquo;coder&amp;amp;rdquo; does not mean tool-aware&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;gemma4:26b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;26B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;9s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;3s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Partial CPU offload&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;glm-4.7-flash&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;30B MoE&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;4s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
			&amp;lt;tr&amp;gt;
					&amp;lt;td&amp;gt;qwen3:30b&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;30B&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;14s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;8s&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;proper&amp;lt;/td&amp;gt;
					&amp;lt;td&amp;gt;Slowest in pool&amp;lt;/td&amp;gt;
			&amp;lt;/tr&amp;gt;
	&amp;lt;/tbody&amp;gt;
&amp;lt;/table&amp;gt;
&amp;lt;p&amp;gt;The latency numbers tell one story — size matters, mostly predictably. The tool-call column tells another: ten models, three different behaviors from the same input, and two of them would silently fail in any agentic loop that expected structured output.&amp;lt;/p&amp;gt;
&amp;lt;h2 id=&amp;#34;what-works-actually-means&amp;#34;&amp;gt;What &amp;amp;ldquo;Works&amp;amp;rdquo; Actually Means&amp;lt;/h2&amp;gt;
&amp;lt;p&amp;gt;The issue with this baseline is that &amp;amp;ldquo;passes&amp;amp;rdquo; hides a lot. A model that returns a tool call in the message content instead of the &amp;lt;code&amp;gt;tool_calls&amp;lt;/code&amp;gt; field looks fine until your application tries to deserialize the response. A model that works at &amp;lt;code&amp;gt;num_predict=300&amp;lt;/code&amp;gt; might silently truncate at &amp;lt;code&amp;gt;num_predict=100&amp;lt;/code&amp;gt;. A model that answers &amp;amp;ldquo;capital of France&amp;amp;rdquo; correctly might write Python &amp;lt;code&amp;gt;is_prime()&amp;lt;/code&amp;gt; that has an off-by-one error nobody noticed because nobody ran it.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;Everything in this phase was manual inspection. I was reading outputs and deciding they looked reasonable. That is not a test; that is a vibe check.&amp;lt;/p&amp;gt;
&amp;lt;p&amp;gt;The only way to actually know whether a model can write working code is to compile and run the code. Which meant building something more serious.&amp;lt;/p&amp;gt;
&amp;lt;hr&amp;gt;
&amp;lt;p&amp;gt;&amp;lt;em&amp;gt;Next up: &amp;lt;a href=&amp;#34;/erichexter/2026/05/28/local-llm-bench-part-2-building-the-harness/&amp;#34;&amp;gt;Part 2&amp;lt;/a&amp;gt; covers building the .NET 10 benchmark harness — including a scoring system that actually executes model-generated C# and runs the tests.&amp;lt;/em&amp;gt;&amp;lt;/p&amp;gt;</content>

      
      <author>
          <name>Eric Hexter</name>
      </author>
      

      

      

      
    </entry>
  
</feed>
