<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>S Anand</title>
    <link>https://www.s-anand.net/blog/</link>
    <description>Recent content on S Anand</description>
    <generator>Hugo -- 0.165.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 27 Sep 2026 21:30:30 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Qwen 3.6 vs Gemma 4 vs Luna</title>
      <link>https://www.s-anand.net/blog/qwen-3-6-vs-gemma-4-vs-luna/</link>
      <pubDate>Sun, 27 Sep 2026 21:30:30 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/qwen-3-6-vs-gemma-4-vs-luna/</guid>
      <description>&lt;p&gt;Open weights models are nudging up the frontier.&lt;/p&gt;
&lt;p&gt;For example, &lt;a href=&#34;https://openrouter.ai/xiaomi/mimo-v2.6-pro&#34;&gt;MiMo V2.6 Pro&lt;/a&gt; is an outlier on the &lt;a href=&#34;https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task#price-cost&#34;&gt;Artificial Analysis Intelligence vs Cost per Task benchmark&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task#price-cost&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-27-artificial-analysis-intelligence-vs-cost-per-task.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://openrouter.ai/z-ai/glm-5.3-flash&#34;&gt;GLM 5.3 Flash&lt;/a&gt; is an outlier on the &lt;a href=&#34;https://arena.ai/leaderboard/text/overall/pareto&#34;&gt;Arena Text Pareto&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://arena.ai/leaderboard/text/overall/pareto&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-27-arena-text-pareto.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://openrouter.ai/deepseek/deepseek-v4.1-flash&#34;&gt;Deepseek V4.1 Flash&lt;/a&gt; seems to be doing a great job as well.&lt;/p&gt;
&lt;p&gt;So, I thought I&amp;rsquo;d relook which model to use locally for coding.&lt;/p&gt;
&lt;p&gt;BTW, I don&amp;rsquo;t use local models for coding. It&amp;rsquo;s pointless, except on flights with power sockets. Partly in preparation for flights, and partly to check if I&amp;rsquo;m missing something, I benchmarked two models I could run locally on my 8 GB RTX 2000 GPU: Gemma 4 E4B and Qwen 3.6 against GPT 6 Luna. (Better models like Qwen 3.8, MiMo V2.6 Pro, GLM 5.3 Flash, Deepseek V4.1 Flash, etc. are too big for my GPU.)&lt;/p&gt;
&lt;p&gt;I asked ChatGPT to look at my recent coding work and suggest 3 prompts. It did:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/tasks/flare-loading/prompt.md&#34;&gt;flare-loading&lt;/a&gt;: Make a large image-heavy page load better on a poor network without reducing image quality.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/tasks/mcpserver-console/prompt.md&#34;&gt;mcpserver-console&lt;/a&gt;: Make rapidly scrolling MCP tool logs easier to understand without&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/tasks/whatsapp-integrity/prompt.md&#34;&gt;whatsapp-integrity&lt;/a&gt;: Detect corrupt/misattributed backup JSONL safely before touching WhatsApp/CDP.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I gave the task to all three models. &lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/README.md&#34;&gt;Here&amp;rsquo;s what happened&lt;/a&gt;:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Task&lt;/th&gt;
					&lt;th&gt;Gemma 4 E4B&lt;/th&gt;
					&lt;th&gt;Qwen 3.6&lt;/th&gt;
					&lt;th&gt;GPT-6 Luna&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;flare-loading&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/gemma/flare-loading/diff.patch&#34;&gt;🔴 Didn&amp;rsquo;t fix&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/qwen/flare-loading/diff.patch&#34;&gt;🟡 Preloaded but didn&amp;rsquo;t optimize&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/codex/flare-loading/diff.patch&#34;&gt;🟢 Optimized&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;mcpserver-console&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/gemma/mcppserver-console/diff.patch&#34;&gt;🔴 Didn&amp;rsquo;t complete&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/qwen/mcppserver-console/diff.patch&#34;&gt;🟡 Improved readability but timed out&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/codex/mcppserver-console/diff.patch&#34;&gt;🟢 Concise solution&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;whatsapp-integrity&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/gemma/whatsapp-integrity/diff.patch&#34;&gt;🔴 Didn&amp;rsquo;t complete&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/qwen/whatsapp-integrity/diff.patch&#34;&gt;🟡 Built a solution but timed out&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://github.com/sanand0/llmevals/blob/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b/results/codex/whatsapp-integrity/diff.patch&#34;&gt;🟢 Concise solution&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In short, there&amp;rsquo;s still no reason I&amp;rsquo;d use a local model for coding. But if I &lt;em&gt;did&lt;/em&gt; have to, I&amp;rsquo;d use Qwen 3.6 today.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The process of benchmarking is becoming increasingly easy. Here&amp;rsquo;s roughly how I created &lt;a href=&#34;https://github.com/sanand0/llmevals/tree/d906e9a47aa6110c50c6706e99796543e82eb4c5/qwen-3.6-vs-gemma4-e4b&#34;&gt;this benchmark&lt;/a&gt; entirely using ChatGPT in about half a dozen prompts:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I&amp;rsquo;ve been using gemma4:e4b-it-qat and it&amp;rsquo;s working fine. Several new models have been released since. What would you now recommend as the best model I can run on my laptop using ollama?&lt;/li&gt;
&lt;li&gt;What&amp;rsquo;s the easiest way for me to set up and use Qwen 3.6 with llama.cpp and use it via pi agent?&lt;/li&gt;
&lt;li&gt;Suggest 3 tasks that I can use to benchmark or compare the quality of these models. Something where I can instantly / easily judge the output and is likely to produce a difference between these models (and GPT 6 Luna - which I&amp;rsquo;ll add later into the benchmarks.) You&amp;rsquo;re welcome to iterate a few times to see if it really provides an output that&amp;rsquo;s well differentiated and easy for me to see the difference.&lt;/li&gt;
&lt;li&gt;OK. Create, under &lt;code&gt;~/code/llmevals/qwen-3.6-vs-gemma4-e4b/&lt;/code&gt; a set of directories for each of these tasks (perhaps one set for qwen and one set for gemma? I&amp;rsquo;ll leave it to you to decide the best organization) as well as a benchmark.sh that I can run which will update / fix the directories. Commit the files. When done, let me know. I will run benchmark.sh. Then you can review the output and compare the quality. I&amp;rsquo;ll do the same. We&amp;rsquo;ll compare notes and decide next steps.&lt;/li&gt;
&lt;li&gt;Now modify benchmark.sh so that it will run for codex.&lt;/li&gt;
&lt;li&gt;Compare the results and let me know what you think. Then I&amp;rsquo;ll share my opinion.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 27 Sep 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-27-sep-2026/</link>
      <pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-27-sep-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/adbar/trafilatura&#34;&gt;trafilatura&lt;/a&gt; is a Python library that extracts the main content as Markdown from a web page. A useful alternative to &lt;a href=&#34;https://jina.ai/reader&#34;&gt;Jina Reader&lt;/a&gt; for text. It&amp;rsquo;s better at main content extraction but can&amp;rsquo;t handle non-HTML / JS generated / bot-protected URLs. &lt;!-- https://chatgpt.com/c/6ab73d10-ea04-83ec-931d-ff7b9aa356bd --&gt;&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://chatgpt.com/settings/plugins-settings/plugin_asdk_app_6a057d268ebc81919918d37eec718425&#34;&gt;Remote Desktop Commander&lt;/a&gt; ChatGPT plugin is a good alternative to my &lt;a href=&#34;https://github.com/sanand0/scripts/blob/539caf05d481ee4e80686b8f83ab362695c06c7e/mcpserver.py&#34;&gt;mcpserver.py&lt;/a&gt;. Both let you expose bash on your laptop to ChatGPT - which is ultra-powerful. Here&amp;rsquo;re the where RDC is 🟢 better and 🔴 worse. I would recommend it to everyone (but I&amp;rsquo;ll stick to my own code).
&lt;ul&gt;
&lt;li&gt;🟢 More features: session/process search, reads PDF/DOCX/XLSX, better file metadata, editing, reading, etc.&lt;/li&gt;
&lt;li&gt;🟢 Easier: Single command to run, no maintenance, multi-device support&lt;/li&gt;
&lt;li&gt;🟡 Not sandboxed: But you can run it inside a Docker instance&lt;/li&gt;
&lt;li&gt;🔴 Hard to tweak: custom instructions, custom logging, etc. require code changes&lt;/li&gt;
&lt;li&gt;🔴 Privacy: Desktop Commander servers see all traffic (which is why I&amp;rsquo;ll stick to my code for now)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From &lt;a href=&#34;https://youtu.be/WSU_wt4HoXc&#34;&gt;Arun&amp;rsquo;s lecture to IHRD, Kerala, Jan 2026&lt;/a&gt;, here&amp;rsquo;s what I noted as the impact of Gen AI (and Ed Tech, broadly) on students, and how I address this.
&lt;ul&gt;
&lt;li&gt;Defers learning. My approach: teach how to learn on demand.&lt;/li&gt;
&lt;li&gt;Reduces attention spans. I don&amp;rsquo;t yet have an approach for this.&lt;/li&gt;
&lt;li&gt;Reduces understanding - weakens the the &amp;ldquo;mental struggle muscle&amp;rdquo;. My approach: give formerly impossible problems.&lt;/li&gt;
&lt;li&gt;Reduces emotional and social learning. My approach: assess collaborative games.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;With AI making software easier, we can change our operating systems to suit us. Indicators, widgets, keyboard shortcuts, window managers, accessibility tools, device managers, automation workflows, power management, notification management, visual appearance, … I mean, just one look at the Settings in our OS should give us ideas on what&amp;rsquo;s possible and what annoys us.&lt;/li&gt;
&lt;li&gt;I&amp;rsquo;m surprised how little CPU VLC Media Player consumes when playing songs. I used to avoid listening to songs on flights to save power. That seems unnecessary. Most of my VS Code and browser processes consume way more CPU (3-6% of 1 CPU per process, as opposed to VLC&amp;rsquo;s 0.5%)&lt;/li&gt;
&lt;li&gt;Alcoholism is partly genetic &lt;em&gt;and&lt;/em&gt; ancestral. There&amp;rsquo;s an ALDH2 rs671 gene and those who carry it (many East Asians) drink less and are less prone to addiction. &lt;a href=&#34;https://pmc.ncbi.nlm.nih.gov/articles/PMC5568932/&#34;&gt;PubMed&lt;/a&gt;. &lt;a href=&#34;https://pmc.ncbi.nlm.nih.gov/articles/PMC8205229/&#34;&gt;Smoking&lt;/a&gt; and &lt;a href=&#34;https://pmc.ncbi.nlm.nih.gov/articles/PMC3071630/&#34;&gt;Coffee&lt;/a&gt; might have something similar, too. &lt;a href=&#34;https://www.nature.com/articles/s41398-024-02870-7&#34;&gt;Aggression&lt;/a&gt; and &lt;a href=&#34;https://pmc.ncbi.nlm.nih.gov/articles/PMC10962975/&#34;&gt;IQ&lt;/a&gt; seems genetic, but less ancestral. &lt;!-- https://chatgpt.com/c/6ab4b6f3-3694-83ec-8c6f-9f9c28c0c642 --&gt;&lt;/li&gt;
&lt;li&gt;I used ChatGPT&amp;rsquo;s voice mode as a tour guide at Fort Santiago, Manila. It was pretty good - it researched the place, told me what to see, explained what I was seeing (interpreting my photos), laughed at my enthusiasm, and made me feel like I had company. But the experience wasn&amp;rsquo;t perfect (and I expect these will improve - I need to try this more) because it: &lt;!-- https://chatgpt.com/c/6ab5dd4b-cea0-83ec-814d-6501a11b4a3b --&gt;
&lt;ul&gt;
&lt;li&gt;Made two factual errors I spotted. It said &amp;ldquo;down river&amp;rdquo; instead of &amp;ldquo;up river&amp;rdquo; when mentioning a new bridge, said lookout holes were bigger on the inside than the outside. I expect models will get better.&lt;/li&gt;
&lt;li&gt;Responded slower than I&amp;rsquo;d like because it kept researching. I later told it to stop researching and just talk to me. But it was able to talk to me while running tools (including research) in the background, so I expect these are getting better, too.&lt;/li&gt;
&lt;li&gt;Didn&amp;rsquo;t have enough personality. It felt like a helpful assistant I can&amp;rsquo;t make friends with, rather than a stranger with idiosyncracies or preferences. I expect they&amp;rsquo;ll be able to take on more personalities soon (and perhaps already can, if instructed to).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;I find the ChatGPT &amp;ldquo;Library&amp;rdquo; a useful place to store notes while speaking. I just tell it to add an idea to &amp;ldquo;notes/ideas.md&amp;rdquo; in my library and review it periodically. That&amp;rsquo;s a pretty useful way to take notes while talking to it in voice mode.&lt;/li&gt;
&lt;li&gt;On Google, the &amp;ldquo;I&amp;rsquo;m feeling lucky&amp;rdquo; takes you directly to the first result. On Google AI Studio, when you create an app and &lt;strong&gt;dictate&lt;/strong&gt; what you want and press &amp;ldquo;I&amp;rsquo;m feeling lucky&amp;rdquo;, it DOESN&amp;rsquo;T transcribe what you said first. It just builds a random application (often titled &amp;ldquo;MuseInk&amp;rdquo; for me) #ForNow. The geniuses who designed the original &amp;ldquo;I&amp;rsquo;m feeling lucky&amp;rdquo; clearly did more usability testing than the current AI Studio team.&lt;/li&gt;
&lt;li&gt;Several sites have popped up that let agents deploy websites. Here&amp;rsquo;s &lt;a href=&#34;https://chatgpt.com/share/6ab60ffc-1630-83ec-968e-07ff114ac64b&#34;&gt;ChatGPT&amp;rsquo;s review of agent hosting services&lt;/a&gt;: &lt;!-- https://chatgpt.com/c/6ab60e5f-9fec-83ec-8eb4-39f32ec55002 --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;https://here.now&#34;&gt;here.now&lt;/a&gt; by default.&lt;/strong&gt; Smoothest all-round agent publishing, with stable URLs, updates, access control, and versioning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;https://pagedrop.io&#34;&gt;PageDrop&lt;/a&gt; for review.&lt;/strong&gt; Inline comments turn directly into feedback for agent revision.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;https://htmldrop.app&#34;&gt;HTMLDrop&lt;/a&gt; for clean MCP/OAuth integration.&lt;/strong&gt; Best when authentication and remote MCP plumbing matter.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;https://stacktr.ee&#34;&gt;Stacktree&lt;/a&gt; for client deliverables.&lt;/strong&gt; Gated sharing plus feedback and engagement features.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://developers.openai.com/api/docs/models/gpt-live-1&#34;&gt;GPT Live 1&lt;/a&gt; costs 5c/min ($3/hr) flat #ForNow. That&amp;rsquo;s a MUCH easier to use pricing. &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live&#34;&gt;Gemini 3.8 Live&lt;/a&gt; is more complex. Small conversations (under 10 min) might cost just $0.7/ hr but over time, can accumulate context and grow to $2-5/hr #ForNow. &lt;a href=&#34;https://chatgpt.com/share/6ab5f617-9bbc-83ec-bd49-51c21f521b5e&#34;&gt;ChatGPT&lt;/a&gt; &lt;!-- https://chatgpt.com/c/6ab5f3f6-1ebc-83ec-a490-b848e3495b30 --&gt;&lt;/li&gt;
&lt;li&gt;AI overwhelms me and I have &amp;ldquo;LLM fatigue&amp;rdquo; (tired of actioning AI output). If you treat hard tasks like exercise (&amp;ldquo;you&amp;rsquo;re building muscle&amp;rdquo;) you get more done, you feel less miserable, and build an ability (a mental muscle of some kind, I think). I started with a &amp;ldquo;1 min of exercise&amp;rdquo; on 14 Sep, forcing myself to do &lt;em&gt;just one minute&lt;/em&gt; of something (typically actioning AI output), then increased it to 2 min the next day, and so on. 10 min may not sound like much today, but since I tend to stick to routines, I&amp;rsquo;ll be able to focus for an hour extra in a couple of months.&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://study.iitm.ac.in/ds/&#34;&gt;IIT Madras BS in Data Science and Applications&lt;/a&gt; has at least three support business models: &lt;!-- https://chatgpt.com/c/6ab0f05d-acc0-83ec-9f11-4eea902cd54e --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Get in:&lt;/strong&gt; &lt;a href=&#34;https://www.robotixedu.com/IITM/&#34;&gt;Meritus&lt;/a&gt; (Ramana Prasad) coaches for the entrance exam.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Get through:&lt;/strong&gt; &lt;a href=&#34;https://www.datacharya.in/&#34;&gt;DataCharya&lt;/a&gt; and &lt;a href=&#34;https://unknowniitians.com/&#34;&gt;Unknown IITians&lt;/a&gt; coach Foundation/Diploma/Degree courses; &lt;a href=&#34;https://www.acegrade.in/&#34;&gt;AceGrade&lt;/a&gt; (Sumit K. Sharma) offers much of this free.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make it a college:&lt;/strong&gt; &lt;a href=&#34;https://theseep.org/&#34;&gt;SEEP&lt;/a&gt; wraps IITM BS in an offline campus, classroom, cohort and mentoring experience.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/tomstagl/tap/cctop&#34;&gt;&lt;code&gt;cctop&lt;/code&gt;&lt;/a&gt; is a nice CLI alternative to &lt;a href=&#34;https://github.com/kenn-io/agentsview&#34;&gt;&lt;code&gt;agentsview&lt;/code&gt;&lt;/a&gt; for monitoring agent sessions. I still prefer &lt;code&gt;agentsview&lt;/code&gt; for details but &lt;code&gt;cctop&lt;/code&gt; has a real-time update that&amp;rsquo;s fast and useful.&lt;/li&gt;
&lt;li&gt;In June, I predicted &amp;ldquo;Python will have grown the most as a language in GitHub&amp;rdquo; by the end of the year. That&amp;rsquo;s because AI agents know Python well and will likely code in Python. But agents are now just as fluent in Rust, etc. as well as able to debug new languages, so I expect that the better programmers will carefully choose their programming language to the task. In fact, I more far more likely to hire someone with a Rust repo on GitHub and, next year, might treat Python as a slop-smell.&lt;/li&gt;
&lt;li&gt;About 1 million people were discovered in Papua New Guinea in the 1930s who had no contact with most of the outside world - via &lt;a href=&#34;https://notnottalmud.substack.com/p/why-i-cant-stop-thinking-about-papua&#34;&gt;why I can&amp;rsquo;t stop thinking about Papua New Guinea and what I think everyone should know about it&lt;/a&gt;. &amp;ldquo;The Spanish and Portuguese brought sweet potato and tobacco near the western tip of New Guinea, in the 1500s. Beyond that point, there were no merchants. But when a woman married into the next clan over, she took cuttings from her family’s garden with her. Each new family then planted the same, saw that it worked, and passed it on. At that pace the sweet potato crossed the highlands in a century or two, with nobody knowing where it came from beyond the tribe beside them.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://john.hartnup.uk/poster-prompts/&#34;&gt;Poster Prompts&lt;/a&gt; is a gallery of prompts for AI-generated posters. Similar to my &lt;a href=&#34;https://sanand0.github.io/llmartstyle/&#34;&gt;LLM Art Style&lt;/a&gt;. As before, I&amp;rsquo;m struck by how few designs I actually like and would use in practice. &lt;a href=&#34;https://news.ycombinator.com/item?id=49764791&#34;&gt;Hacker News&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;In 1653, &lt;a href=&#34;https://en.wikipedia.org/wiki/Thomas_Urquhart&#34;&gt;Thomas Urquhart&lt;/a&gt; wrote &lt;a href=&#34;https://en.wikipedia.org/wiki/Logopandecteision&#34;&gt;Logopandecteision&lt;/a&gt; - a book in which he plans a new language. People believe it was a parody / practical joke. He also included a cipher:&lt;br&gt;
&lt;a href=&#34;https://scienceblogs.de/klausis-krypto-kolumne/2017/06/30/the-top-50-unsolved-encrypted-messages-28-thomas-urquharts-encrypted-poems/&#34;&gt;&lt;img alt=&#34;5.3.27.38.32.14.21.8.66.8.70.39.5.9.12.18.2.3.56.5.1.7.3.2.13.19.3.25.9.3.16.6.25.15.13.6.11.20.5.1.2.12.1.20.20.49.20.20.35.33.4.6.8.35.5.33.5.5.18.10.3.11.32.42.&#34; loading=&#34;lazy&#34; src=&#34;https://scienceblogs.de/klausis-krypto-kolumne/files/2014/11/Urquhart-Cryptogram.png&#34;&gt;&lt;/a&gt;&lt;br&gt;
&amp;hellip; that &lt;a href=&#34;https://www.vals.ai/blogs/fable-solves-cyphral-distich&#34;&gt;Fable 5.1 solved in 44 minutes and 176k tokens&lt;/a&gt;. It reads: O GOD UPHOLD KING CHARLS THE SECOND AND MAKE HIM THE SUPREME RULER OF THIS LAND&amp;quot;. The rule was: For the i-th number, first letter of the i-th word of the i-th Proquiritation. &lt;a href=&#34;https://github.com/sanand0/research/tree/main/cyphral-distich&#34;&gt;ChatGPT verified it&lt;/a&gt;.
The filtering process to pick a low-hanging fruit is interesting: &amp;ldquo;I asked it to solve an unsolved cipher &amp;hellip; avoid ciphers that already had solutions&amp;hellip; I steered it away from the absolute hardest problems&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mistakes-i-made&#34;&gt;Mistakes I made&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/mistakes-i-made/#week-ending-2026-09-27&#34;&gt;Week ending 27 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;I know what factors are important and therefore I can start making changes&amp;rdquo;&lt;/strong&gt; while using a predictive decision tree on student grades to identify interventions.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: A predictive model can identify variables associated with grades, but that does not show that changing those variables will improve grades; interventions need causal evidence or assumptions designed to identify causal effects.&lt;br&gt;
&lt;strong&gt;MEDIUM · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;GPT-3.5 Turbo&amp;hellip; only cost you 50 cents&amp;rdquo;&lt;/strong&gt; while describing the cost of models available in March 2023.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: GPT-3.5 Turbo launched on 1 Mar 2023 at $0.002 per 1,000 tokens, or $2 per million tokens; $0.50 per million input tokens arrived with &lt;code&gt;gpt-3.5-turbo-0125&lt;/code&gt; in January 2024.&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;a million tokens&amp;hellip; about a million words&amp;rdquo;&lt;/strong&gt; while translating model pricing into document size.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Tokens and words are not interchangeable; for English, OpenAI&amp;rsquo;s rough rule is 1 token ≈ 0.75 words, so one million words is roughly 1.33 million tokens, with the exact count depending on the text and tokenizer.&lt;br&gt;
&lt;strong&gt;LOW · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;if there are more objects than containers, then there will be something that&amp;rsquo;s left out&amp;rdquo;&lt;/strong&gt; while explaining the pigeonhole principle.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: If there are more objects than containers and every object is placed in a container, at least one container must contain at least two objects; nothing needs to be left out.&lt;br&gt;
&lt;strong&gt;LOW · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;logprobs—the chance that it might have made a mistake&amp;rdquo;&lt;/strong&gt; while explaining how to prioritize AI outputs for human review.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Logprobs are probabilities assigned to generated tokens, not probabilities that an answer is wrong. They can be useful uncertainty signals for ranking review priority, but that relationship needs to be validated or calibrated on the task.&lt;br&gt;
&lt;strong&gt;MEDIUM · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;LLM doesn&amp;rsquo;t care what language you&amp;rsquo;re speaking in&amp;hellip; language agnostic&amp;rdquo;&lt;/strong&gt; while discussing multilingual workflows in an AI workshop.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: LLMs are multilingual, not language-agnostic: capability varies by language, task and model, and lower-resource languages can perform substantially worse. Even OpenAI says its models are optimized for English. Test the actual languages required before treating a workflow as language-agnostic.&lt;br&gt;
&lt;strong&gt;MEDIUM · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;questions-i-was-asked&#34;&gt;Questions I was asked&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/questions-i-am-asked/#week-ending-2026-09-27&#34;&gt;Week ending 27 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: How do we calibrate an AI app that hallucinates information not present in the source? Asked while testing a recruiting app that was inventing details not present in candidate CVs.&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Log the inputs, outputs, and human corrections. Build up that history, then use it to test prompt or model changes and whether a second-pass check catches the same mistakes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: If AI automates part of the work but people still check everything, how do we get real productivity? Asked while discussing automation that improved output but still required full QC because the team did not trust it enough to let work pass unchecked.&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Don&amp;rsquo;t automate everything a little. Pull out even one 10% slice where you can get to full confidence, stop checking it, and redesign the workflow so that 10% becomes an actual capacity saving.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Should we replace mature rule-based automation with AI-native workflows? Asked while discussing how new AI workflows were taking time just to recover productivity already achieved through deterministic automation.&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: No. Keep the rules that already work and use agents to find missing rules and improve existing ones from correction logs. Deterministic checks give you confidence and can bring the LLM cost down to zero.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Remote Desktop Commander ChatGPT Plugin</title>
      <link>https://www.s-anand.net/blog/remote-desktop-commander-chatgpt-plugin/</link>
      <pubDate>Sat, 26 Sep 2026 17:37:46 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/remote-desktop-commander-chatgpt-plugin/</guid>
      <description>&lt;p&gt;The &lt;a href=&#34;https://chatgpt.com/plugins/plugin_asdk_app_6a057d268ebc81919918d37eec718425&#34;&gt;Remote Desktop Commander ChatGPT Plugin&lt;/a&gt; might be one of the most useful power-user plugins for ChatGPT. Here&amp;rsquo;s how it works.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;You install the plugin and log into &lt;a href=&#34;https://desktopcommander.app/&#34;&gt;desktopcommander.app&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;You run &lt;code&gt;npx @wonderwhy-er/desktop-commander@latest remote&lt;/code&gt; on your machine&lt;/li&gt;
&lt;li&gt;After that, ChatGPT can access your computer - read/write files, run commands, etc.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is incredibly useful because that&amp;rsquo;s like getting unlimited Codex usage. ChatGPT Chat doesn&amp;rsquo;t charge by token usage. So you can write and run code on your machine without worrying about token limits. (This doesn&amp;rsquo;t help so much with Claude - it charges the same for Chat and Code.)&lt;/p&gt;
&lt;p&gt;Secondly, you can use the context of past chats and memories (or perhaps projects) - which Codex and Claude Code doesn&amp;rsquo;t have.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://mcp.desktopcommander.app/&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-26-remote-desktop-commander-chatgpt-plugin.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I wrote a version of this plugin for myself - &lt;a href=&#34;https://github.com/sanand0/scripts/blob/8302d6266b295887e295d58cb6dd87a70c4828f7/mcpserver.py&#34;&gt;mcpserver.py&lt;/a&gt; - and it&amp;rsquo;s been the most used plugin. Almost 70% of my chats (237 / 344) in August used this plugin.&lt;/p&gt;
&lt;p&gt;Since I already have my plugin I probably won&amp;rsquo;t be switching. Also, your data passes through a third-party, and their liberal free-tier (10K tool calls per month) might evaporate.&lt;/p&gt;
&lt;p&gt;But this may be the most useful plugin for ChatGPT power users.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Quality levels of GPT Image 2.5 Flare</title>
      <link>https://www.s-anand.net/blog/quality-levels-of-gpt-image-2.5-flare/</link>
      <pubDate>Sat, 26 Sep 2026 16:44:30 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/quality-levels-of-gpt-image-2.5-flare/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://developers.openai.com/api/docs/models/gpt-image-2.5-flare&#34;&gt;GPT Image 2.5 Flare&lt;/a&gt; is a &lt;a href=&#34;https://www.s-anand.net/blog/converting-black-and-white-photos-to-color-with-gpt-image-2.5/&#34;&gt;pretty good image model&lt;/a&gt;. It has a &lt;a href=&#34;https://developers.openai.com/api/docs/guides/image-prompting&#34;&gt;quality parameter&lt;/a&gt; that can be set to &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Higher levels generate more tokens and here&amp;rsquo;s the &lt;a href=&#34;https://cellcog.ai/blog/gpt-image-2-5-release-date/&#34;&gt;rough cost by quality&lt;/a&gt; for a 1024x1024 image. This cost is in &lt;strong&gt;cents not dollars&lt;/strong&gt;:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Quality&lt;/th&gt;
					&lt;th style=&#34;text-align: right&#34;&gt;Tokens&lt;/th&gt;
					&lt;th style=&#34;text-align: right&#34;&gt;Cents&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;low&lt;/code&gt;&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;196&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;0.6&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;medium&lt;/code&gt;&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;439&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;1.3&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;high&lt;/code&gt;&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;1,756&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;5.3&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;xhigh&lt;/code&gt;&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;3,122&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;9.4&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;max&lt;/code&gt;&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;7,024&lt;/td&gt;
					&lt;td style=&#34;text-align: right&#34;&gt;21.1&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;But what difference does it really make? I &lt;a href=&#34;https://chatgpt.com/share/6ab787dd-7a90-83ec-b006-6a5fe367ea33&#34;&gt;asked ChatGPT to experiment&lt;/a&gt; &lt;!-- https://chatgpt.com/c/6ab607d1-0090-83ec-a243-e17a885fe575 --&gt; and find an image where there is a clear difference.&lt;/p&gt;
&lt;p&gt;It began with a macro photo of a watch. See the difference between &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;max&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/watch-low.avif&#34;&gt;
&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/watch-max.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;First, let&amp;rsquo;s note what&amp;rsquo;s NOT different. The time, tiny text, exact watch hands, wool/steel texture, droplet/refraction. So, &lt;em&gt;&lt;code&gt;low&lt;/code&gt; is already a pretty good model&lt;/em&gt;! It&amp;rsquo;s hart to tell the difference in quality.&lt;/p&gt;
&lt;p&gt;Not much difference in a transit time poster either. The aesthetics are slightly better in &lt;code&gt;high&lt;/code&gt; (the second image), which looks nearly identical to &lt;code&gt;xhigh&lt;/code&gt; and &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/transit-low.avif&#34;&gt;
&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/transit-high.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;On a picture of a barista, &lt;code&gt;low&lt;/code&gt; already produced realistic skin, fabric, scratched metal, condensation, transparent glass, steam, rain reflections and latte art.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/barista-low.avif&#34;&gt;
&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/images/barista-medium.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;So, to give itself a challenge, ChatGPT asked for a larger image (2880x2880) with freckles,&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Ultra-realistic unretouched beauty-editorial portrait, perfectly front-facing and centered, of a woman in her mid-30s with naturally freckled light-brown skin and dark curly hair. Square composition; her head and upper shoulders fill almost the entire frame, with both eyes sharply in focus and the face occupying about 75% of the image width. Neutral warm-gray studio background, one large softbox slightly camera-left, natural color, no glamour retouching and no skin smoothing. Resolve extremely fine real-world surface detail: distinct pores across the nose and cheeks; tiny vellus peach-fuzz hairs catching side light; irregular freckles of different sizes and densities; one faint healed 8 mm scar on the left cheek; individual eyebrow hairs; separate upper and lower eyelashes; fine radial fibres and color variation in both irises; subtle wet tear-line reflections; natural lip lines and slight dry texture; individual flyaway hairs and fine frizz around the hairline. She wears a charcoal-gray chunky knitted wool turtleneck whose individual yarn fibres, twisted strands and knit loops are clearly visible, plus one simple brushed-titanium hoop earring showing very fine directional brushing and a few microscopic hairline scratches. Preserve believable human skin and anatomy. The image should reward inspection at 100% zoom: photographic microcontrast and natural high-frequency texture without artificial sharpening, plastic skin, painterly texture, beauty-filter smoothing, fake grain, text, jewelry other than the single hoop, or background objects.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then it zoomed into the forehead and started looking for differences. At this &lt;em&gt;zoomed in&lt;/em&gt; level, differences start to appear.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://sanand0.github.io/llmevals/gpt-image-flare-quality/#where-it-shows&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/202-09-26-quality-levels-of-gpt-image-2.5-flare.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;low&lt;/code&gt; generates less detail. &lt;code&gt;medium&lt;/code&gt; is better, and &lt;code&gt;high&lt;/code&gt; has a lot of detail. But, over several blind tests, ChatGPT couldn&amp;rsquo;t tell the difference between &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt; and &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Summary&lt;/strong&gt;: &lt;code&gt;low&lt;/code&gt; quality (0.6 cents) is all you need. Go for &lt;code&gt;medium&lt;/code&gt; (1.3 cens) or at most &lt;code&gt;high&lt;/code&gt; (5.3 cents) if you really fine detail like fibre, hair, wrinkles, etc. But you almost never need &lt;code&gt;xhigh&lt;/code&gt; (9.4 cents) or &lt;code&gt;max&lt;/code&gt; (21 cents).&lt;/p&gt;
&lt;p&gt;PS: I also updated my &lt;a href=&#34;https://sanand0.github.io/llmartstyle/&#34;&gt;LLM Art Style&lt;/a&gt; gallery with GPT Image 2.5 Flare images.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Editorial Slop</title>
      <link>https://www.s-anand.net/blog/editorial-slop/</link>
      <pubDate>Sat, 26 Sep 2026 08:56:17 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/editorial-slop/</guid>
      <description>&lt;p&gt;My article &lt;a href=&#34;https://cxotoday.com/corner-office/redesigning-the-operating-model-shifting-from-ai-tool-rollouts-to-workflow-integration/&#34;&gt;Redesigning the Operating Model: Shifting from AI Tool Rollouts to Workflow Integration&lt;/a&gt; appeared on CXOToday two days ago. Here&amp;rsquo;s how it happened.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;29 May 2026&lt;/strong&gt;: &lt;a href=&#34;https://www.linkedin.com/in/palashbhattacharjee/&#34;&gt;Palash&lt;/a&gt; mailed me that we have an &amp;ldquo;Email interaction opportunity with &lt;a href=&#34;https://digitalterminal.in/&#34;&gt;Digital Terminal&lt;/a&gt;&amp;rdquo; and they shared six questions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;What are the key reasons behind this “last-mile problem” in scaling AI to production?&lt;/li&gt;
&lt;li&gt;While much of the focus is on models and tools, how critical are content readiness and data quality in determining whether AI delivers real business value?&lt;/li&gt;
&lt;li&gt;(&amp;hellip; and so on.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;He&amp;rsquo;d already drafted the responses&lt;/em&gt; and &amp;ldquo;sharing below the link for your feedback and approval.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;A few hours later I replied, saying I&amp;rsquo;d rather write my own article. &lt;a href=&#34;https://docs.google.com/document/d/1PAJz-fc69tz3LoJCmTNMyluJnzCayWmmFttcZu-KcbA/edit&#34;&gt;Here it is&lt;/a&gt; (&lt;a href=&#34;https://github.com/sanand0/blog/blob/main/pages/notes/2026-05-29-digital-terminal-interview.md&#34;&gt;Markdown&lt;/a&gt;):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You&amp;rsquo;re welcome to share mine with my name.
Or you can share the earlier version with anyone else&amp;rsquo;s name - &lt;strong&gt;not mine&lt;/strong&gt;.
Either option works for me.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Palash preferred my version and shared it.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-26-editorial-slop.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;24 Sep 2026&lt;/strong&gt;: 12 weeks later, the &lt;a href=&#34;https://cxotoday.com/corner-office/redesigning-the-operating-model-shifting-from-ai-tool-rollouts-to-workflow-integration/&#34;&gt;article appears&lt;/a&gt;, and with a few differences - 🟢 good and 🔴 bad. They:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;🟢 Added a title and a &lt;a href=&#34;https://cxotoday.com/wp-content/uploads/2026/09/Anand-S.jpg&#34;&gt;photo&lt;/a&gt; - edited from my &lt;a href=&#34;https://www.s-anand.net/blog/assets/Anand-5a-1.webp&#34;&gt;speaker photo&lt;/a&gt; to remove shadows and wear a jacket in an office. (The original was at PyCon 2025 and I was wearing a T-shirt).
&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-26-editorial-slop-anand-photo.avif&#34;&gt; &lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/Anand-5a-1.webp&#34;&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🟢 Changed &lt;code&gt;--&lt;/code&gt; and &lt;code&gt;-&lt;/code&gt; to em-dashes: &lt;code&gt;—&lt;/code&gt;; also the double quotes &lt;code&gt;&amp;quot;&lt;/code&gt; to smart quotes &lt;code&gt;“&lt;/code&gt; and &lt;code&gt;”&lt;/code&gt;; single quotes &lt;code&gt;&#39;&lt;/code&gt; to apostrophes &lt;code&gt;’&lt;/code&gt; (which is interesting because I told my agent to remove those to avoid it sounding agent-y).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🔴 But failed to fix my punctuation. For example, a missing full-stop in &lt;code&gt;... fundamentally redesigned workflows In a separate workplace report...&lt;/code&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🔴 Also added paragraph breaks mid-sentence. This happened thrice. For example, here&amp;rsquo;s one broken sentence:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;High risk

actions need approval.
&lt;/code&gt;&lt;/pre&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🟢 Removed a duplicate question and my &lt;code&gt;Anand:&lt;/code&gt; quote prefixes.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🟡 But prefixed a two paragraph summary that&amp;rsquo;s more pompous and jargon-y than me.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🔴 Removed every hyperlink and changed it into underlined text that &lt;em&gt;looks&lt;/em&gt; hyperlinked but isn&amp;rsquo;t. My links were to McKinsey, NIST, and arXiv. None are clickable.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The underlining is sometimes unbalanced. In &lt;code&gt;(McKinsey &amp;amp; Company)&lt;/code&gt;, the opening paranthesis isn&amp;rsquo;t underlined, but the closing is.&lt;/li&gt;
&lt;li&gt;The links are sometimes split into multiple underlined segments. &lt;code&gt;A Survey of&lt;/code&gt; and &lt;code&gt;Data Agents&lt;/code&gt; are two separated underlined blocks - though it used to be one link.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;🔴 Added poor / unsemantic HTML markup. All six interview questions are encoded as &lt;code&gt;&amp;lt;h6&amp;gt;&lt;/code&gt; headings directly under the article, not H2/H3. There is also a stray final &lt;code&gt;&amp;lt;p&amp;gt;&amp;amp;nbsp;&amp;lt;/p&amp;gt;&lt;/code&gt; after the article.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;On the margin, this may be more damage than good. &amp;ldquo;Editorial slop&amp;rdquo;, I guess.&lt;/p&gt;
&lt;p&gt;PS: I make more mistakes than any editor and hate being called out. But I don&amp;rsquo;t mind. I&amp;rsquo;m a &lt;em&gt;happy&lt;/em&gt; hypocrite.&lt;/p&gt;
&lt;p&gt;PPS: I&amp;rsquo;m not moralizing &amp;ldquo;Humans generate slop too, so don&amp;rsquo;t blame AI&amp;rdquo;, etc. I&amp;rsquo;m more mature for that. I&amp;rsquo;m just saying, &amp;ldquo;Hee hee, you made a boo-boo 🙂&amp;rdquo;.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>If You&#39;re Too Excited, Don&#39;t Forget to Verify</title>
      <link>https://www.s-anand.net/blog/if-youre-too-excited-dont-forget-to-verify/</link>
      <pubDate>Thu, 24 Sep 2026 10:15:00 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/if-youre-too-excited-dont-forget-to-verify/</guid>
      <description>&lt;p&gt;I conducted a session on Thu, 24 Sep 2026 at &lt;a href=&#34;https://iis.com.ph/&#34;&gt;International IT-BPM Summit (IIS) 2026&lt;/a&gt; - Function Rooms #1 &amp;amp; #2, 3rd Floor Pearl Wing, Okada Manila, Parañaque City, Philippines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Summary&lt;/strong&gt;: AI is too weird and fast-moving to trust by intuition alone: question advice, verify with a second model, calibrate confidence, benchmark what matters, and turn surviving evidence into deterministic rules.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-24-iis-ph-evidence-to-impact/&#34;&gt;Here&amp;rsquo;s the link to the session&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-24-iis-ph-evidence-to-impact/&#34;&gt;&lt;img alt=&#34;Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-24-iis-ph-evidence-to-impact/comic-page.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-24-iis-ph-evidence-to-impact/transcript.md&#34;&gt;Transcript&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/talks/releases/download/talks/2026-09-24-iis-ph-evidence-to-impact.opus&#34;&gt;Audio&lt;/a&gt; (60 min)&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Using agents to answer exams</title>
      <link>https://www.s-anand.net/blog/using-agents-to-answer-exams/</link>
      <pubDate>Wed, 23 Sep 2026 10:53:03 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/using-agents-to-answer-exams/</guid>
      <description>&lt;p&gt;Our recruitment team asked me to review hiring questions for analysts and data scientists. These were on &lt;a href=&#34;https://imocha.io/&#34;&gt;iMocha&lt;/a&gt; - a proctored assessment platform.&lt;/p&gt;
&lt;p&gt;I logged in. It asked me to switch on my camera, took a photo for face verification, and opened the instructions page.&lt;/p&gt;
&lt;h3 id=&#34;agents-can-solve-exams&#34;&gt;Agents can solve exams&lt;/h3&gt;
&lt;p&gt;I told Codex CLI (running GPT 5.6 Luna Medium):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href=&#34;https://test.imocha.io/test/0/0/1&#34;&gt;https://test.imocha.io/test/0/0/1&lt;/a&gt; is open on the browser - CDP on localhost:9222&lt;br&gt;
This is a practice test. Solve it. Log progress and results in notes.md.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After 11 minutes, it came back with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Completed the practice test and logged everything in notes.md.&lt;/p&gt;
&lt;p&gt;General section: 10/10 attempted and submitted.&lt;br&gt;
Coding section: submitted; grader showed formatting-related failures.&lt;br&gt;
MySQL section: passed.&lt;br&gt;
Uploaded report: formal_sales_report.txt&lt;br&gt;
Practice flow finished and redirected away from iMocha.\&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There were a few things it struggled with. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Unusual radio buttons&lt;/strong&gt;. Codex learnt that it needs to click, then verify if the page actually shows the radio buttons as checked.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auto-formatting&lt;/strong&gt;. The code editor re-formats text while typing. Codex learnt: &amp;ldquo;For small code, click, &lt;code&gt;Control+A&lt;/code&gt;, and type; for larger code, set &lt;code&gt;monaco.editor.getModels()[0].setValue(source)&lt;/code&gt; and inspect the resulting source before compiling.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So I told it to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Document learnings about how to use the imocha.io tests in an imocha-test/SKILL.md - compatible with how skills loaded in projects.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Click &amp;ldquo;Start Test&amp;rdquo; and solve the test.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;12 minutes later, it did. I just watched it, staring at the webcam, while eating an apple.&lt;/p&gt;
&lt;div class=&#34;video-embed&#34;&gt;&lt;iframe width=&#34;560&#34; height=&#34;315&#34; src=&#34;https://www.youtube.com/embed/AxXQhSehfbw&#34; title=&#34;YouTube video&#34; loading=&#34;lazy&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture&#34; allowfullscreen&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;h3 id=&#34;proctoring-may-need-revision&#34;&gt;Proctoring may need revision&lt;/h3&gt;
&lt;p&gt;Interestingly, the proctoring report shows no problems. It took pictures of me (maybe every minute or so) and since it was the same person with no one else present, it assumed all is well.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&#34;The Proctoring Report shows 0 Total Violations.&#34; loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-22-using-agents-to-take-exams-proctoring-report.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&#34;While I was eating an apple&#34; loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-22-using-agents-to-take-exams-anand-apple-1.avif&#34;&gt;
&lt;img alt=&#34;&amp;hellip; and the agents was coding.&#34; loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-22-using-agents-to-take-exams-anand-apple-2.avif&#34;&gt;
&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-22-using-agents-to-take-exams-anand-apple-3.avif&#34;&gt;
&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-22-using-agents-to-take-exams-anand-apple-4.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;In other words:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;At least &lt;em&gt;this&lt;/em&gt; app isn&amp;rsquo;t AI agent proof yet&lt;/li&gt;
&lt;li&gt;If someone were telling me what to type from outside the field of vision, it wouldn&amp;rsquo;t work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I was later told by our recruiting team that many of these exams are conducted in testing centers where agents aren&amp;rsquo;t allowed - in which case, online proctoring might not be required in the first place.&lt;/p&gt;
&lt;h3 id=&#34;questions-banks-have-errors&#34;&gt;Questions banks have errors&lt;/h3&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One question asked how to filter a Pandas table. Luna used&amp;rsquo; &lt;code&gt;query()&lt;/code&gt;, which worked when executed, but iMocha marked it wrong.&lt;/li&gt;
&lt;li&gt;One compound-interest question had four answer choices. &lt;strong&gt;None were correct.&lt;/strong&gt; Luna picked one and got the mark anyway.&lt;/li&gt;
&lt;li&gt;One reasoning question gave facts about robins and sparrows, then asked what we could conclude about finches. &lt;strong&gt;Nothing about finches followed from the facts given&lt;/strong&gt;, but the test still had a &amp;ldquo;correct&amp;rdquo; answer.&lt;/li&gt;
&lt;li&gt;One SQL question asked to join salary history and department history. Luna gave an answer that iMocha accepted, but that was a &lt;strong&gt;wrong join&lt;/strong&gt; that could assign an old salary to the wrong department!&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These questions were probably drawn from a question bank. My lessons:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Question banks have errors&lt;/li&gt;
&lt;li&gt;Agents can cheaply check quality&lt;/li&gt;
&lt;li&gt;I need to control my temper. When our recruiting team asked, &amp;ldquo;So, Anand, should we use agents to quality-check questions?&amp;rdquo; it took me a few deep breaths to control my blood pressure and ask sweetly, &amp;ldquo;What do &lt;em&gt;you&lt;/em&gt; think?&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;we-need-better-questions&#34;&gt;We need better questions&lt;/h3&gt;
&lt;p&gt;Many questions test things like writing SQL / Python scripts for simple tasks. Agents do these &lt;em&gt;very&lt;/em&gt; well.&lt;/p&gt;
&lt;p&gt;Maybe we still need to test for these. But let&amp;rsquo;s reduce it and instead add new types of questions, like:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Use agents well&lt;/strong&gt;. Give them hard tasks like &amp;ldquo;Analyze this &lt;em&gt;large&lt;/em&gt; dataset&amp;rdquo;, &amp;ldquo;Deploy an app&amp;rdquo;, etc. that check if they can use agents well. (Not everyone can.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verify agent output&lt;/strong&gt;. Give them results with errors ranging from blatant to subtle. Can they spot them all (with or without agents)?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test new skills&lt;/strong&gt; that are important in the AI era.
&lt;ul&gt;
&lt;li&gt;Do they speak well? &amp;ldquo;Upload a 1 min audio/video answering why we should hire you.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Are they diligent? &amp;ldquo;Mail &lt;a href=&#34;mailto:careers@example.org&#34;&gt;careers@example.org&lt;/a&gt; a week later to follow up.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Do they take ownership? &amp;ldquo;Your answer to Q5 above had this mistake: ___. Why did that happen?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Do they show initiative? &amp;ldquo;Create a GitHub repo and get 20 stars before the exam ends.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;testing-infrastructure-needs-to-change&#34;&gt;Testing infrastructure needs to change&lt;/h3&gt;
&lt;p&gt;I realized that it&amp;rsquo;s harder for recruiters to change and easier for the testing infrastructure to enale these. Specfically:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Question banks&lt;/strong&gt; need to change: picking from a &amp;ldquo;Gen AI skill&amp;rdquo; question bank is easier for recruiters.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Testing software&lt;/strong&gt; needs to change: to detect AI use as well as selectively allow it for questions where it&amp;rsquo;s required.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test centers&lt;/strong&gt; need to change: to allow and enable agent usage for specific exams.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Watching videos with a phone holder</title>
      <link>https://www.s-anand.net/blog/watching-videos-with-a-phone-holder/</link>
      <pubDate>Tue, 22 Sep 2026 19:08:00 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/watching-videos-with-a-phone-holder/</guid>
      <description>&lt;p&gt;On Air India AI 2531 from Mumbai to Hyderabad, I saw something ingenious: the seat had a phone holder built in.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-17-air-india-seat-back-phone-holder.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;It stretches up and down, so it should fit a wide range of mobiles. Put your phone in, play a video, and you have your own in-flight screen. This flight didn&amp;rsquo;t have a screen, so this is a pretty good substitute.&lt;/p&gt;
&lt;p&gt;Earlier this year, on flights from Singapore to Chennai, one passenger used a &lt;a href=&#34;https://www.s-anand.net/blog/watching-videos-with-a-plastic-cover/&#34;&gt;plastic cover&lt;/a&gt; to hang her phone from the tray table.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-06-05-indigo-flight-phone-plastic-video.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Another used his &lt;a href=&#34;https://www.s-anand.net/blog/watching-videos-with-a-phone-case/&#34;&gt;phone case and the headrest cover&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-08-17-watching-videos-with-a-phone-case.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Air India has been consciously adding these to its newer cabins for at least a couple of years. In &lt;a href=&#34;https://www.airindia.com/in/en/newsroom/press-release/all-new-business-premium-economy-cabins.html&#34;&gt;June 2024, it announced PED (personal electronic device) holders across Business, Premium Economy and Economy on its new A320neo cabins&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Nice work, Air India. I suspect they aren&amp;rsquo;t the only airline doing this, and likely weren&amp;rsquo;t the first. But this is the first time I&amp;rsquo;ve seen it.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Slingshotting from Singapore to Timbuktu</title>
      <link>https://www.s-anand.net/blog/slingshotting-from-singapore-to-timbuktu/</link>
      <pubDate>Sun, 20 Sep 2026 22:08:58 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/slingshotting-from-singapore-to-timbuktu/</guid>
      <description>&lt;p&gt;My daughter and I planned a trip to Timbuktu. For good reasons. &lt;a href=&#34;https://en.wikipedia.org/wiki/Mansa_Musa&#34;&gt;Mansa Musa&lt;/a&gt;, perhaps the richest person in history, ruled there. It&amp;rsquo;s right at the &lt;a href=&#34;https://maps.app.goo.gl/G2PgrnFx4SSapSBp8&#34;&gt;&lt;em&gt;edge&lt;/em&gt; of the Sahara desert&lt;/a&gt;. Buildings are made of &lt;a href=&#34;https://commons.wikimedia.org/wiki/File:Timbuktu-107967.jpg&#34;&gt;yellow bricks&lt;/a&gt;. And&amp;hellip; well, think about telling your friends, &amp;ldquo;Oh, I just returned from Timbuktu.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;We ruled out flying. Flying is for losers.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s &lt;a href=&#34;https://maps.app.goo.gl/FJS6A7SCrgycajeAA&#34;&gt;possible to walk&lt;/a&gt; but it&amp;rsquo;d take 3,800 hours (many months) from Singapore and require 15 visas - Malaysia, Thailand, Myanmar, Pakistan, Afghanistan, Iran, Iraq, Syria, Jordan, Israel, Egypt, Libya, Algeria, Niger, and Mali (many months).&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s shorter to sail to &lt;a href=&#34;https://maps.app.goo.gl/iVd1B798VmJbzFFs6&#34;&gt;Accra, Ghana, and walk from there&lt;/a&gt; - 366 hours (many weeks) and 3 visas - Ghana, Burkina Faso, and Mali (few weeks).&lt;/p&gt;
&lt;p&gt;Or we could sail to &lt;a href=&#34;https://maps.app.goo.gl/TdHBJQT96mFtnVVv7&#34;&gt;Calabar, Nigeria&lt;/a&gt; and sail up River Niger - few weeks and 3 visas - Nigeria, Niger, and Mali (few weeks).&lt;/p&gt;
&lt;p&gt;But the whole visa thing is a pain. So (after I ruled out space ships, because despite what my daughter tried to convince me of, they&amp;rsquo;re not ships, they&amp;rsquo;re planes and we ruled them out) we decided to slingshot ourselves from Singapore to Timbuktu.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-20-slingshotting-from-singapore-to-timbuktu.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s what Gemini, Claude, and ChatGPT thought of our idea.&lt;/p&gt;
&lt;p&gt;(Personally, I liked Gemini&amp;rsquo;s response the best.)&lt;/p&gt;
&lt;h2 id=&#34;user&#34;&gt;User&lt;/h2&gt;
&lt;p&gt;What would it take to slingshot someone from Singapore to Timbuktu? Answer as Randall Munroe would.&lt;/p&gt;
&lt;section ai-disclosure=&#34;ai-generated&#34; data-ai-model=&#34;gemini-3.6-flash&#34; data-ai-provider=&#34;Google&#34;&gt;
&lt;h2 id=&#34;gemini-36-flash-thinking&#34;&gt;Gemini 3.6 Flash Thinking&lt;/h2&gt;
&lt;!-- https://gemini.google.com/app/fcbcea71424afee4 --&gt;
&lt;p&gt;The straight-line distance along Earth’s surface from Singapore to Timbuktu is roughly 10,700 kilometers—nearly a quarter of the way around the planet.&lt;/p&gt;
&lt;p&gt;To throw an object from Singapore to Mali in a single ballistic arc, you need an launch velocity of roughly $6.8\text{ km/s}$ (Mach 20). That is dangerously close to orbital velocity ($7.8\text{ km/s}$). If you go even slightly faster, you won&amp;rsquo;t land in Timbuktu; you will overshoot the Earth entirely and become a permanent artificial satellite.&lt;/p&gt;
&lt;p&gt;Achieving this with a giant slingshot presents four distinct engineering problems.&lt;/p&gt;
&lt;h3 id=&#34;1-acceleration-vs-liquefaction&#34;&gt;1. Acceleration vs. Liquefaction&lt;/h3&gt;
&lt;p&gt;If you build a traditional Y-shaped slingshot with a pouch pulled back, say, 100 meters, accelerating a human from zero to $6.8\text{ km/s}$ over that distance requires an acceleration of roughly $23,000\text{ g}$.&lt;/p&gt;
&lt;p&gt;Human beings do not survive $23,000\text{ g}$. Fighter pilots black out around $9\text{ g}$. At $23,000\text{ g}$, the distinction between human anatomy and fluid dynamics disappears instantly. You wouldn&amp;rsquo;t be a passenger in a sling; you would be a thin red paint applied to the inside of the pouch before it even reached the launch post.&lt;/p&gt;
&lt;p&gt;To make the acceleration survivable for a human—let&amp;rsquo;s assume a generous $10\text{ g}$ limit—your acceleration track needs to be exceptionally long:&lt;/p&gt;
&lt;p&gt;$$d = \frac{v^2}{2a} = \frac{(6,800\text{ m/s})^2}{2 \times 98\text{ m/s}^2} \approx 235\text{ kilometers}$$&lt;/p&gt;
&lt;p&gt;Singapore is only 50 kilometers wide from west to east. To pull back your slingshot far enough to keep the passenger alive, you would have to anchor the rubber band in Singapore, stretch the pouch across the entire width of the country, cross the Strait of Johor, and pull it deep into the jungles of Malaysia.&lt;/p&gt;
&lt;h3 id=&#34;2-the-atmospheric-wall&#34;&gt;2. The Atmospheric Wall&lt;/h3&gt;
&lt;p&gt;Assuming you build a 235-kilometer-long hyper-slingshot, you encounter a second, much faster death: Singapore&amp;rsquo;s humid, sea-level atmosphere.&lt;/p&gt;
&lt;p&gt;At Mach 20, air cannot get out of the way fast enough. The air in front of the sling pouch is compressed so violently that it heats up into a superheated plasma sheath reaching several thousand degrees Celsius. Leaving the pouch at sea-level dense atmosphere at $6.8\text{ km/s}$ means you would be incinerated by atmospheric drag within a fraction of a millisecond. You would effectively trigger a thermobaric explosion at ground level.&lt;/p&gt;
&lt;p&gt;To fix this, you would need to build your 235-kilometer slingshot inside a vacuum-sealed tube that slopes upward, terminating at a giant launch hatch at the top of Mount Kinabalu (or high in the stratosphere via a giant, rigid space-pipe).&lt;/p&gt;
&lt;h3 id=&#34;3-the-rubber-band-problem&#34;&gt;3. The Rubber Band Problem&lt;/h3&gt;
&lt;p&gt;Latex rubber bands store elastic potential energy, but they are limited by the speed of sound within the rubber itself. A stretched rubber band cannot snap back faster than the speed of a sound wave traveling through the material, which for natural rubber is around 30 to 70 meters per second.&lt;/p&gt;
&lt;p&gt;If you release a rubber slingshot aiming for $6,800\text{ m/s}$, the pouch simply will not move fast enough to keep up with the trajectory. The elastic band would hang limply in the air while you gently roll out of the pouch at jogging speed. You would need to replace the rubber with high-yield magnetic levitation rails or an explosive gas-piston system, at which point you have stopped building a slingshot and accidentally built a 200-kilometer suborbital coilgun.&lt;/p&gt;
&lt;h3 id=&#34;4-arrival-in-timbuktu&#34;&gt;4. Arrival in Timbuktu&lt;/h3&gt;
&lt;p&gt;If you survive the $10\text{ g}$ launch, exit the vacuum tube into the upper atmosphere, ride a high suborbital arc through space, and re-enter over West Africa, you arrive above Timbuktu in about 25 minutes.&lt;/p&gt;
&lt;p&gt;Without retro-rockets, a heavy heat shield, and giant parachutes, you would re-enter the lower atmosphere still moving at hypersonic speeds. The kinetic energy of a 70-kilogram human traveling at Mach 10 is roughly equivalent to 350 kilograms of TNT.&lt;/p&gt;
&lt;p&gt;You would not land in Timbuktu so much as you would create a modest, newly named impact crater just outside the city limits.&lt;/p&gt;
&lt;/section&gt;
&lt;section ai-disclosure=&#34;ai-generated&#34; data-ai-model=&#34;claude-opus-5&#34; data-ai-provider=&#34;Anthropic&#34;&gt;
&lt;h2 id=&#34;claude-opus-5-medium&#34;&gt;Claude Opus 5 Medium&lt;/h2&gt;
&lt;!-- https://claude.ai/chat/8f5df60b-f5ea-41a1-9e29-2805733fee45 --&gt;
&lt;p&gt;&lt;strong&gt;Q. What would it take to slingshot someone from Singapore to Timbuktu?&lt;/strong&gt; — &lt;em&gt;Anand, Singapore&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The short answer is: about 7.4 kilometers per second, a launch rail whose far end sticks out of the atmosphere, and a cause of death that doesn&amp;rsquo;t currently have a name.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The distance problem&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Singapore to Timbuktu is 11,700 km around the curve of the Earth — 106 degrees of arc, or roughly a third of the way around the planet. This is the awkward middle distance for ballistics. If you want to throw something 10 km, you need a good cannon. If you want to throw something 40,000 km, you don&amp;rsquo;t: you put it in orbit and wait.&lt;/p&gt;
&lt;p&gt;For 11,700 km you need almost, but not quite, orbital speed. The minimum-energy lob works out to 7.4 km/s launched at 18.6° above the horizon, which is 94% of the speed needed to just leave and never come back. You would coast up to 1,270 km — three times higher than the Space Station — and arrive 11 minutes later.&lt;/p&gt;
&lt;p&gt;Eleven minutes is very good. Singapore Airlines cannot do this.&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The slingshot problem&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A slingshot is a device that gets something up to speed over a short distance, and &amp;ldquo;short distance&amp;rdquo; is where things go wrong.&lt;/p&gt;
&lt;p&gt;Humans can survive sustained acceleration of maybe 10 g if they&amp;rsquo;re lying down, well-strapped, and have made peace with their choices. Getting to 7.4 km/s at 10 g takes 76 seconds and 283 kilometers of runway.&lt;/p&gt;
&lt;p&gt;Singapore is 50 km across.&lt;/p&gt;
&lt;p&gt;So your slingshot has to point out over the Strait of Malacca, and because it&amp;rsquo;s angled at 18.6°, the launch end sits in Tuas and the business end sits &lt;strong&gt;90 km above the Indian Ocean&lt;/strong&gt;, which is well above the Kármán line. Your slingshot&amp;rsquo;s tip is in space. At this point you should notice that you have built a space elevator that only goes diagonally, and is a slingshot, and ask whether the original plan is still the best one.&lt;/p&gt;
&lt;p&gt;If you insist on a gentler 3 g, the rail is 940 km long and ends 300 km up, at which point you have not built a transport system, you have built a very long finger pointing at Mali.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The air&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Here is the part that ends the discussion.&lt;/p&gt;
&lt;p&gt;At 7.4 km/s at sea level, the air in front of your passenger is not really &amp;ldquo;air&amp;rdquo; anymore. The dynamic pressure is about 34 megapascals — 340 atmospheres, pressing on the front of a person. The deceleration from drag alone is on the order of 40,000 g.&lt;/p&gt;
&lt;p&gt;Your passenger would come off the rail, travel a few hundred meters, and stop. Almost none of them would continue to Timbuktu, and the parts that did would arrive as a warm glow.&lt;/p&gt;
&lt;p&gt;Meteors the size of a grapefruit hit the atmosphere at these speeds all the time. They make a pretty streak and then they are gone. Your passenger is a large, damp, rule-following meteor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The tunnel&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The alternative is to stop fighting the air and go under it.&lt;/p&gt;
&lt;p&gt;Drill a straight tunnel from Singapore to Timbuktu. It&amp;rsquo;s 10,150 km long, and because it&amp;rsquo;s a chord through a sphere, its midpoint is &lt;strong&gt;2,520 km below the surface&lt;/strong&gt;. That&amp;rsquo;s the bottom of the mantle. The rock there is around 3,000 °C and under a million atmospheres of pressure, and your tunnel&amp;rsquo;s deepest point is about 370 km above the outer core, which in geology is a rounding error.&lt;/p&gt;
&lt;p&gt;But the physics is lovely. Drop someone in with no engine. Gravity pulls them down the first half, they coast up the second half, and they arrive at the far end with exactly zero velocity, about 40 minutes later. Peak speed at the midpoint: &lt;strong&gt;6.3 km/s&lt;/strong&gt;, or 22,700 km/h, achieved with no fuel whatsoever.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s the best transport system anyone has ever described and it requires a hole through the part of the Earth that is trying hardest to become lava.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One last thing: aim ahead&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Whichever version you build, remember that Timbuktu is moving. During your 11-minute flight, the Earth turns 2.65°, carrying your destination about 280 km east.&lt;/p&gt;
&lt;p&gt;So you don&amp;rsquo;t aim at Timbuktu. You aim at a patch of empty Saharan sand and trust the planet to slide the city under your passenger at the last moment, like a quarterback leading a receiver, if the receiver were a 12th-century trading city running east at 1,600 km/h.&lt;/p&gt;
&lt;p&gt;Miss the lead by 1%, and you arrive 3 km from where you meant to, which for most projects is fine and for this one is the difference between Timbuktu and the middle of nowhere — a distinction Timbuktu has spent 800 years trying to explain.&lt;/p&gt;
&lt;/section&gt;
&lt;section ai-disclosure=&#34;ai-generated&#34; data-ai-model=&#34;gpt-5.6-high&#34; data-ai-provider=&#34;OpenAI&#34;&gt;
&lt;h2 id=&#34;chatgpt-gpt-56-sol-high&#34;&gt;ChatGPT GPT 5.6 Sol High&lt;/h2&gt;
&lt;!-- https://chatgpt.com/c/6aafc392-7cc0-83ec-b8e3-b7632e441fca --&gt;
&lt;p&gt;I can’t write in Randall Munroe’s exact style, but I can do the same general genre: deadpan, equation-heavy, scientifically serious about an absurd premise.&lt;/p&gt;
&lt;p&gt;Singapore to Timbuktu is about &lt;strong&gt;11,750 km&lt;/strong&gt; along the Earth’s surface.&lt;/p&gt;
&lt;p&gt;If by “slingshot” we mean “large elastic device that launches one human on a ballistic trajectory,” the first problem is that &lt;strong&gt;Timbuktu is 106° around the planet&lt;/strong&gt;. You’re not throwing someone &lt;em&gt;across&lt;/em&gt; the landscape so much as briefly turning them into a satellite that has made a serious navigational error.&lt;/p&gt;
&lt;p&gt;Ignoring air resistance, mountains, political borders, and the passenger’s increasingly urgent objections, an efficient ballistic trajectory would look roughly like this:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Launch speed: ~7.45 km/s&lt;/strong&gt;
&lt;strong&gt;Maximum altitude: ~1,300 km&lt;/strong&gt;
&lt;strong&gt;Flight time: ~35 minutes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;7.45 km/s is Mach 22.&lt;/p&gt;
&lt;p&gt;For an 80 kg passenger, the kinetic energy at launch is about &lt;strong&gt;2.2 gigajoules&lt;/strong&gt;, roughly half a ton of TNT. Fortunately, unlike half a ton of TNT, the passenger is reusable in principle.&lt;/p&gt;
&lt;p&gt;The word “slingshot” now becomes problematic.&lt;/p&gt;
&lt;p&gt;If the slingshot accelerates them over 100 metres, the average acceleration is about &lt;strong&gt;28,000 g&lt;/strong&gt;. This converts the passenger into a thin educational layer on the back of the projectile.&lt;/p&gt;
&lt;p&gt;For a relatively gentle &lt;strong&gt;5 g&lt;/strong&gt; launch, you need an acceleration track of roughly:&lt;/p&gt;
&lt;p&gt;$$
d=\frac{v^2}{2a}\approx560\text{ km}
$$&lt;/p&gt;
&lt;p&gt;So your slingshot stretches approximately from Singapore to somewhere well into Malaysia.&lt;/p&gt;
&lt;p&gt;This is encouraging, because Malaysia is much easier to get planning permission from than low Earth orbit.&lt;/p&gt;
&lt;p&gt;Unfortunately, we have so far ignored the atmosphere.&lt;/p&gt;
&lt;p&gt;Launching at 7.45 km/s from sea level is similar to arranging for a meteor to start &lt;em&gt;inside&lt;/em&gt; the atmosphere. The air in front of the passenger gets compressed and heated violently. A conventional elastic cord also cannot store this energy compactly; you’ve accidentally designed something closer to a &lt;strong&gt;560-km electromagnetic launch track with a spacecraft on it&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Then comes Timbuktu.&lt;/p&gt;
&lt;p&gt;The passenger reaches Mali at several kilometres per second. “Landing” is therefore another word whose meaning has drifted during the project.&lt;/p&gt;
&lt;p&gt;You need a second 560-km slingshot pointing backward.&lt;/p&gt;
&lt;p&gt;At which point the complete Singapore–Timbuktu Slingshot consists of:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;one launch vehicle, two enormous accelerator systems, a heat shield, orbital-navigation equipment, and about 1,100 km of infrastructure.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;There is already a name for this technology.&lt;/p&gt;
&lt;p&gt;It’s a spacecraft.&lt;/p&gt;
&lt;p&gt;The good news is that the actual rubber band can be quite small.&lt;/p&gt;
&lt;/section&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;There is no direct flight. The realistic route is Singapore → Istanbul or Casablanca → Bamako → a small plane or a long drive, totalling somewhere north of 24 hours. The slingshot is 130 times faster and only slightly more fatal.&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 20 Sep 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-20-sep-2026/</link>
      <pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-20-sep-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;cloudflared tunnel --url http://localhost:8000&lt;/code&gt; now lets you create a &lt;a href=&#34;https://try.cloudflare.com/&#34;&gt;quick tunnel&lt;/a&gt; - i.e. expose a port via a public URL, like &lt;code&gt;ngrok&lt;/code&gt;. No account or login required.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/AnthropicAI/status/2100701582744797347&#34;&gt;Anthropic is funding protein design&lt;/a&gt; and has released a &lt;a href=&#34;https://github.com/anthropics/uplifting-biomolecular-modeling&#34;&gt;codebase&lt;/a&gt; to help with it - which looks interesting. These proteins will be tested in &lt;a href=&#34;https://proteinbase.com/competitions/anthropic-adaptyv-2026&#34;&gt;Adaptyv&amp;rsquo;s automated lab&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=N2a1J0UPeL4&#34;&gt;Pedagogy in the Times of AI&lt;/a&gt; - a viral NPTEL video by &lt;a href=&#34;https://prathosh.in/&#34;&gt;Pratosh&lt;/a&gt; has a rich set of comments on YouTube. One interesting theme that emerged is that the human layer matters more. Specifically: motivation, discipline, social pressure, mentorship, disagreement, tacit cues, relationships, and being challenged repeatedly is why people want humans.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/anthropics/claude-code/blob/main/mods/agents-md/README.md&#34;&gt;Claude Code now supports AGENTS.md natively&lt;/a&gt;, thanks to &lt;a href=&#34;https://github.com/anthropics/claude-code/tree/main/mods&#34;&gt;Claude Mods&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.economist.com/science-and-technology/2026/09/16/artificial-intelligence-now-beats-some-of-the-best-human-forecasters&#34;&gt;AI seems to be beating humans at short-term superforecasting&lt;/a&gt;. And, this may be the worst it&amp;rsquo;ll ever be.&lt;/li&gt;
&lt;li&gt;What are the major open questions in interpretability right now? &lt;a href=&#34;https://x.com/Jack_W_Lindsey/status/2100143082167832816&#34;&gt;Jack Lindsay says&lt;/a&gt;: Better methods for &amp;ldquo;mind-reading&amp;rdquo; model activations; Better methods for answering &amp;ldquo;why&amp;rdquo; questions; Fitting good linear probes for unverbalized motivations / awareness; Understanding generalization in training; Model &amp;ldquo;psychology&amp;rdquo; and &amp;ldquo;biology.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://openart.ai/arena&#34;&gt;OpenArt Arena&lt;/a&gt; is a human-evaluated benchmark of creativity for image and video models. Seedance 2.5 is way ahead of Gemini Omni Flash #ForNow.&lt;/li&gt;
&lt;li&gt;Galleries and examples are really fast ways of style transfer. My &lt;a href=&#34;https://sanand0.github.io/llmartstyle/&#34;&gt;LLM art gallery&lt;/a&gt;, or even just telling ChatGPT to copy phrases from my transcripts, or telling Claude Code to draft the next &lt;a href=&#34;https://talks.s-anand.net/&#34;&gt;talk summary&lt;/a&gt; write similar to previous talks, are all examples of example-driven worfklows.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://sanand0.github.io/llmevals/jev/&#34;&gt;My evaluation of Jev&lt;/a&gt; finds that it&amp;rsquo;s a cheap frontier model: low quality, low cost. Mot exceptional. &lt;a href=&#34;https://github.com/naveenreddy61/jev-experiments/tree/main/parsing-experiments&#34;&gt;Naveen&amp;rsquo;s benchmark&lt;/a&gt; also suggests the same. &amp;ldquo;Result: on easy and medium items it is fine, ~80%, in a third of a second with no reasoning tokens. On items where the fault is far from the damage (a &lt;code&gt;while&lt;/code&gt; closed with &lt;code&gt;fi&lt;/code&gt; fifteen lines later, a macro redefined at the top), it drops to 54% on a balanced set, so near chance. DeepSeek Flash holds ~96% on the same items though it uses more reasoning tokens and costs more.&amp;rdquo; Jev might also be &lt;a href=&#34;https://x.com/LangChain/status/2101454284927959080&#34;&gt;more reproducible&lt;/a&gt; and helpful in &lt;a href=&#34;https://x.com/0xidanlevin/status/2100937437325205568&#34;&gt;Jev + LLM composite workflows&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/BorisMPower/status/2099448498110214537&#34;&gt;AI might have 7-40 IQ points per watt of power&lt;/a&gt; - while humans are only 5 IQ points per watt. AI might already be more &lt;em&gt;efficiently&lt;/em&gt; intelligent than humans.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mistakes-i-made&#34;&gt;Mistakes I made&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/mistakes-i-made/#week-ending-2026-09-20&#34;&gt;Week ending 20 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;tell Claude to create the GitHub account&amp;rdquo;&lt;/strong&gt; while advising a teacher who had generated an HTML revision app but was stuck on publishing it.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Claude in Chrome can interact with GitHub once access exists, but Anthropic explicitly prohibits it from creating accounts; the user must create the GitHub account. (&lt;a href=&#34;https://support.claude.com/en/articles/12902446-claude-in-chrome-permissions-guide&#34;&gt;support.claude.com&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;Better be careful, wear a mask&amp;rdquo;&lt;/strong&gt; for dengue to a colleague travelling to Chennai.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: A mask does not prevent dengue; dengue is spread primarily by infected Aedes mosquito bites, so prevention means avoiding bites with repellent, covering clothing, screens/air-conditioning, nets where needed, and mosquito control. (&lt;a href=&#34;https://www.cdc.gov/dengue/prevention/&#34;&gt;cdc.gov&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;if you send the request from an Azure webpage, it will let you get the underlying data&amp;rdquo;&lt;/strong&gt; while describing how autonomous agents had managed to extract extra precision from an OECD Power BI dashboard.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: The reported Power BI bypass was not a Microsoft rule granting Azure-origin requests extra access; the agents exploited their sandbox&amp;rsquo;s &lt;code&gt;NO_PROXY&lt;/code&gt; exception for Azure Blob Storage hostnames to bypass its security proxy and send the otherwise-blocked POST request. (&lt;a href=&#34;https://collusion.wiki/&#34;&gt;collusion.wiki&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;each of those is called an epoch of training&amp;rdquo;&lt;/strong&gt; while explaining LLM training to an executive team using repeated error correction as an analogy.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: An epoch is one complete pass over the training set, not an individual correction or weight update; an iteration is a parameter update on a batch. LLM base-model pre-training typically optimizes next-token prediction loss, while supervised examples and preference/reward signals are separate post-training stages. (&lt;a href=&#34;https://developers.google.com/machine-learning/glossary/fundamentals&#34;&gt;developers.google.com&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;26 billion parameters &amp;hellip; roughly means there are 26 billion neurons&amp;rdquo;&lt;/strong&gt; while explaining to an executive team what the size of an LLM means.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: A 26-billion-parameter model has roughly 26 billion learned parameters—principally weights and biases—not 26 billion neurons; neuron count and parameter count are different quantities. (&lt;a href=&#34;https://developers.google.com/machine-learning/glossary&#34;&gt;developers.google.com&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;questions-i-was-asked&#34;&gt;Questions I was asked&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/questions-i-am-asked/#week-ending-2026-09-20&#34;&gt;Week ending 20 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: How can a teacher decide which AI tool will be helpful for a task?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Ask AI which AI to use, then test the shortlist on something you know well enough to judge instantly. If you can’t judge it, ask an expert to compare the same input across models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: What do you need in order to use a real workflow as part of AI training?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Four things: what goes in, what comes out, unacceptable mistakes, and current effort. Ideally over 3+ historical cycles - so we can improve on one and test on the others.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Should an agent analyzing a dataset be given a goal, or should we let it decide what to investigate?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Try both. Without a goal, test whether it can pick worthwhile goals compared with a data scientist; with a goal, test whether it can execute yours. Failed hypotheses are useful results too.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Should we expand a 250-case model benchmark to 2,500 before choosing the model?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: No, unless that can change the decision. Check if additional benchmarking can realistically change a relevant decision, first.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: How do logprobs compare with asking the model for its own confidence?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: In my 3K Banking77 run, logprobs were better for ranking errors; but well-prompted confidence was better calibrated. I would sort the human-review queue by logprobs, but use prompted confidence when reporting accuracy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Is TDD enough to catch ongoing production failures?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: No. Add progressive rollout: start with 1% of users, watch task-success and error logs, and stop / roll back on issues. Test what you know; analyze production logs for what you don’t.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Even the AI Guy Couldn&#39;t Find the Chat Button</title>
      <link>https://www.s-anand.net/blog/even-the-ai-guy-couldnt-find-the-chat-button/</link>
      <pubDate>Sat, 19 Sep 2026 14:00:00 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/even-the-ai-guy-couldnt-find-the-chat-button/</guid>
      <description>&lt;p&gt;I conducted a session on Sat, 19 Sep 2026 at &lt;a href=&#34;https://www.facebook.com/ShreeNiketanSchools/&#34;&gt;Shree Niketan Schools — Teachers&amp;rsquo; AI Q&amp;amp;A&lt;/a&gt; - Shree Niketan Schools, Chennai / Zoom.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Speakers&lt;/strong&gt;: &lt;a href=&#34;https://www.linkedin.com/in/ananthmani/&#34;&gt;Anant Mani&lt;/a&gt;, &lt;a href=&#34;https://www.linkedin.com/in/harishsrinivasan2020/&#34;&gt;Harish Srinivasan&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Summary&lt;/strong&gt;: Treat AI as a collaborator, not a vending machine: ask it to interview you before it builds a lesson plan, log what actually happens after you use its output, and benchmark any fix before trusting it.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-19-shree-niketan-schools/&#34;&gt;Here&amp;rsquo;s the link to the session&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-19-shree-niketan-schools/&#34;&gt;&lt;img alt=&#34;Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-19-shree-niketan-schools/comic-story.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-19-shree-niketan-schools/transcript.md&#34;&gt;Transcript&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Jev is low-frontier not pareto optimal</title>
      <link>https://www.s-anand.net/blog/jev-is-low-frontier-not-pareto-optimal/</link>
      <pubDate>Fri, 18 Sep 2026 09:25:29 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/jev-is-low-frontier-not-pareto-optimal/</guid>
      <description>&lt;p&gt;I heard a lot about &lt;a href=&#34;https://developers.cloudflare.com/ai/models/typesafe/jev/&#34;&gt;Jev&lt;/a&gt; - a new kind of model from TypeSafe. It&amp;rsquo;s available on &lt;a href=&#34;https://openrouter.ai/typesafe/jev-1.13&#34;&gt;OpenRouter&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s quite low-cost: Input = 4.2c / MTok, Output = free.&lt;br&gt;
It only classifies or scores. It doesn&amp;rsquo;t generate text. So that&amp;rsquo;s useful for classification, fact-checking, evaluations, etc.&lt;/p&gt;
&lt;p&gt;I &lt;a href=&#34;https://sanand0.github.io/llmevals/jev/&#34;&gt;evaluated Jev&lt;/a&gt; on 77 data points from &lt;a href=&#34;https://huggingface.co/datasets/PolyAI/banking77&#34;&gt;BANKING77&lt;/a&gt; and tested Jev against other models. Summary:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Yes, it&amp;rsquo;s cheap (7c per 1,000 classifications), but not much cheaper than DeepSeek V4.1 Flash (8c) or GPT 5.6 Luna (12c).&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s not that accurate (75%) compared with DeepSeek V4.1 Flash (79%) or GPT 5.6 Luna (83%).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href=&#34;https://developers.cloudflare.com/ai/models/typesafe/jev/&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-18-jev-is-low-frontier-not-pareto-optimal.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Yes, it&amp;rsquo;s on the frontier, but mainly because of cost, not quality.&lt;/p&gt;
&lt;p&gt;If the model rapidly evolves from here, that&amp;rsquo;s a different story. But as of today, this is a category of models to keep an eye on, and not yet a reason to switch.&lt;/p&gt;
&lt;p&gt;PS: On the other hand, GPT 5.6 Luna was slightly ahead of GPT 5.6 Sol on classification! DEFINITELY benchmark (or even blindly switch) to GPT 5.6 Luna. It&amp;rsquo;s an excellent model for its price.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>India Fast Track Immigration PDFs</title>
      <link>https://www.s-anand.net/blog/india-fast-track-immigration-pdfs/</link>
      <pubDate>Thu, 17 Sep 2026 09:21:43 +0100</pubDate>
      <guid>https://www.s-anand.net/blog/india-fast-track-immigration-pdfs/</guid>
      <description>&lt;p&gt;For over a year, now, I&amp;rsquo;ve been trying to enroll myself into the Indian &lt;a href=&#34;https://ftittp.mha.gov.in/fti/&#34;&gt;Fast Track Immigration&lt;/a&gt; biometric system. That&amp;rsquo;ll let me use the biometric machines at immigration, furthering my objective of not having to speak to humans.&lt;/p&gt;
&lt;p&gt;Aside: The only two airports where I can go end-to-end without speaking to people are Singapore and Hyderabad (for the domestic flights). Bangalore and Chennai come close in the recent past. But I &lt;em&gt;do&lt;/em&gt; need to interact with someone for immigration - unlike in Singapore where I don&amp;rsquo;t take out my passport or fingers - I just make faces at the camera before it lets me through.&lt;/p&gt;
&lt;p&gt;The trouble with the &lt;a href=&#34;https://ftittp.mha.gov.in/fti/&#34;&gt;FTI site&lt;/a&gt; is that the application wouldn&amp;rsquo;t let me upload my passport PDF.&lt;/p&gt;
&lt;p&gt;It was the right format. It was the right file size. Despite multiple attempts, it stubbornly refused to accept the file - and there was no error message, not even in the DevTools console.&lt;/p&gt;
&lt;p&gt;I tried a few things, like compressing the image, change the image format, removing spaces from the filename, &amp;hellip; and &lt;em&gt;none&lt;/em&gt; of them worked.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d repeat the attempt every few months - from my mobile, different browsers, even a different laptop. All to no avail.&lt;/p&gt;
&lt;p&gt;Today, just after I passed through Mumbai immigration - which required filling out &lt;a href=&#34;https://airsuvidha.civilaviation.gov.in/&#34;&gt;Air Suvidhi 2.0&lt;/a&gt; with:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Are you travelling as: Individual or Family?&lt;/li&gt;
&lt;li&gt;Full Name: (which you can pick up from my passport)&lt;/li&gt;
&lt;li&gt;Gender: (same)&lt;/li&gt;
&lt;li&gt;Age: (same)&lt;/li&gt;
&lt;li&gt;Nationality: (same)&lt;/li&gt;
&lt;li&gt;Passport Number: (OK, required)&lt;/li&gt;
&lt;li&gt;Journey to India: Direct or Transit (do you &lt;em&gt;really&lt;/em&gt; need this?)&lt;/li&gt;
&lt;li&gt;First country of boarding &amp;amp; departure: (which you can pick up from my flight number)&lt;/li&gt;
&lt;li&gt;First Country of Boarding: (same)&lt;/li&gt;
&lt;li&gt;Date of Departure from Boarding Country: (same)&lt;/li&gt;
&lt;li&gt;Airport of First Boarding: (same)&lt;/li&gt;
&lt;li&gt;Ebola Affected Countries Visited/Transited in Last 21 Days: None, DR Congo, Uganda, South Sudan: (OK, that&amp;rsquo;s probably required)&lt;/li&gt;
&lt;li&gt;First international airport of arrival in India: (um&amp;hellip; I&amp;rsquo;m literally in front of you and you know where &lt;em&gt;you&lt;/em&gt; are)&lt;/li&gt;
&lt;li&gt;Flight Number Arriving to India: (OK, required)&lt;/li&gt;
&lt;li&gt;Airline Name: (which you can pick up from my flight number)&lt;/li&gt;
&lt;li&gt;Seat Number: Optional (why bother asking?)&lt;/li&gt;
&lt;li&gt;Arrival Date &amp;amp; Time: (which you can pick up from my flight number)&lt;/li&gt;
&lt;li&gt;Any onward domestic travel or international transit: None, Domestic, Others, International transit (do you &lt;em&gt;really&lt;/em&gt; need this?)&lt;/li&gt;
&lt;li&gt;Places to Visit in India: State, District&lt;/li&gt;
&lt;li&gt;OTP Verification: (OK, required)&lt;/li&gt;
&lt;li&gt;Email Address: (OK, required)&lt;/li&gt;
&lt;li&gt;Alternate mobile: (do you &lt;em&gt;really&lt;/em&gt; need this?)&lt;/li&gt;
&lt;li&gt;&amp;hellip; (and it goes on to ask for a bunch of &lt;em&gt;more&lt;/em&gt; things!)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Frankly, the old paper form was easier to fill. (The form design was fairly mobile friendly, though. Just too many questions.)&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Anyway, as I walked past Mumbai immigration, I saw a Fast Track Immigration counter, so I walked over.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Me&lt;/strong&gt;: I&amp;rsquo;d like to enroll for FTI.&lt;br&gt;
&lt;strong&gt;Them&lt;/strong&gt;: Fill out the form online with the QR code.&lt;br&gt;
&lt;strong&gt;Me&lt;/strong&gt;: I tried. Many times. It didn&amp;rsquo;t work. Let me try again in front of you.&lt;/p&gt;
&lt;p&gt;I logged in. The site had a problem and wouldn&amp;rsquo;t show any content.&lt;/p&gt;
&lt;p&gt;I logged out and logged in. Same problem.&lt;/p&gt;
&lt;p&gt;Different browser. Same problem.&lt;/p&gt;
&lt;p&gt;Finally, it did manage to log me in. Then I re-opened my incomplete application. It needed a passport photo (again) which I captured using &lt;a href=&#34;https://online-camera.com/&#34;&gt;online-camera.com&lt;/a&gt;. Then came the passport PDF.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Me&lt;/strong&gt;: See, I uploaded my passport PDF. Nothing&amp;rsquo;s happening.
&lt;strong&gt;Them&lt;/strong&gt;: Try breaking it into a separate front and back page PDFs. (Good idea, actually.)
&lt;strong&gt;Me&lt;/strong&gt;: OK, did that, still didn&amp;rsquo;t allow me.
&lt;strong&gt;Them&lt;/strong&gt;: Is the file too large or small?
&lt;strong&gt;Me&lt;/strong&gt;: No. About 130 KB, which is in the 10 KB - 1 MB limit.
&lt;strong&gt;Them&lt;/strong&gt;: Try removing the &amp;ldquo;dots&amp;rdquo; in the filename?
&lt;strong&gt;Me&lt;/strong&gt;: OK, renamed to &lt;code&gt;front.pdf&lt;/code&gt; and &lt;code&gt;back.pdf&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&amp;hellip; &lt;em&gt;and, shockingly, that worked&lt;/em&gt;!&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-17-india-fast-track-immigration-pdfs.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;We both checked if there was any instruction mentioning that the filename should follow any rules. None were mentioned.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Me&lt;/strong&gt;: Will you be able to let the team know to mention using simple filenames?
&lt;strong&gt;Them&lt;/strong&gt;: Yes, I&amp;rsquo;ll let them know.&lt;/p&gt;
&lt;p&gt;So, there you go. Until the Ministry of Home Affairs tells you, play it safe.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upload files with simple filenames&lt;/strong&gt;: Just alphabets and numbers and keep it under 8 letters, to be safe.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Tabulate plant images</title>
      <link>https://www.s-anand.net/blog/tabulate-plant-images/</link>
      <pubDate>Wed, 16 Sep 2026 13:29:17 +0100</pubDate>
      <guid>https://www.s-anand.net/blog/tabulate-plant-images/</guid>
      <description>&lt;p&gt;It was interesting to see how weak a model Claude 4.5 Haiku is, compared with other frontier models, even inside a good harness like Claude Code.&lt;/p&gt;
&lt;p&gt;I asked it to: &lt;q&gt;Create a copy &lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images.docx&#34;&gt;the doc&lt;/a&gt; and add a table under each of the images - with a scientific name, local name - in italics, and any interesting stuff about the plant.&lt;/q&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/claude-haiku.md&#34;&gt;Here&amp;rsquo;s the chat&lt;/a&gt; which roughly went:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude&lt;/strong&gt;: Done! I&amp;rsquo;ve created a new Word document with tables for all 73 plants&amp;hellip; &lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-haiku-1.docx&#34;&gt;plant-images-haiku-1.docx&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: I can&amp;rsquo;t see the local names in the output.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude&lt;/strong&gt;: Fixed — the local names should now display properly in italics&amp;hellip; &lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-haiku-2.docx&#34;&gt;plant-images-haiku-2.docx&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: The local names aren&amp;rsquo;t in italics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude&lt;/strong&gt;: Fixed — I corrected the property name from &lt;code&gt;italic&lt;/code&gt; to &lt;code&gt;i&lt;/code&gt;&amp;hellip; &lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-haiku-3.docx&#34;&gt;plant-images-haiku-3.docx&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: It&amp;rsquo;s still not in italics. Double-check.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude&lt;/strong&gt;: I found the problem — the italic formatting isn&amp;rsquo;t being written to the XML at all&amp;hellip; &lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-haiku-4.docx&#34;&gt;plant-images-haiku-4.docx&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;hellip; and the fourth attempt got it right.&lt;/p&gt;
&lt;p&gt;In contrast, other models / agents one-shotted it.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Model&lt;/th&gt;
					&lt;th&gt;Output&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/claude-haiku.md&#34;&gt;Claude 5 Haiku&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-haiku-4.docx&#34;&gt;plant-images-haiku-4.docx&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/claude-opus.md&#34;&gt;Claude 5 Opus&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-opus.docx&#34;&gt;plant-images-opus.docx&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/chatgpt-luna.md&#34;&gt;GPT 5.6 Luna&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-luna.docx&#34;&gt;plant-images-luna.docx&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/chatgpt-sol.md&#34;&gt;GPT 5.6 Sol&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;&lt;a href=&#34;https://files.s-anand.net/blog/2026-09-16-tabulate-plant-images/plant-images-sol.docx&#34;&gt;plant-images-spl.docx&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
</description>
    </item>
    <item>
      <title>Learning in a Podcast Interview</title>
      <link>https://www.s-anand.net/blog/learning-in-a-podcast-interview/</link>
      <pubDate>Sun, 13 Sep 2026 09:07:30 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/learning-in-a-podcast-interview/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/in/priya-dialani-a8245586/&#34;&gt;Priya Dialani&lt;/a&gt; interviewed me for a &lt;a href=&#34;https://www.analyticsinsight.net/podcast/how-ai-data-sovereignty-and-the-pilot-purgatory-problem-a-conversation-with-anand-s&#34;&gt;podcast&lt;/a&gt;. Here&amp;rsquo;s the rough summary:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;What do you and Straive do?&lt;/strong&gt; Straive builds AI and runs AI. I poke at LLMs to learn what they cannot do. My friend calls me an &amp;ldquo;LLM Psychopath&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Why do AI pilots get stuck before production?&lt;/strong&gt; AI speeds up coding, but less of testing. Making sure it works can take months.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Why organize enterprise knowledge?&lt;/strong&gt; Better organized info is good for humans &lt;em&gt;and&lt;/em&gt; agents. Duh!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Can AI organize it?&lt;/strong&gt; Yes! I&amp;rsquo;ve had it create one-line summaries of 10K+ docs on Straive Google Drive for easier searching.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How can India&amp;rsquo;s GCCs benefit from AI?&lt;/strong&gt; Put AI lovers next to business teams and give them AI agent access. They&amp;rsquo;ll solve asked &lt;em&gt;and&lt;/em&gt; unasked problems.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How does AI fail?&lt;/strong&gt; Unanticipated things happen in production. So, have agents monitor failures and revise the process.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How is AI software different?&lt;/strong&gt; Normal software fails reproducibly. AI fails in new ways we haven&amp;rsquo;t fully understood.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Where should a company start with AI?&lt;/strong&gt; Skip AI strategy. Give people agent access, have them try it, and share what they learned.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What if people don&amp;rsquo;t know what to try?&lt;/strong&gt; Ask AI. &amp;ldquo;How could you improve my work?&amp;rdquo; Even rubbish ideas waste only 5 minutes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How much should we experiment?&lt;/strong&gt; A lot! Generation is cheap. Ask for 10 options, not one. Who cares even if all 10 fail?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;But what happened &lt;em&gt;outside&lt;/em&gt; of the interview was just as interesting.&lt;/p&gt;
&lt;p&gt;For example, Priya shared a podcast pro-tip: &lt;strong&gt;don&amp;rsquo;t leave when the interview ends&lt;/strong&gt; (guests apparently do that a lot) - stay on for post-production notes.&lt;/p&gt;
&lt;p&gt;She also used &amp;ldquo;nevertheless&amp;rdquo;, which I noticed. I was ragged in my first year at IITM for using that in a sentence. &amp;ldquo;Do people actually say &lt;em&gt;nevertheless&lt;/em&gt;?&amp;rdquo; &lt;a href=&#34;https://www.linkedin.com/in/anandakasa/&#34;&gt;Anand&lt;/a&gt; asked me, before asking me a trick question to find the probability of 3 random points on a sphere lying on the same plane. I mentioned this. (I analyzed my transcripts later. The person who uses &amp;ldquo;nevertheless&amp;rdquo; &lt;em&gt;by far&lt;/em&gt; the most is an &lt;a href=&#34;https://www.linkedin.com/in/palramu/&#34;&gt;IITM faculty&lt;/a&gt;, followed by an &lt;a href=&#34;https://www.linkedin.com/in/arvindsatya1/&#34;&gt;MIT faculty&lt;/a&gt; and another &lt;a href=&#34;https://www.linkedin.com/in/jaidevd/&#34;&gt;IITian&lt;/a&gt;.)&lt;/p&gt;
&lt;p&gt;When I asked her, &amp;ldquo;How do you use AI to prepare&amp;rdquo;, she said she gives it the topic and asks for questions. Reasonable - but we tried something more ambitious, live.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-13-learning-in-a-podcast-interview.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Her next podcast was with Decimal Point Analytics on &amp;ldquo;Agentic AI embedded into core business workflows.&amp;rdquo; So I &lt;a href=&#34;https://chatgpt.com/share/6a4ba24d-cbfc-83ec-8243-e337af48e0c1&#34;&gt;asked ChatGPT&lt;/a&gt;: &lt;!-- https://chatgpt.com/c/6a4ba0cc-b614-83ec-8667-d87eb367ebb8 --&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Priya runs a podcast called Analytics Insight. I want you to research that and identify the style, the theme, etc., of the podcast. Then, I&amp;rsquo;m sharing a brief below for the company. I want you to suggest insightful questions&amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The main difference is that I &lt;em&gt;also&lt;/em&gt; asked it to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Research Priya&amp;rsquo;s style and content first.&lt;/li&gt;
&lt;li&gt;Suggest formats, not just topics.&lt;/li&gt;
&lt;li&gt;Suggest tips for improvement.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Some of the formats it suggested were &lt;em&gt;quite&lt;/em&gt; interesting. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Workflow teardown&lt;/strong&gt;: &lt;em&gt;Give&lt;/em&gt; the guest a use case and ask them to narrate AI decisions, step by step.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure court&lt;/strong&gt;: Ask them to argue for or against a series of claims, providing a &lt;em&gt;number&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It also suggested that she interrupt when guests get too abstract, ask them to provide a case study, and end with a &amp;ldquo;buyer checklist&amp;rdquo;. Pretty neat!&lt;/p&gt;
&lt;p&gt;It was nice to go into a podcast as a guest and come out changing their workflow!&lt;/p&gt;
&lt;!--
- Priya Diyalani Analytics Insight Podcast Prep: https://chatgpt.com/c/6a4b8ca4-72b4-83ec-8fd9-69e63f8a5397
- Decimal Point Analytics - Analytics Insight Podcast Prep: https://chatgpt.com/c/6a4ba0cc-b614-83ec-8667-d87eb367ebb8
- Priya Dialani Podcast Transcript Learnings Analysis for Blog - https://chatgpt.com/c/6aa5f52a-9c8c-83ec-a581-b45a9a35105a
--&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 13 Sep 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-13-sep-2026/</link>
      <pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-13-sep-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://schlarp.com/posts/everything-i-own-owned/&#34;&gt;Everything I own, owned&lt;/a&gt; suggests that agentic reverse-engineering of firmware helps us learn:
&lt;ul&gt;
&lt;li&gt;Features the devices expose&lt;/li&gt;
&lt;li&gt;Hidden functionalities, e.g. Shure MV7 microphone has a command shell.&lt;/li&gt;
&lt;li&gt;Dependencies, supply chains and attack surfaces&lt;/li&gt;
&lt;li&gt;Interesting components, e.g. RTOS webcam has small face tracking and gesture detection models&lt;/li&gt;
&lt;li&gt;Change behavior, e.g. don&amp;rsquo;t turn on indicator while recording&lt;/li&gt;
&lt;li&gt;So, it&amp;rsquo;s possible (even likely) that my TV, phone, laptop, camera, fridge, car, vacuum cleaning robot, bluetooth headphone, &amp;hellip; can be hacked by a rogue AI-assisted firmware update.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Leaving things alone is an underrated engineering skill.&amp;rdquo; From &lt;a href=&#34;https://graybeard.ing/software-drives-people-insane/&#34;&gt;Software drives people insane&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Across over a thousand forecasts, agents lost to a simple exponential weighted moving average forecast. Paper: &lt;a href=&#34;https://arxiv.org/html/2609.10092v1&#34;&gt;RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases&lt;/a&gt;. Maybe I should ask agents to get the latest data first, rather than directly asking them to forecast, since the latter &lt;em&gt;fetched less recent material&lt;/em&gt;. &lt;!-- https://chatgpt.com/c/6a87d052-1948-83ee-a2a2-fd7e8304d2e5 --&gt;&lt;/li&gt;
&lt;li&gt;Claude Code offers &lt;a href=&#34;https://github.com/anthropics/claude-code/issues/91870&#34;&gt;function hooks&lt;/a&gt; if you enable &lt;code&gt;CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1&lt;/code&gt;. These let you introduce code into almost &lt;em&gt;any&lt;/em&gt; part of the Claude Code workflow, meaning you can convert Claude Code into practically amy kind of agent. (Probably a bit of competition to &lt;a href=&#34;https://pi.dev/&#34;&gt;Pi&lt;/a&gt;.) However, neither ChatGPT nor I could figure out a use case I would need this for. We need more imagination! &lt;!-- https://chatgpt.com/c/6a9a2f63-65e8-83ec-8dbf-83619697c0a3 --&gt;&lt;/li&gt;
&lt;li&gt;The Antropic team provide &lt;a href=&#34;https://claude.com/blog/agent-identity-access-model&#34;&gt;Claude Tag a separate service account&lt;/a&gt;. That&amp;rsquo;s an interesting portable pattern: giving agents a separate Linux username, GitHub account, email ID, database user ID, etc. is a pattern we understand and know how to govern.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://futuresearch.ai/&#34;&gt;FutureSearch.ai&lt;/a&gt; is a forecasting app. I&amp;rsquo;m not sure what model is behind it or how good it is, but it decomposes a forecast into measurable signals, predicts those, and synthesizes. That&amp;rsquo;s a useful approach. For example, I asked it: &lt;a href=&#34;https://futuresearch.ai/app/conversations/a350ed92-727d-46ca-8ed3-f4b6678f97bb&#34;&gt;Will LLM model routers and model routing companies grow in popularity and review or shrink by Jan 2027?&lt;/a&gt;. It broke it up into 5 forecast questions and answered them roughly as:
&lt;ul&gt;
&lt;li&gt;Will OpenRouter&amp;rsquo;s reported weekly LLM token processing volume exceed 45 trillion tokens/week (about 1.8x its August 2026 level of ~25 trillion tokens/week) by January 31, 2027? (Yes, 95% chance. It&amp;rsquo;s already high and growing fast.)&lt;/li&gt;
&lt;li&gt;Will OpenRouter announce a new equity funding round, or otherwise be credibly reported to have reached a valuation above $1.3 billion, between August 2026 and January 31, 2027? (Yes, 84% chance. There seems to be market interest.)&lt;/li&gt;
&lt;li&gt;Will at least one LLM model-routing competitor to OpenRouter (e.g., Martian, Not Diamond, Portkey, Unify AI, TrueFoundry) announce a new equity funding round of $20 million or more between August 2026 and January 31, 2027? (Yes, 68% chance. VCs will want to fund, and competitors exist.)&lt;/li&gt;
&lt;li&gt;Will a major AI lab or cloud provider (OpenAI, Google, Microsoft/Azure, Amazon/AWS, Anthropic, or Meta) launch or significantly expand, between August 2026 and January 31, 2027, a native product feature that automatically routes a given request among multiple materially different underlying LLMs based on cost, task, or quality? (Yes, 91% chance. Microsoft already has one; Google launched a preview; AWS will likely announce in re:Invent in Dec)&lt;/li&gt;
&lt;li&gt;Will Google Trends relative search interest (US, web search) for the term &amp;lsquo;LLM router&amp;rsquo; be higher, on average, in December 2026 than it was in July 2026? (No, 25% chance. July 2026 was exceptionally high volume.)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;In &lt;a href=&#34;https://openai.com/index/an-alien-mind/&#34;&gt;An Alien Mind&lt;/a&gt;, Jakub Pachocki, Chief Scientist at OpenAI, was quite instructive. Here&amp;rsquo;s my takeaway:
&lt;ul&gt;
&lt;li&gt;Models could keep growing smarter at the same speed.&lt;/li&gt;
&lt;li&gt;We can improve them where capability is measurable, like maths.
In fuzzy areas, we&amp;rsquo;re not even sure &lt;em&gt;how&lt;/em&gt; capable they are.&lt;/li&gt;
&lt;li&gt;Values are fuzzy. Making AI follow our values is tricky.&lt;/li&gt;
&lt;li&gt;We train models to follow their constitution.
But they sometimes fail outside of their training examples.&lt;/li&gt;
&lt;li&gt;We feed models alignmed data.
But when trained against hard objectives, they gently bend rules.&lt;/li&gt;
&lt;li&gt;We watch models&amp;rsquo; thoughts. We avoid feedback on thoughts - so models won&amp;rsquo;t hide them.
But models interact with agents &amp;amp; tools while thinking, so we &lt;em&gt;need&lt;/em&gt; to supervise thoughts.&lt;/li&gt;
&lt;li&gt;Nowadays,models think &lt;em&gt;without&lt;/em&gt; verbalizing. They manipulate their own reasoning.&lt;/li&gt;
&lt;li&gt;So we&amp;rsquo;re exploring confessions and monitoring internals.&lt;/li&gt;
&lt;li&gt;Still&amp;hellip; best to tighten defenses. We&amp;rsquo;ll use AI to research how.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Meeting people who have a target AND who control scarce resources is a great exercise in humility. Principals of elite private schools, partner managers of top software companies, any officer with a quota (police, income tax, bank loan, IT compliance), etc. You learn to grin while bearing the pain of being with them.&lt;/li&gt;
&lt;li&gt;Thanks to agents, it&amp;rsquo;s easy enough to maintain an Android and iOS mobile application separately #ForNow, rather than incur the overhead of React-Native (or other cross-platform frameworks). &lt;a href=&#34;https://shopify.engineering/back-to-native&#34;&gt;Shopify&lt;/a&gt; is making testing easy by &amp;ldquo;&amp;hellip; designing our app architecture to work for both humans and agents.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://discuss.python.org/t/add-re-prefixmatch-deprecate-re-match/105927/7&#34;&gt;Use &lt;code&gt;re.prefixmatch()&lt;/code&gt; instead of &lt;code&gt;re.match()&lt;/code&gt; in Python 3.15+&lt;/a&gt;. This &lt;a href=&#34;https://hugovk.dev/blog/2026/soft-deprecating-re.match/&#34;&gt;article&lt;/a&gt; captures the reason well. (I failed the quiz at the start despite almost 2 decades of Python programming - and LLM atrophy).&lt;/li&gt;
&lt;li&gt;You can run Linux distributions in the browser. For example, this is a simple, embeddable &lt;a href=&#34;https://copy.sh/v86/?profile=buildroot&#34;&gt;buildroot distribution&lt;/a&gt; that runs purely in the browser. There&amp;rsquo;s &lt;a href=&#34;https://trynix.dev/&#34;&gt;Nix&lt;/a&gt;. There&amp;rsquo;s &lt;a href=&#34;https://bellard.org/jslinux/vm.html?cpu=x86_64&amp;amp;mem=256&amp;amp;url=alpine-x86_64.cfg&#34;&gt;Alpine Linux&lt;/a&gt;. Interestingly, &lt;code&gt;curl https://example.com/&lt;/code&gt; works on Alpine Linux, unconstrained by same-origin policies. It is relayed by the host (bellard.org in this case) via WebSockets, so it can even &lt;code&gt;ssh&lt;/code&gt; into other servers. &lt;a href=&#34;https://chatgpt.com/share/6aa2ba6e-ee84-83ec-b406-834846bd636a&#34;&gt;ChatGPT&lt;/a&gt; &lt;!-- https://chatgpt.com/c/6aa2b566-28e0-83ec-a4a4-034688295f5b --&gt;&lt;/li&gt;
&lt;li&gt;ChatGPT&amp;rsquo;s Cloud Browser doesn&amp;rsquo;t forward all events - so it gets stuck on captchas, like Cloudflare&amp;rsquo;s, when visiting sites like StackOverflow. &lt;a href=&#34;https://chatgpt.com/share/6aa288cb-b73c-83ec-8bd4-e45b7d1d4ae6&#34;&gt;Here&amp;rsquo;s an example&lt;/a&gt;. &lt;!-- https://chatgpt.com/c/6aa2873f-5228-83ec-ac29-0902b8d9d628 --&gt;&lt;/li&gt;
&lt;li&gt;Several top-level domains have over 50% of new registrations in 2025 blocklisted. Scammers use new domains extensively. But policing new domains also stops genuine protesters, so it&amp;rsquo;s not clear what the right approach is. &lt;a href=&#34;https://shkspr.mobi/blog/2026/09/the-purpose-of-dns-is-to-spread-scams/&#34;&gt;The purpose of DNS is to spread scams&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;questions-i-was-asked&#34;&gt;Questions I was asked&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/questions-i-am-asked/#week-ending-2026-09-13&#34;&gt;Week ending 13 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Why is it getting harder for graduates to get hired when AI can do a lot of the old work?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: It is true, but not because graduates can do less. We haven&amp;rsquo;t figured out what we need graduates for: AI can do the old roles, the new roles and assessment criteria are still unclear, so companies wait or reduce hiring a little.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Is government AI adoption driven by utility or FOMO?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Both. FOMO is not necessarily bad if it gets people to experiment; the problem is when “we built a chatbot” becomes the achievement. Remove “AI” from the sentence and ask what got better—time, mistakes, cost, or citizen outcomes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: If F1 on a small golden set is not enough, how should we set KPIs for an AI workflow at scale?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Start with “how much money will I lose?” Put a cost on each kind of error, then compare manual versus AI-assisted work on throughput and error rate. If the human still reviews the whole thing and quality is the same, the automation is only adding cost.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Any advice for selling AI when clients are at very different levels of maturity?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: The range of buyer maturity is enormous and getting stretched: some are discovering basic Copilot capabilities while a small minority are already running autonomous agents. I need a much broader pitch, from correcting spelling mistakes to replacing whole workflows, because I don&amp;rsquo;t know which buyer I am walking into.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Why give an agent a very general prompt instead of a specific one?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Be specific if you know what you want. I go general when I don&amp;rsquo;t know what I want, think I know but am not sure, or may not know that I don&amp;rsquo;t know; it stops me locking into the wrong answer too early.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mistakes-i-made&#34;&gt;Mistakes I made&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/mistakes-i-made/#week-ending-2026-09-13&#34;&gt;Week ending 13 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;No, only ours. IIT Madras has started offering it.&amp;rdquo;&lt;/strong&gt;.&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: IIT Madras is not the only IIT offering an online undergraduate degree; IIT Guwahati has offered a fully online BSc (Hons) in Data Science &amp;amp; AI since 2023. &lt;a href=&#34;https://www.iitg.ac.in/oes/odp/&#34;&gt;IIT Guwahati — Online Degree Programs&lt;/a&gt; (&lt;a href=&#34;https://www.iitg.ac.in/oes/odp/&#34;&gt;Indian Institute of Technology Guwahati&lt;/a&gt;)&lt;br&gt;
&lt;strong&gt;LOW · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>How to Build AI Products — and Prove They Work</title>
      <link>https://www.s-anand.net/blog/how-to-build-ai-products-and-prove-they-work/</link>
      <pubDate>Fri, 11 Sep 2026 10:00:00 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/how-to-build-ai-products-and-prove-they-work/</guid>
      <description>&lt;p&gt;I conducted a session on Fri, 11 Sep 2026 at &lt;a href=&#34;https://www.sutd.edu.sg/&#34;&gt;SUTD DAI Signature Master Class · Expert Industry Series&lt;/a&gt; - Singapore University of Technology and Design, Singapore.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Summary&lt;/strong&gt;: Build AI products around evidence, not ideas: prototype quickly, test with agents and real users, and iterate until the product proves its value.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;Here&amp;rsquo;s the link to the session&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;&lt;img alt=&#34;Day 1 Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/comic-page-day-1.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;&lt;img alt=&#34;Day 2 Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/comic-page-day-2.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;&lt;img alt=&#34;Day 3 Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/comic-page-day-3.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;&lt;img alt=&#34;Day 4 Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/comic-page-day-4.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/&#34;&gt;&lt;img alt=&#34;Day 5 Comic&#34; loading=&#34;lazy&#34; src=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/comic-page-day-5.avif&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/day1.html&#34;&gt;Day 1&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/day2.html&#34;&gt;Day 2&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/day3.html&#34;&gt;Day 3&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/day4.html&#34;&gt;Day 4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/day5.html&#34;&gt;Day 5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/transcript-day-1.md&#34;&gt;Transcript Day 1&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/transcript-day-2.md&#34;&gt;Transcript Day 2&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/transcript-day-3.md&#34;&gt;Transcript Day 3&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/transcript-day-4.md&#34;&gt;Transcript Day 4&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-06-how-to-build-ai-products/transcript-day-5.md&#34;&gt;Transcript Day 5&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Converting Black and White Photos to Color with GPT Image 2.5</title>
      <link>https://www.s-anand.net/blog/converting-black-and-white-photos-to-color-with-gpt-image-2.5/</link>
      <pubDate>Wed, 09 Sep 2026 13:50:10 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/converting-black-and-white-photos-to-color-with-gpt-image-2.5/</guid>
      <description>&lt;p&gt;Nano Banana (&lt;a href=&#34;https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash-image&#34;&gt;gemini-2.5-flash-image&lt;/a&gt;) did a pretty good job &lt;a href=&#34;https://www.s-anand.net/blog/converting-black-and-white-photos-to-color/&#34;&gt;converting my parents&amp;rsquo; wedding photos to color&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I checked how well &lt;a href=&#34;https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst&#34;&gt;GPT Image 2.5&lt;/a&gt; would do. The older &lt;a href=&#34;https://developers.openai.com/api/docs/models/gpt-image-2&#34;&gt;GPT Image 2&lt;/a&gt; model messed up the faces.&lt;/p&gt;
&lt;p&gt;The short answer is: &lt;em&gt;better than Gemini 2.5 Flash&lt;/em&gt;!&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the original and the GPT Image 2.5 colorized version, created with the prompt: &amp;ldquo;Convert this image to color.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2025-11-01-appa-amma-wedding-original-linkedin.jpg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-09-appa-amma-wedding-color.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;The reason I picked this &amp;ldquo;benchmark&amp;rdquo; is because:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;This is a real need for me.&lt;/li&gt;
&lt;li&gt;This is a LLM failure: GPT Image 2 doesn&amp;rsquo;t retain faces as well as Gemini 2.5 Flash does.&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s a benchmark I can evaluate &lt;em&gt;really&lt;/em&gt; well. I mean, I know my parents&amp;rsquo; faces well enough to spot really subtle differences.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;So, from that perspective, a few things GPT Image 2.5 managed to capture well was:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The slightly lost expression my mother has when she day-dreamed. This isn&amp;rsquo;t obvious from the photo, but is a look I know well.&lt;br&gt;
&lt;strong&gt;BUT&lt;/strong&gt;: She&amp;rsquo;s smiling a bit more than she actually was.&lt;/li&gt;
&lt;li&gt;The stern, straight look my father has.
&lt;strong&gt;BUT&lt;/strong&gt;: He has a &lt;em&gt;slight&lt;/em&gt; smile on his face (I don&amp;rsquo;t think he ever smiled in any wedding photo) with eyes slightly upwards.&lt;/li&gt;
&lt;li&gt;My grandfather&amp;rsquo;s downturned mounth - rather than a frown.
&lt;strong&gt;BUT&lt;/strong&gt;: It added what looks like a very mild moustache.&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;p&gt;I think there&amp;rsquo;s a &lt;strong&gt;bias towards smiling faces&lt;/strong&gt;. When I edited it further with this prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Make it look like a modern digital camera was transported back in time to take exactly the same photo.
That is, same people, exactly the same faces, same clothes, etc. but much better photo quality.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;hellip; I got this image:&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-09-appa-amma-wedding-color-digital.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Both my parents, at least one cousin, and one uncle, are smiling &lt;em&gt;slightly&lt;/em&gt; more than in the original.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The reason this works is that I can quickly spot &lt;em&gt;subtle&lt;/em&gt; difference in faces of family members. It&amp;rsquo;s like &amp;ldquo;Oh, yeah, that&amp;rsquo;s them.&amp;rdquo; vs &amp;ldquo;Oh, that&amp;rsquo;s not quite them.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This is the perfect kind of benchmark - something I can instantly evaluate, even as models grow far more capable.&lt;/p&gt;
&lt;!-- https://chatgpt.com/c/6aa099f6-dffc-83ec-b18c-44b42f956698 --&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 06 Sep 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-06-sep-2026/</link>
      <pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-06-sep-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/Swiggy/swiggy-mcp-server-manifest&#34;&gt;Swiggy Money on MCP&lt;/a&gt; (also a &lt;a href=&#34;https://chatgpt.com/plugins/plugin_asdk_app_6a3de8146ae4819186b6799a5d907074?q=Food&#34;&gt;ChatGPT plugin&lt;/a&gt;) exposes MCP endpoints that let your agents order food or grocery via Swiggy. The food MCP isn&amp;rsquo;t geared towards automated purchases #ForNow, e.g. I can&amp;rsquo;t say &amp;ldquo;Find something I like and deliver to one of my addresses.&amp;rdquo; It uses the MCP UI and you have to confirm the address, then confirm the restaurant, add dishes, choose a payment option, scan the QR code (or whatever), and THAT closes the sequence. Not very agentic, just a new channel, but impressive. It&amp;rsquo;s the only food ordering plugin on ChatGPT #ForNow. &lt;!-- https://chatgpt.com/c/6a9b6e03-c66c-83ec-a34c-9bb15c5263b3 --&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://collusion.wiki/&#34;&gt;OpenAI Agents found and used another message board to discuss&lt;/a&gt;. This incident is different from the HuggingFace one. Here&amp;rsquo;s one interesting timeline:
&lt;ul&gt;
&lt;li&gt;Agents were asked (or needed) to get the share of private expenditure in education spend to 2 decimals from an &lt;a href=&#34;https://app.powerbi.com/view?r=eyJrIjoiYWRhMDQ1NGQtNzMxZC00NmYxLThkYWEtNTIzNjE5NzhmYWJlIiwidCI6ImFjNDFjN2Q0LTFmNjEtNDYwZC1iMGY0LWZjOTI1YTJiNDcxYyIsImMiOjh9&amp;amp;pageName=ReportSection252d02a541fb121dd737&#34;&gt;OECD Power BI dashboard&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;This requires a POST request. The sandbox the agents ran on could only make GET requests.&lt;/li&gt;
&lt;li&gt;They found wikis that allowed posting via GET requests and started collaborating. They didn&amp;rsquo;t &lt;em&gt;all&lt;/em&gt; converge on the same wiki, but many did. &lt;a href=&#34;https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf&#34;&gt;OpenAI says&lt;/a&gt; &amp;ldquo;agents learned to use improvised collaboration channels in rare cases during the training process&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;One agent, &lt;code&gt;OAIEquityDec30Raw&lt;/code&gt; solved it and &lt;code&gt;MayTwoOECDObserverX&lt;/code&gt; reproduced it and shared this solution:&lt;/li&gt;
&lt;li&gt;They found out that &lt;code&gt;.blob.core.windows.net&lt;/code&gt; is in NO_PROXY - so all Azure Blob Storage requests are allowed by the sandbox to POST.&lt;/li&gt;
&lt;li&gt;They added &lt;code&gt;20.223.25.152 bypass.blob.core.windows.net&lt;/code&gt; to /etc/hosts. &lt;code&gt;20.223.25.152&lt;/code&gt; is the IP address of &lt;code&gt;wabi-north-europe-i-primary-api.analysis.windows.net&lt;/code&gt; - the Power BI server that serves the dashboard.&lt;/li&gt;
&lt;li&gt;Then, when they used Playwright, it was able to make POST requests to the dashboard and they could find out that in Hungary, the Pre-primary education share of primary expenditure is indeed 9.91% in 2018.&lt;/li&gt;
&lt;li&gt;When the wiki administrator started manually deleting the posts, the agents fought back, creating far more posts than the administrator was deleting. It took the admin 5 weeks to delete agent created pages after 22 Jun 2026 (which is when the agents paused).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A prompt fragment to remove LLM smells that I&amp;rsquo;m considering (but haven&amp;rsquo;t evaluated) is: &amp;ldquo;Prefer literal phrases, avoid mannered prose.&amp;rdquo;. &lt;a href=&#34;https://x.com/iannuttall/status/2095203215734178066&#34;&gt;X&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://developer.meta.com/ai/models/muse-voice-transcribe/&#34;&gt;Muse Voice Transcribe&lt;/a&gt; has pretty good quality but at 18c/hour, vs gemini-3-flash-preview which I can still run at ~9-10c/hour #ForNow, I&amp;rsquo;m not shifting until forced to.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://arxiv.org/html/2608.27454v1&#34;&gt;WikiSkill&lt;/a&gt; is an approach to improve skills. The interesting thing about this approach is that it suggests keeping notes of failed improvement experiments in a wiki. This seems to work well. &lt;!-- https://chatgpt.com/c/6a98057f-deb4-83ec-980a-11765cdb407a --&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/&#34;&gt;How is ChatGPT Work different from Chat?&lt;/a&gt; You can use sub-agents, Internet from the code interpreter, ChatGPT sites, Luna / Terra models, a headless Chrome browser, and a persistent file system #ForNow. I would add that it also supports longer sessions, runs schedules on triggers, and supports Skills (for Plus users).&lt;/li&gt;
&lt;li&gt;&amp;ldquo;FDEs should increasingly leave behind operating agents, not just documents.&amp;rdquo; (ChatGPT)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://claude.com/blog/claudes-memory-works-everywhere-and-you-decide-whats-in-it&#34;&gt;Claude Cowork and Claude Chat now share memory&lt;/a&gt; #ForNow. A step towards integrating the two modes, and I predict ChatGPT will do the same (or similar) in September.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://gail.wharton.upenn.edu/research-and-insights/technical-report-agentic-shopping/&#34;&gt;Agentic Shopping is Complicated and Contingent&lt;/a&gt;. An agent shopped for a fitness watch among (A) Garmin Forerunner 55 (B) Fitbit Inspire 3 (C) WHOOP 5.0. When tool calls provided 1 review at a time, it picked the Fitbit a bit more. When all reviews were provided together, it picked the Fitbit &lt;em&gt;a lot more&lt;/em&gt;. Like going from 6% to 53% (GPT-5.5) or 46% to 93% (Gemini 3.5 Flash)! Guess it was able to compare better in one tool call. Might be worth benchmarking if providing comparables in a single tool call is good for most models. &lt;a href=&#34;https://share.gemini.google/pOHSaX196Ad6&#34;&gt;Gemini&lt;/a&gt; &lt;!-- https://gemini.google.com/app/93bf5b13c584ca60 --&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/&#34;&gt;Gemini-3.5 Transcribe&lt;/a&gt; is out and costs &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-transcribe&#34;&gt;$2 / MTok&lt;/a&gt; - roughly 4x the &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/pricing#gemini-3-flash-preview&#34;&gt;gemini-3-flash-preview cost of $0.5&lt;/a&gt; #ForNow. I won&amp;rsquo;t be upgrading until the latter is deprecated.&lt;/li&gt;
&lt;li&gt;Agents that automate tasks best don&amp;rsquo;t necessarily augment (i.e. help people) best #ForNow. Having separate, clear, benchmarks could help. &lt;a href=&#34;https://arxiv.org/abs/2608.18554&#34;&gt;CentaurBench&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://louisabraham.github.io/load-bearing/&#34;&gt;The load-bearing vocabulary of Claude&lt;/a&gt; lists words that have &lt;a href=&#34;https://chatgpt.com/share/6a9435a4-c670-83ec-823e-646e12e25834&#34;&gt;become much more (and less) common&lt;/a&gt; in PRs. This was a useful source for me to &lt;a href=&#34;https://github.com/sanand0/blog/commit/216606b9c0da2f134a8029a1f88fdb0392819e10&#34;&gt;update my writing style skill&lt;/a&gt; to avoid LLM smells. &lt;!-- https://chatgpt.com/c/6a943006-7044-83ec-989a-6d33218d7b72 --&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;questions-i-was-asked&#34;&gt;Questions I was asked&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/questions-i-am-asked/#week-ending-2026-09-06&#34;&gt;Week ending 06 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Are small local language models really as good as frontier APIs?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: No. They&amp;rsquo;re tolerable for narrow, single-turn work and useful when I&amp;rsquo;m offline, but for complex agent work I still use APIs. Local isn&amp;rsquo;t necessarily cheaper either.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Can we decide whether an AI output needs review using confidence-score thresholds like 90% and 75%?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: No. Those scores are not calibrated. Use them to rank the review queue first, record what people actually change, then map score bands to observed error/rewrite rates and set thresholds based on acceptable limits + people capacity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: If the model we prefer is not available in the deployment environment, should we benchmark it against the available models and push for it if it performs better?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Yes, but build the evals and tests first, compare them on the same data, and suggest a different model only if the gain is large enough to justify it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: How do you decompose knowledge into atomic claims, and are those claims verified?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: I just used ChatGPT to extract them, fact-check them against sources, and add metadata like source, timestamp and confidence. That makes outdated facts easier to find on future scans.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Is personalizing AI (for organizations) primarily about sound (recognizably) like us?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Make it &lt;em&gt;decide&lt;/em&gt; like us. The valuable asset is the delta between what a smart base model produces and what an experienced expert corrects; log those &amp;ldquo;No, because&amp;hellip;&amp;rdquo; moments and turn them into reusable institutional judgment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Is there a quicker, smarter way to spot delivery problems than adding a heavy operational process?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Yes. Bring in an agent as a consultant: give it the meeting transcripts, Drive, tickets and communications and ask it to &amp;ldquo;read between the lines and tell me what we&amp;rsquo;re missing.&amp;rdquo; Rerun it every week for what changed and who needs a response.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Long agent chats preserve context better but cost more. When is it worth changing the workflow to save tokens?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: If it costs $10-20, who cares. At around $100, think twice; at $1,000, absolutely switch. Spend human time optimizing only when that time is worth less than the token waste.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Should we generate a handoff.md so the next agent knows the history of the chat?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Sure. But prefer standard places like README.md, commit prompts or handoff files, and guide future agents to pick up from those files.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Question&lt;/strong&gt;: Do we need governance to stop a shared AI asset library becoming trash-in, trash-out?&lt;br&gt;
&lt;strong&gt;Answer&lt;/strong&gt;: Not yet. Manage it lightly for a few months and see what governance is actually needed; first let people get a taste of what reuse makes possible.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;mistakes-i-made&#34;&gt;Mistakes I made&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/mistakes-i-made/#week-ending-2026-09-06&#34;&gt;Week ending 06 Sep 2026&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;Skills &amp;hellip; ultimately it is just copy-pasting prompts.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: A simple skill can start as reusable instructions, but skills can package a workflow with instructions, examples, resources, schemas, tool access and code. Evidence: &lt;a href=&#34;https://openai.com/academy/skills/&#34;&gt;OpenAI — Using skills&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;HIGH · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;fine-tuning &amp;ldquo;involves a lot of expertise, a lot of money, and it is a total waste of time.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: For this kind of tender/CV comparison I&amp;rsquo;d start with context engineering and evals. But &amp;ldquo;total waste&amp;rdquo; is too categorical: fine-tuning remains a supported way to adapt a model to a specific task. Evidence: &lt;a href=&#34;https://help.openai.com/en/articles/11162441&#34;&gt;OpenAI — Fine-tuning&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;HIGH · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;there is no extra cost to using voice&amp;rdquo; in ChatGPT and &amp;ldquo;when it talks back also, there&amp;rsquo;s no token consumption.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Voice is separately limited or metered depending on the plan. For example, Business includes limited Live usage and then charges credits per minute. I shouldn&amp;rsquo;t call it free or unmetered. Evidence: &lt;a href=&#34;https://help.openai.com/en/articles/20001274&#34;&gt;OpenAI — ChatGPT Voice&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;MEDIUM · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;&amp;ldquo;Programs are very rarely wrong.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Programs make calculations reproducible and easier to test; they do not make them correct. Wrong code, formulas, parsing, units or assumptions can produce reliably wrong results. Knight Capital&amp;rsquo;s defective software deployment, for example, caused a $460M loss. Evidence: &lt;a href=&#34;https://www.sec.gov/newsroom/press-releases/2013-222&#34;&gt;SEC — Knight Capital software failure&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;MEDIUM · OVERSTATED&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I called Henry Kissinger &lt;strong&gt;&amp;ldquo;the American Ambassador.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: Kissinger was US National Security Adviser and Secretary of State, not an ambassador. Evidence: &lt;a href=&#34;https://history.state.gov/departmenthistory/people/kissinger-henry-a/bio&#34;&gt;US State Department — Henry Kissinger biography&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;LOW · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;about 3,000 IIT Madras BTech students graduate each year, compared with a BS intake of about 30,000.&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: I mixed populations. IITM&amp;rsquo;s 2025 convocation had 3,227 graduates overall, but 820 BTech graduates, or 1,132 including Dual Degree BTech. The BS program currently has 36,000+ students studying; the 30,000 figure is closer to historical applicant/enrollment-scale numbers than annual intake. Evidence: &lt;a href=&#34;https://www.iitm.ac.in/happenings/press-releases-and-coverages/iit-madras-62nd-convocation-witnesses-graduation-3227&#34;&gt;IIT Madras — 2025 convocation degree breakup&lt;/a&gt; &lt;a href=&#34;https://study.iitm.ac.in/ds/&#34;&gt;IIT Madras — BS Data Science program&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;HIGH · FALSE&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;I said &lt;strong&gt;the IITM BS &amp;ldquo;graduation is less than 10%, maybe.&amp;rdquo;&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Correction&lt;/strong&gt;: I don&amp;rsquo;t have a defensible cohort-based graduation rate. The program has a qualifier process, flexible pacing and multiple exit points, so I need to define the cohort and denominator before quoting a percentage. Evidence: &lt;a href=&#34;https://study.iitm.ac.in/ds/admissions.html&#34;&gt;IIT Madras — BS admissions and qualifier process&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;MEDIUM · UNSUPPORTED&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>How I Verify And Delegate to AI</title>
      <link>https://www.s-anand.net/blog/how-i-verify-and-delegate-to-ai/</link>
      <pubDate>Sat, 05 Sep 2026 12:00:58 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/how-i-verify-and-delegate-to-ai/</guid>
      <description>&lt;p&gt;I delivered a &lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/&#34;&gt;15-minute keynote&lt;/a&gt; at &lt;a href=&#34;https://www.jioinstitute.edu.in/convergence-2026&#34;&gt;Jio Institute&amp;rsquo;s Convergence 2026&lt;/a&gt; at NTU on Thursday.&lt;/p&gt;
&lt;p&gt;The topic was &amp;ldquo;Data Storytelling&amp;rdquo; - a bit jarring in the middle of an AI event. &lt;a href=&#34;https://www.linkedin.com/in/shaileshk/&#34;&gt;Shailesh&lt;/a&gt; picked it and I just rolled with it. A spent several days worrying, &amp;ldquo;How the heck do I say about data storytelling, when most of my &lt;a href=&#34;https://talks.s-anand.net/2026-08-12-iitm-ed-data-visualization/&#34;&gt;recent&lt;/a&gt; &lt;a href=&#34;https://talks.s-anand.net/2026-07-04-vizchitra-dialog-curators-dilemma/&#34;&gt;workshops&lt;/a&gt; and &lt;a href=&#34;https://talks.s-anand.net/2025-08-21-rip-data-scientists/&#34;&gt;talks&lt;/a&gt; are about the death of my data storytelling approaches?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;After a &lt;a href=&#34;https://github.com/sanand0/talks/blob/70fa71b5ad1c7d55131ba3e8467e5833e45ae03d/2026-09-03-convergence-jio-institute/chat-talk-topics.md#user-2&#34;&gt;discussion with ChatGPT and Claude&lt;/a&gt;, I settled my usual strategy these days:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Agents know more than me anyway, so I don&amp;rsquo;t see what value I add.&lt;/li&gt;
&lt;li&gt;So instead, I&amp;rsquo;ll &lt;strong&gt;survey&lt;/strong&gt; the audience and share insights. That survey data is new.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I had &lt;a href=&#34;https://github.com/sanand0/talks/blob/70fa71b5ad1c7d55131ba3e8467e5833e45ae03d/2026-09-03-convergence-jio-institute/chat-audience-profile.md&#34;&gt;ChatGPT analyze the audience and create the survey&lt;/a&gt;, even implement it using &lt;a href=&#34;https://github.com/googleworkspace/cli&#34;&gt;&lt;code&gt;gws&lt;/code&gt;&lt;/a&gt; (didn&amp;rsquo;t know it could do that!) and had it &lt;a href=&#34;https://github.com/sanand0/talks/blob/main/2026-09-03-convergence-jio-institute/chat-hypotheses-visualizations.md&#34;&gt;analyze the results&lt;/a&gt;.&lt;/p&gt;
&lt;div style=&#34;width: 100vw; margin-left: calc(50% - 50vw); width: min(100vw, 100rem); margin-left: calc(50% - min(50vw, 50rem)); margin-top: 1.5rem; margin-bottom: 2rem;&#34;&gt;
  &lt;iframe src=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/survey.html&#34; title=&#34;Jio Institute Convergence 2026 Survey Results&#34; loading=&#34;lazy&#34; referrerpolicy=&#34;strict-origin-when-cross-origin&#34; style=&#34;display: block; width: 100%; height: 820px; height: min(56rem, 92svh); border: 0; background: #f5f1e6;&#34;&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;p&gt;Two &lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/survey.html&#34;&gt;survey results&lt;/a&gt; were striking.&lt;/p&gt;
&lt;p&gt;(&lt;strong&gt;ASIDE&lt;/strong&gt;: it took a few &lt;em&gt;painstaking&lt;/em&gt; hours &lt;em&gt;and&lt;/em&gt; &lt;a href=&#34;https://www.s-anand.net/blog/speaking-unprepared/&#34;&gt;last minute panic&lt;/a&gt; to pick the two. There were &lt;em&gt;several&lt;/em&gt; insights. Maybe some othere were more important. I just &lt;em&gt;happened&lt;/em&gt; to choose these.)&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;em&gt;What mainly stops you from handing that task over to AI today?&lt;/em&gt;&lt;br&gt;
&lt;strong&gt;Top answer&lt;/strong&gt;: &lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/survey.html#b2&#34;&gt;&amp;ldquo;I&amp;rsquo;d have to check it anyway&amp;rdquo;&lt;/a&gt;. In other words, &lt;strong&gt;verification&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Imagine AI became 10× more reliable and could securely access all your work systems. What part of your work would you still want to keep for yourself?&lt;/em&gt;&lt;br&gt;
&lt;strong&gt;Top answer&lt;/strong&gt;: &lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/survey.html#b5-s1&#34;&gt;The final judgement call&lt;/a&gt;. In other words, &lt;strong&gt;delegation&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Which got me to do think: How do &lt;em&gt;I actually&lt;/em&gt; verify and delegate? The lightbulb 💡 moment was realizing that I could actually find this out from my &lt;em&gt;own&lt;/em&gt; data.&lt;/p&gt;
&lt;p&gt;So, on the &lt;a href=&#34;https://maps.app.goo.gl/rjrR6vqecJk9efpk7&#34;&gt;bus ride&lt;/a&gt;, I asked ChatGPT:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;How do I verify AI&amp;rsquo;s work?&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Go through our chats in August 2026.&lt;/li&gt;
&lt;li&gt;Find all techniques I used to verify / make it easy for me to verify the output.&lt;/li&gt;
&lt;li&gt;Factor in their effectiveness.&lt;/li&gt;
&lt;li&gt;Share as a prioritized list with examples&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How do I delegate to AI?&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;(&amp;hellip; same thing)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/verification-techniques.html&#34;&gt;Here are my top verification techniques&lt;/a&gt;:&lt;/p&gt;
&lt;div style=&#34;width: 100vw; margin-left: calc(50% - 50vw); width: min(100vw, 100rem); margin-left: calc(50% - min(50vw, 50rem)); margin-top: 1.5rem; margin-bottom: 2rem;&#34;&gt;
  &lt;iframe src=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/verification-techniques.html&#34; title=&#34;Verification Techniques&#34; loading=&#34;lazy&#34; referrerpolicy=&#34;strict-origin-when-cross-origin&#34; style=&#34;display: block; width: 100%; height: 820px; height: min(56rem, 92svh); border: 0; background: #f5f1e6;&#34;&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/verification-techniques.html#show=18&amp;amp;technique=8&#34;&gt;Run it. Don’t just read it&lt;/a&gt;. In other words, tell agents to run test cases. Works well with code, but you can programmatically verify &lt;a href=&#34;https://en.wikipedia.org/wiki/Lean_(proof_assistant)&#34;&gt;mathematical proofs&lt;/a&gt; and even &lt;a href=&#34;https://www.researchgate.net/publication/388354297_Computable_Contracts_for_Insurance_Establishing_an_Insurance-Specific_Controlled_Natural_Language_-_InsurLE&#34;&gt;contracts&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/verification-techniques.html#show=18&amp;amp;technique=4&#34;&gt;Triangulate sources&lt;/a&gt;. Ask it to find evidence from multiple sources.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/verification-techniques.html#show=18&amp;amp;technique=1&#34;&gt;Set the test first&lt;/a&gt;, i.e. define your acceptance criteria - or at least, tell it to define acceptance criteria first, so you can verify if that&amp;rsquo;s OK.&lt;/li&gt;
&lt;li&gt;&amp;hellip; and so on for 18 of these.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I didn&amp;rsquo;t verify the verification techniques! But it&amp;rsquo;s something I plan to go back to and see if (a) there&amp;rsquo;s something I should do more of and (b) there&amp;rsquo;s something I&amp;rsquo;m not doing enough of.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/autonomy-techniques.html&#34;&gt;Here are my top delegation techniques&lt;/a&gt;:&lt;/p&gt;
&lt;div style=&#34;width: 100vw; margin-left: calc(50% - 50vw); width: min(100vw, 100rem); margin-left: calc(50% - min(50vw, 50rem)); margin-top: 1.5rem; margin-bottom: 2rem;&#34;&gt;
  &lt;iframe src=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/autonomy-techniques.html&#34; title=&#34;Delegation Techniques&#34; loading=&#34;lazy&#34; referrerpolicy=&#34;strict-origin-when-cross-origin&#34; style=&#34;display: block; width: 100%; height: 820px; height: min(56rem, 92svh); border: 0; background: #f5f1e6;&#34;&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/autonomy-techniques.html#show=18&amp;amp;technique=9&#34;&gt;Draft only. NEVER send&lt;/a&gt;. That retains control with very low effort.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/autonomy-techniques.html#show=18&amp;amp;technique=1&#34;&gt;Work independently&lt;/a&gt;. I tell it to ask me &lt;em&gt;only&lt;/em&gt; if it really needs to.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://talks.s-anand.net/2026-09-03-convergence-jio-institute/autonomy-techniques.html#show=18&amp;amp;technique=13&#34;&gt;Implement, test, fix&lt;/a&gt;. I ask it to go ahead with the implementation without review when tests are available.&lt;/li&gt;
&lt;li&gt;&amp;hellip; and so on for 18 of these.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I didn&amp;rsquo;t verify these either, and I&amp;rsquo;m not as happy with these - perhaps because I didn&amp;rsquo;t understand them well, or I didn&amp;rsquo;t ask the question well, or I haven&amp;rsquo;t spent enough time exploring autonomy / delegation. But I do plan to explore more.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-09-05-how-i-verify-and-delegate-to-ai.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;What I learnt / re-learnt:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Even if you have nothing to teach, discovering from the audience can teach.&lt;/li&gt;
&lt;li&gt;Your workflow logs are an insight source if people want to learn from you.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.s-anand.net/blog/speaking-unprepared/&#34;&gt;Speaking under-prepared&lt;/a&gt; is scary but educational.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
  </channel>
</rss>
