<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Daniel Lemire&#039;s blog</title>
	<atom:link href="https://lemire.me/blog/feed/" rel="self" type="application/rss+xml" />
	<link>https://lemire.me/blog</link>
	<description>Daniel Lemire is a software performance expert. He ranks among the top 2% of scientists globally (Stanford/Elsevier 2025) and is one of GitHub&#039;s top 1000 most followed developers.</description>
	<lastBuildDate>Thu, 08 Oct 2026 17:20:08 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://lemire.me/blog/wp-content/uploads/2026/09/cropped-IMG_9858-2-32x32.jpeg</url>
	<title>Daniel Lemire&#039;s blog</title>
	<link>https://lemire.me/blog</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Mastering SIMD with Java Vector API</title>
		<link>https://lemire.me/blog/2026/10/07/mastering-simd-with-java-vector-api/</link>
					<comments>https://lemire.me/blog/2026/10/07/mastering-simd-with-java-vector-api/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 02:07:05 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=23008</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/mastering-simd-java-vector-api-150x150.webp" class="webfeedsFeaturedVisual wp-post-image" alt="Cover of Mastering SIMD with Java Vector API by Roman Snytsar" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" />I write a lot about data parallelism. About how our processors have special instructions that can do several operations at once. They are often critical to get the best performance out of a CPU. Compilers are sometimes able to take advantage of these instructions. But it is hard to tell when it will work. For &#8230; <a href="https://lemire.me/blog/2026/10/07/mastering-simd-with-java-vector-api/" class="more-link">Continue reading <span class="screen-reader-text">Mastering SIMD with Java Vector API</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/mastering-simd-java-vector-api-150x150.webp" class="webfeedsFeaturedVisual wp-post-image" alt="Cover of Mastering SIMD with Java Vector API by Roman Snytsar" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" /><p>I write a lot about data parallelism. About how our processors have special instructions that can do several operations at once. They are often critical to get the best performance out of a CPU.</p>
<p>Compilers are sometimes able to take advantage of these instructions. But it is hard to tell when it will work.</p>
<p>For Java programmers, you can write directly for data parallelism since JDK 16. With enough skill, you might multiply the performance of a few key functions. </p>
<p>The Vector API is in the incubator module <code>jdk.incubator.vector</code>. It still not officially supported, but you can enable it by passing <code>--add-modules=jdk.incubator.vector</code>. I expect that it will soon become a mainstream feature in Java.</p>
<p>Roman Snytsar wrote a book about it: <em>Mastering SIMD with Java Vector API</em>, and the subtitle is <a href="https://www.amazon.com/dp/B0GPBQNTJ5"><em>Unlocking Single-Core Performance Through SIMD Optimization</em></a>. I had the honor of being the technical reviewer. I learned a few things while reading the book and I enjoyed it.</p>
<p>Who has time for books, especially technical books?</p>
<p>What a book offers is time to reflect. Reading technical books today is probably just as relevant as it ever was. The author takes you on a story&#8230; and gets you to think. </p>
<p><img decoding="async" class="alignnone size-large" alt="Cover of Mastering SIMD with Java Vector API by Roman Snytsar" src="https://lemire.me/blog/wp-content/uploads/2026/10/mastering-simd-java-vector-api.webp" /></p>
<p>Snytsar&#8217;s book works from  problems like sum an array, compute a mean and a standard deviation, remove duplicates, merge two sorted arrays, and so forth. A lot of them are similar to the problems you will encounter as a programmer.  He starts from the ordinary solution and rebuilds it with the Java Vector API. He then runs benchmarks. He looks at the assembly code.</p>
<p>The book covers performance issues such as unrolling, asymmetric loops, structural hazards, dependencies. Snytsar shows us that, in some instances, data parallelism can fail to bear fruits. The negative lessons are just as important as the positive ones.</p>
<p>The Java Vector API is a thick layer of abstractions, but Snytsar shows that some tricks work better on some hardware than others.  I liked the chapter on removing duplicates, which is built on <code>compress</code>. It is one instance where the specific hardware matters. And you are unlikely to just find out about it yourself.</p>
<p>It is a book worth buying if you are a Java programmer with a focus on performance. </p>
<p>The book: Roman Snytsar, <a href="https://www.amazon.com/dp/B0GPBQNTJ5"><em>Mastering SIMD with Java Vector API</em>,</a> Apress, 2026.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/07/mastering-simd-with-java-vector-api/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Faster software linking with mold</title>
		<link>https://lemire.me/blog/2026/10/06/linking-node-js-with-mold/</link>
					<comments>https://lemire.me/blog/2026/10/06/linking-node-js-with-mold/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 08:00:20 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=23017</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8cr2v98cr2v98cr2-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" />When you build a program, the compiler turns each source file into an object file. Then a linker stitches all the object files and libraries into one executable. On Linux, the default linker is usually GNU ld (also called bfd). The GNU binutils also ship gold, an alternative linker that was designed to be faster. &#8230; <a href="https://lemire.me/blog/2026/10/06/linking-node-js-with-mold/" class="more-link">Continue reading <span class="screen-reader-text">Faster software linking with mold</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8cr2v98cr2v98cr2-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>When you build a program, the compiler turns each source file into an object file. Then a linker stitches all the object files and libraries into one executable.</p>
<p>On Linux, the default linker is usually GNU ld (also called bfd). The GNU binutils also ship gold, an alternative linker that was designed to be faster. Gold was built a Google and first made available in 2008.</p>
<p>The <a href="https://github.com/rui314/mold">mold</a> linker is a newer linker written by Rui Ueyama, it was first released in 2021. It is meant as a drop-in replacement for GNU ld, and its main selling point is speed. The latest release is version 3.</p>
<p>How much faster is it on a large project? I built <a href="https://github.com/nodejs/node">Node.js</a> from its main branch on an Intel Xeon Gold 6548N server (Emerald Rapids, two sockets, 64 cores and 128 threads) running Linux with GCC 14.3. The final <code>node</code> executable weighs about 160 MB.</p>
<p>I captured the command that links the <code>node</code> executable and ran it six times with each linker. I report the median.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">linker</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">time to link node</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">GNU ld (bfd) 2.41</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">2.52 s</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">GNU gold 2.41</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">1.46 s</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">mold 3.0.0, 1 thread</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">0.49 s</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">mold 3.0.0, 8 threads</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">0.13 s</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">mold 3.0.0, 128 threads</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">0.11 s</td>
</tr>
</tbody>
</table>
<p>
<a href="https://lemire.me/blog/wp-content/uploads/2026/10/link_big4.png"><img loading="lazy" decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/10/link_big4-1024x614.png" alt="" width="660" height="396" class="alignnone size-large wp-image-23023" srcset="https://lemire.me/blog/wp-content/uploads/2026/10/link_big4-1024x614.png 1024w, https://lemire.me/blog/wp-content/uploads/2026/10/link_big4-300x180.png 300w, https://lemire.me/blog/wp-content/uploads/2026/10/link_big4-768x461.png 768w, https://lemire.me/blog/wp-content/uploads/2026/10/link_big4.png 1050w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>With multithreading, mold is about 24 times faster than GNU ld and about 14 times faster than gold.</p>
<p>Even if I restrict mold to a single thread, it links Node.js in half a second, five times faster than GNU ld. With only 8 threads, you get nearly all of the benefit.</p>
<p>Using mold is easy. You can pass <code>-fuse-ld=mold</code> to GCC or clang. Or you can wrap your whole build:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>mold<span style="color: #f8f8f8;"> </span>-run<span style="color: #f8f8f8;"> </span>make<span style="color: #f8f8f8;"> </span>-j128
</code></pre>
</div>
<p>The <code>-run</code> option intercepts every call to the default linker and redirects it to mold. You do not need to change the build scripts.</p>
<p>Of course, two seconds saved on a link does not matter much if you build Node.js once. But a developer who edits a file and rebuilds dozens of times a day pays the link cost each time. With mold, the link step becomes nearly free.</p>
<p>My scripts and raw results are <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/10/mold">available</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/06/linking-node-js-with-mold/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
			</item>
		<item>
		<title>Ephemeral testing</title>
		<link>https://lemire.me/blog/2026/10/05/ephemeral-testing/</link>
					<comments>https://lemire.me/blog/2026/10/05/ephemeral-testing/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 08:00:51 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=23013</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-150x150.jpg 150w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-300x300.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-768x768.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9.jpg 1024w" sizes="auto, (max-width: 150px) 100vw, 150px" />We have many ways to ensure software quality. Unit testing. Fuzz testing. Integration testing. And so forth.   I&#8217;d like to propose a method that was unthinkable before: ephemeral testing. (Ephemeral is a fancy word for &#8216;throw away&#8217; or &#8216;temporary&#8217;.)   You write your code. You build your software component. Or the AI agent does &#8230; <a href="https://lemire.me/blog/2026/10/05/ephemeral-testing/" class="more-link">Continue reading <span class="screen-reader-text">Ephemeral testing</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-150x150.jpg 150w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-300x300.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9-768x768.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/10/Gemini_Generated_Image_8mi98l8mi98l8mi9.jpg 1024w" sizes="auto, (max-width: 150px) 100vw, 150px" /><div class="" data-block="true" data-editor="51ceu" data-offset-key="9ps6d-0-0">
<div data-offset-key="9ps6d-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9ps6d-0-0">We have many ways to ensure software quality. Unit testing. Fuzz testing. Integration testing. And so forth.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="bkkb0-0-0">
<div data-offset-key="bkkb0-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="bkkb0-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="n0f7-0-0">
<div data-offset-key="n0f7-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="n0f7-0-0">I&#8217;d like to propose a method that was unthinkable before: </span><span data-offset-key="n0f7-0-1">ephemeral testing</span><span data-offset-key="n0f7-0-2">. (Ephemeral is a fancy word for &#8216;throw away&#8217; or &#8216;temporary&#8217;.)</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="6f3i2-0-0">
<div data-offset-key="6f3i2-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="6f3i2-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="5q75-0-0">
<div data-offset-key="5q75-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5q75-0-0">You write your code. You build your software component. Or the AI agent does it for you, it does not matter.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="43dkq-0-0">
<div data-offset-key="43dkq-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="43dkq-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="f23fg-0-0">
<div data-offset-key="f23fg-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="f23fg-0-0">Then you ask an AI agent to build on it: an application, another layer, maybe several. You have it test what it built. You do not assess the original work directly. You assess how good the software built on top of it is.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="9k6bp-0-0">
<div data-offset-key="9k6bp-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9k6bp-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="9m917-0-0">
<div data-offset-key="9m917-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9m917-0-0">It is a form of integration testing. The difference is that the software on top is entirely ephemeral. You throw it away when you are done.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="n0mf-0-0">
<div data-offset-key="n0mf-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="n0mf-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="5r99t-0-0">
<div data-offset-key="5r99t-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5r99t-0-0">A library with a clean API, stable invariants, and useful errors lets the agent produce something that works quickly. A library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures. The failures are evidence about your code, not about the agent.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="5klhj-0-0">
<div data-offset-key="5klhj-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5klhj-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="ac4qd-0-0">
<div data-offset-key="ac4qd-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="ac4qd-0-0">You can repeat it. Different agents, different tasks, same foundation.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="3bk3f-0-0">
<div data-offset-key="3bk3f-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="3bk3f-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="7uv4n-0-0">
<div data-offset-key="7uv4n-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="7uv4n-0-0">In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.</span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="3t8kv-0-0">
<div data-offset-key="3t8kv-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="3t8kv-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="2hmec-0-0">
<div data-offset-key="2hmec-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="2hmec-0-0">Of course, you could argue that with AI, you can rebuild everything whenever you need to. But that&#8217;s not practical. You need some form of stability. </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="c1lk0-0-0">
<div data-offset-key="c1lk0-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="c1lk0-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="51ceu" data-offset-key="a1c0t-0-0">
<div data-offset-key="a1c0t-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="a1c0t-0-0">I have been applying this trick to various projects. As I consider a new feature, I ask my AI to prototype quickly what I might later build based on what I am doing it. Ephemeral testing works for me thus far.</span></div>
</div>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/05/ephemeral-testing/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Research paper overload: submissions capped at two a month</title>
		<link>https://lemire.me/blog/2026/10/04/arxiv-capped-submissions-at-two-a-month/</link>
					<comments>https://lemire.me/blog/2026/10/04/arxiv-capped-submissions-at-two-a-month/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 17:07:27 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=23004</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/arxiv-limit-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Much of my research is on arXiv. I was one of the early adopters. It is simply a repository of research papers with PDF and metadata. It makes it convenient to find research without any paywall. Physicists started it in 1991. Math and computer science followed gradually. It relies on volunteers to moderate it, because &#8230; <a href="https://lemire.me/blog/2026/10/04/arxiv-capped-submissions-at-two-a-month/" class="more-link">Continue reading <span class="screen-reader-text">Research paper overload: submissions capped at two a month</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/arxiv-limit-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Much of my research is on <a href="https://arxiv.org">arXiv</a>. I was one of the early adopters. It is simply a repository of research papers with PDF and metadata.</p>
<p>It makes it convenient to find research without any paywall.</p>
<p>Physicists started it in 1991. Math and computer science followed gradually.</p>
<p>It relies on volunteers to moderate it, because we don&#8217;t want garbage or spam.</p>
<p>On October 1, they <a href="https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/">capped the submissions to two per person per month</a>.</p>
<p>Why?</p>
<p>arXiv got over 40,000 submissions in September.</p>
<p>Two years ago, it was about 20,000.</p>
<p>Doubling every two years. It is unsustainable for human beings.</p>
<p><img decoding="async" class="alignnone size-large" alt="Papers posted to arXiv each month" src="https://lemire.me/blog/wp-content/uploads/2026/10/arxiv-monthly.webp" /></p>
<p>In related news, <a href="https://x.com/GoogleVRP/status/2105689195180179605">Google stopped taking new bug reports in its open-source bounty program</a>. They could not cope with the influx.</p>
<p>The writing is on the wall.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/04/arxiv-capped-submissions-at-two-a-month/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Google&#8217;s first orbital data center is in orbit</title>
		<link>https://lemire.me/blog/2026/10/04/googles-first-orbital-data-center-is-in-orbit/</link>
					<comments>https://lemire.me/blog/2026/10/04/googles-first-orbital-data-center-is-in-orbit/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 16:49:29 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=23002</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/orbital-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Google&#8217;s first orbital data center is in orbit (October 1st). They say it is about the size of a fridge. It has one kilowatt of solar power. It is about the power of a household. It was launched on a Falcon 9 rocket (SpaceX). It should run for about a year. Google plans to launch &#8230; <a href="https://lemire.me/blog/2026/10/04/googles-first-orbital-data-center-is-in-orbit/" class="more-link">Continue reading <span class="screen-reader-text">Google&#8217;s first orbital data center is in orbit</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/orbital-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p><a href="https://blog.google/innovation-and-ai/models-and-research/google-research/project-suncatcher-prototype/">Google&#8217;s first orbital data center</a> is in orbit (October 1st). They say it is about the size of a fridge.</p>
<p>It has one kilowatt of solar power. It is about the power of a household.</p>
<p>It was launched on a Falcon 9 rocket (SpaceX).</p>
<p>It should run for about a year. Google plans to launch two more orbital data centres next year. The plan is to use laser communications.</p>
<p>For people complaining that the latency is going to be a problem with orbital data centres…</p>
<p>An orbital data centre can be effectively in your line of sight and it can be relatively close (much closer than the diameter of the USA).</p>
<p><img decoding="async" class="alignnone size-large" alt="Orbital latency compared with California to New York" src="https://lemire.me/blog/wp-content/uploads/2026/10/orbital-latency.webp" /></p>
<p><a href="https://lemire.me/blog/wp-content/uploads/2026/10/image.gif"><img loading="lazy" decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/10/image.gif" alt="" width="800" height="450" class="alignnone size-full wp-image-23028" /></a><br />
Further if you have a cluster of space data centres, they can remain close and in a line of sight.</p>
<p>So you are not going to run your video games or do high frequency trading from space, but for compute, latency is unlikely to be the bottleneck.</p>
<p>The great features are space (no need to wait five years for a permit) and power (no need to negotiate with a local government for electricity). Solar power in space is great because it runs 24h a day.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/04/googles-first-orbital-data-center-is-in-orbit/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>Parsing compressed JSON at 40 GB/s</title>
		<link>https://lemire.me/blog/2026/10/01/parsing-compressed-json-at-40-gb-s/</link>
					<comments>https://lemire.me/blog/2026/10/01/parsing-compressed-json-at-40-gb-s/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 03:10:29 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22991</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/jsonstream-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Parsing compressed JSON at 40 GB/s" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />A common way to store JSON data is to write one document per line. We call it NDJSON or JSON Lines. Log files, database exports and machine-learning datasets often come in this format. The files can be large, so we may compress them. {"id":1,"active":true,"user":{"name":"user_1","tags":["guest"]},"score":-3862,"note":"..."} {"id":2,"active":false,"user":{"name":"user_2","tags":["staff","admin"]},"score":8123,"note":"..."} How fast can you read such a file? I wrote &#8230; <a href="https://lemire.me/blog/2026/10/01/parsing-compressed-json-at-40-gb-s/" class="more-link">Continue reading <span class="screen-reader-text">Parsing compressed JSON at 40 GB/s</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/10/jsonstream-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Parsing compressed JSON at 40 GB/s" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>A common way to store JSON data is to write one document per line. We call it NDJSON or JSON Lines. Log files, database exports and machine-learning datasets often come in this format. The files can be large, so we may compress them.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000; font-weight: bold;">{</span><span style="color: #204a87; font-weight: bold;">"id"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"active"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #204a87; font-weight: bold;">true</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"user"</span><span style="color: #000; font-weight: bold;">:{</span><span style="color: #204a87; font-weight: bold;">"name"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #4e9a06;">"user_1"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"tags"</span><span style="color: #000; font-weight: bold;">:[</span><span style="color: #4e9a06;">"guest"</span><span style="color: #000; font-weight: bold;">]},</span><span style="color: #204a87; font-weight: bold;">"score"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #0000cf; font-weight: bold;">-3862</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"note"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #4e9a06;">"..."</span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">{</span><span style="color: #204a87; font-weight: bold;">"id"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #0000cf; font-weight: bold;">2</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"active"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #204a87; font-weight: bold;">false</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"user"</span><span style="color: #000; font-weight: bold;">:{</span><span style="color: #204a87; font-weight: bold;">"name"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #4e9a06;">"user_2"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"tags"</span><span style="color: #000; font-weight: bold;">:[</span><span style="color: #4e9a06;">"staff"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #4e9a06;">"admin"</span><span style="color: #000; font-weight: bold;">]},</span><span style="color: #204a87; font-weight: bold;">"score"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #0000cf; font-weight: bold;">8123</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #204a87; font-weight: bold;">"note"</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #4e9a06;">"..."</span><span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>How fast can you read such a file? I wrote a small demo with <a href="https://github.com/simdjson/simdjson">simdjson</a>. My test file has 5 million records: 812 MB of NDJSON. For each record, I read a few fields. I count the active records, and I sum the scores of the active records whose user has the <code>admin</code> tag.</p>
<p>You could decompress the whole file, but it might be better to process the compressed file. So you decompress a chunk, you parse the complete lines in that chunk, and you move the incomplete last line to the front of the buffer. In NDJSON, a newline can never appear inside a document: a newline in a string must be escaped as <code>n</code>. So you can always cut the buffer right after its last newline character. The simdjson library has a function for many documents in one buffer (<code>iterate_many</code>):</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">simdjson</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">ondemand</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">parser</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">parser</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">while</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87;">true</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// fill the buffer with decompressed bytes...</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">size_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">eof</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">?</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">len</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">last_newline</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">len</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">simdjson</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">ondemand</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">document_stream</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">stream</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">parser</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">iterate_many</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #000; font-weight: bold;">).</span><span style="color: #000;">get</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">stream</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">auto</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">doc</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">stream</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">accumulate</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">doc</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">value_unsafe</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">result</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">eof</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">break</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">memmove</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">buf</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">len</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">len</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cut</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>I ran my benchmarks on an Intel Xeon Gold 6548N server (Emerald Rapids) with two sockets, 64 cores and 128 threads. I use GCC 14.</p>
<p>A gzip file is one long compressed stream. The decompressor needs the previous 32 KiB of output to decode what comes next. So you must decompress the file from the start, with one thread. I can still parse in other threads: one thread decompresses chunks and puts them in a queue, and the other threads parse them. It doubles the speed to 2.5 GB/s.</p>
<p>Other formats do better. A zstd or an lz4 file can be made of many independent <em>frames</em>, one after the other. It is still a regular <code>.zst</code> or <code>.lz4</code> file: the usual command-line tools decompress it as usual.</p>
<p>My program writes a new frame every 256 KiB of JSON, always after a newline. Each frame stores its decompressed size and a checksum. To find where a frame ends, you only need to read a few block headers. You do not need to decompress anything. So the threads can take frames one by one, and each thread decompresses and parses its own frames.</p>
<p><img decoding="async" alt="Decompressing and parsing NDJSON" src="https://lemire.me/blog/wp-content/uploads/2026/10/jsonstream-formats.webp" class="alignnone size-large" /></p>
<p>With one thread, zstd and lz4 are no faster than gzip: about 1.1 GB/s. With 64 threads, I get 40 GB/s with zstd and 34 GB/s with lz4. That is 16 times faster than the best I can do with gzip.</p>
<p>The files are not much larger. Small frames compress a bit worse, since each frame starts from scratch, but the zstd file is only 6% larger than the gzip file.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">file</th>
<th style="text-align: right;">size</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">NDJSON</td>
<td style="text-align: right;">812.0 MB</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">gzip (one stream)</td>
<td style="text-align: right;">56.9 MB</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">zstd (256 KiB frames)</td>
<td style="text-align: right;">60.6 MB</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">lz4 (256 KiB frames)</td>
<td style="text-align: right;">108.5 MB</td>
</tr>
</tbody>
</table>
<p>My data is synthetic. The records are very repetitive, which makes decompression fast. Your data might decompress more slowly.</p>
<p>If you control how your JSON files are written, consider using zstd with many frames. You get files about as small as with gzip, and you can read them many times faster with multiple threads.</p>
<p><strong>Source code</strong>: <a href="https://github.com/simdjson/simdjson_compressed_demo">https://github.com/simdjson/simdjson_compressed_demo</a>. The benchmark numbers and the plotting script are in <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/jsonstream">my blog repository</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/10/01/parsing-compressed-json-at-40-gb-s/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Transcoding UTF-8 to UTF-16 with replacement at gigabytes per second</title>
		<link>https://lemire.me/blog/2026/09/30/transcoding-utf-8-to-utf-16-with-replacement-at-gigabytes-per-second/</link>
					<comments>https://lemire.me/blog/2026/09/30/transcoding-utf-8-to-utf-16-with-replacement-at-gigabytes-per-second/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 12:33:27 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22987</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/withreplacement-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Transcoding UTF-8 to UTF-16 with replacement" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Our software represents strings using the UTF-16 or the UTF-8 formats. Most text on the web is UTF-8, but Java, C# or JavaScript represents the strings as UTF-16 to the programmer. Sometimes we need to transcode (convert) strings. Your browser probably uses the simdutf library for validating or transcoding. It is part of the widely &#8230; <a href="https://lemire.me/blog/2026/09/30/transcoding-utf-8-to-utf-16-with-replacement-at-gigabytes-per-second/" class="more-link">Continue reading <span class="screen-reader-text">Transcoding UTF-8 to UTF-16 with replacement at gigabytes per second</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/withreplacement-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Transcoding UTF-8 to UTF-16 with replacement" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Our software represents strings using the UTF-16 or the UTF-8 formats. Most text on the web is UTF-8, but Java, C# or JavaScript represents the strings as UTF-16 to the programmer.</p>
<p>Sometimes we need to transcode (convert) strings. Your browser probably uses the simdutf library for validating or transcoding. It is part of the widely used V8 JavaScript engine.</p>
<p>Up until a few days ago, the simdutf library missed one key feature: if the UTF-8 input is invalid, it did not know how to transcode it to UTF-16. The objective is to replace ill-formed UTF-8 sequences with the replacement character U+FFFD. It is the character that shows up as a weird question mark sometimes. You may have seen it in a broken web site.</p>
<p>The new function <code>convert_utf8_to_utf16_with_replacement</code> always succeeds. It must be used in conjunction with the function <code>utf16_length_from_utf8_with_replacement</code>, which scans and validates the input, recording the offsets of errors if there are some. The usage is as follows.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">source</span><span style="color: #000; font-weight: bold;">[]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #4e9a06;">'c'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'a'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'f'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'\xff'</span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #204a87; font-weight: bold;">size_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">length</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000;">simdutf</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">utf8_to_utf16_result</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">res</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">simdutf</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">utf16_length_from_utf8_with_replacement</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">source</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">length</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">unique_ptr</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">char16_t</span><span style="color: #000; font-weight: bold;">[]</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">utf16</span><span style="color: #000; font-weight: bold;">{</span><span style="color: #204a87; font-weight: bold;">new</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">char16_t</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">res</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">count</span><span style="color: #000; font-weight: bold;">]};</span>
<span style="color: #204a87; font-weight: bold;">size_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">written</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">simdutf</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">convert_utf8_to_utf16_with_replacement</span><span style="color: #000; font-weight: bold;">(</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">source</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">length</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">utf16</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">get</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">res</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>Because we have a list of errors before even beginning the transcoding, the bytes between the recorded offsets are valid UTF-8, so we can transcode faster. Each recorded error becomes one U+FFFD.</p>
<p>I timed a scalar decoder against the library on Japanese Wikipedia. The file is 531 KB. I repeated it so the timed buffer is 1.1 MB. I use one core of a Xeon Gold 6548N. GCC 14.3.1 at <code>-O3</code>, best of eight runs.</p>
<p><img decoding="async" alt="UTF-8 to UTF-16 with replacement against a scalar decoder, Japanese Wikipedia" src="https://lemire.me/blog/wp-content/uploads/2026/09/withreplacement-japanese.webp" class="alignnone size-large" /></p>
<p>With no errors, simdutf reaches 7.0 GB/s and the scalar loop 1.7 GB/s.</p>
<p>On valid input, the new approach may even be marginally faster than the previous method, because our initial validation pass allows us to transcode faster afterward.</p>
<p><img decoding="async" alt="Valid UTF-8 to UTF-16, ordinary conversion and conversion with replacement" src="https://lemire.me/blog/wp-content/uploads/2026/09/withreplacement-valid.webp" class="alignnone size-large" /></p>
<p><strong>Compiler</strong>: GCC 14.3.1, <code>-O3</code>, one core (<code>taskset -c 2</code>), simdutf icelake kernel. The functions are on the branch <a href="https://github.com/simdutf/simdutf/tree/utf8-to-utf16-with-replacement"><code>utf8-to-utf16-with-replacement</code></a>.</p>
<p><strong>Credit</strong>: The most non-trivial part of this routine was coded by Benjamin Bucher over the summer.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/30/transcoding-utf-8-to-utf-16-with-replacement-at-gigabytes-per-second/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>How do you deal with people who strongly disagree with your views?</title>
		<link>https://lemire.me/blog/2026/09/30/how-do-you-deal-with-people-who-strongly-disagree-with-your-views/</link>
					<comments>https://lemire.me/blog/2026/09/30/how-do-you-deal-with-people-who-strongly-disagree-with-your-views/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 01:19:54 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22978</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/ohs7r-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />  1. Begin by genuinely listening to them. Make sure that you understand them. It is often useful to ask for confirmation: restate what you think they are saying and make sure that you understand. Physically, it helps to turn your shoulders so that you face them. Don&#8217;t look at your screen while they talk. &#8230; <a href="https://lemire.me/blog/2026/09/30/how-do-you-deal-with-people-who-strongly-disagree-with-your-views/" class="more-link">Continue reading <span class="screen-reader-text">How do you deal with people who strongly disagree with your views?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/ohs7r-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><div class="" data-block="true" data-editor="3lgi9" data-offset-key="37v46-0-0">
<div data-offset-key="37v46-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="9ph0p-0-0">
<div data-offset-key="9ph0p-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9ph0p-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="98jvh-0-0">
<div data-offset-key="98jvh-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="98jvh-0-0">1. Begin by genuinely listening to them. Make sure that you understand them. It is often useful to ask for confirmation: restate what you think they are saying and make sure that you understand. Physically, it helps to turn your shoulders so that you face them. Don&#8217;t look at your screen while they talk.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="4jbfc-0-0">
<div data-offset-key="4jbfc-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="4jbfc-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="dq3hi-0-0">
<div data-offset-key="dq3hi-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="dq3hi-0-0">2. Ask more questions. You can often effortlessly weaken the stance of someone who has not given much thought to an issue just by asking them questions. Don&#8217;t &#8216;challenge them&#8217;. Ask genuine questions. The objective is not to break their views with questions, but to establish the playing ground. Many people are unclear about what they think they believe. They may simply be repeated what they heard.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="5ahpf-0-0">
<div data-offset-key="5ahpf-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5ahpf-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="3aesg-0-0">
<div data-offset-key="3aesg-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="3aesg-0-0">3. Focus on what you share in common. Often, the disagreement is smaller than you think. Furthermore, you can sometimes entirely avoid a conflict by establishing strong agreement on many points.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="b6leb-0-0">
<div data-offset-key="b6leb-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="b6leb-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="22jv9-0-0">
<div data-offset-key="22jv9-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="22jv9-0-0">4. Unless you are clearly a knowledgeable person, do not immediately express your opinion. Begin by establishing your credibility. Refer to hard facts, data. Refer to people or documents you have consulted. This, in turn, should have been done ahead of time. Never engage in a debate without having done at least a bit of research and reflection.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="1vnie-0-0">
<div data-offset-key="1vnie-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="1vnie-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="bpsvg-0-0">
<div data-offset-key="bpsvg-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="bpsvg-0-0">5. Use language that fits the issue. You need to learn the vocabulary. Learn the technical terms and use them, although you may want to remain accessible, so explain them as well.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="5jluv-0-0">
<div data-offset-key="5jluv-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5jluv-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="ef9md-0-0">
<div data-offset-key="ef9md-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="ef9md-0-0">6. Tell stories. People understand by emotions as much as by their reason. Ideally, the story should be true and it should be backed by hard evidence.</span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="1bhuq-0-0">
<div data-offset-key="1bhuq-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="1bhuq-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="3lgi9" data-offset-key="3pr8k-0-0">
<div data-offset-key="3pr8k-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="3pr8k-0-0">7. Avoid personal attacks and do not tolerate them against you and others. If someone is insulting you, if it is fair to say so. You may say that you feel offended. It is the one case where you can and maybe should go on the offensive. Don&#8217;t answer by an insult, but attack the fact that they insulted you. You may ask for an apology.</span></div>
</div>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/30/how-do-you-deal-with-people-who-strongly-disagree-with-your-views/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>simdjson 5.0 is out</title>
		<link>https://lemire.me/blog/2026/09/28/simdjson-5-0-is-out/</link>
					<comments>https://lemire.me/blog/2026/09/28/simdjson-5-0-is-out/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 12:38:25 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22975</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/simdjson5-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="simdjson 5.0 is out" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />The simdjson library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, &#8230; <a href="https://lemire.me/blog/2026/09/28/simdjson-5-0-is-out/" class="more-link">Continue reading <span class="screen-reader-text">simdjson 5.0 is out</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/simdjson5-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="simdjson 5.0 is out" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>The <a href="https://github.com/simdjson/simdjson">simdjson</a> library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, and back, without writing any glue code.</p>
<p>Today we are releasing version 5.0.</p>
<ol>
<li>
<p>Static reflection is no longer guarded and is an officially supported feature. When your compiler has reflection enabled (e.g., <code>g++ -std=c++26 -freflection</code> with GCC 16), simdjson detects it by itself and the reflection-based functions become available.</p>
</li>
<li>
<p>When deserializing a C++ structure through reflection, simdjson now uses <em>key selectors</em> (see below) by default: it reads the object in a single pass, whatever the order of the keys. You can return to the previous approach (one lookup per member) with <code>-DSIMDJSON_DISABLE_KEY_SELECTOR_REFLECTION=1</code>.</p>
</li>
<li>
<p>A positive integer in <code>[2^64, 10^20)</code> is now reported as a big integer (<code>BIGINT_NUMBER</code>), like other overflowing integers, instead of a malformed number. When big integers are parsed as strings, a token such as <code>123456789123456789123x</code> is now rejected.</p>
</li>
</ol>
<p>One major change is the key selectors. A common task is to extract a few fields from a JSON object. In simdjson 5.0 (C++20 or better), you can name the keys at compile time and visit the object once:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">using</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">namespace</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">simdjson</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">auto</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">json</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">R"({ "name": "Daniel", "age": 42, "city": "Montreal" })"</span><span style="color: #000;">_padded</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000;">ondemand</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">parser</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">parser</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">auto</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">doc</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">parser</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">iterate</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">json</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">string_view</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">name</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">city</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">age</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">auto</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">result</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">doc</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">get_object</span><span style="color: #000; font-weight: bold;">().</span><span style="color: #000;">for_each</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #4e9a06;">"name"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"city"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"age"</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">name</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">city</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">age</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #8f5902; font-style: italic;">// name == "Daniel", city == "Montreal", age == 42</span>
</code></pre>
</div>
<p>The keys can appear in any order. At compile time, simdjson builds a perfect hash function for your set of keys. At run time, recognizing a key takes a hash computed from a couple of bytes and one comparison. You can also pass one callback per key instead of variables. The iteration stops as soon as all keys have been found.</p>
<p>We added annotations for the data structures for automated (C++26) serialization and deserialization: <code>rename</code>, <code>rename_all</code>, <code>alias</code>, <code>skip</code>, <code>default_value</code>, <code>flatten</code>, <code>deny_unknown_fields</code>, <code>transparent</code>, and so forth.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">struct</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">[[</span><span style="color: #a40000; border: 1px solid #EF2929;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #c4a000;">simdjson</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #c4a000;">rename_all</span><span style="color: #a40000; border: 1px solid #EF2929;">&lt;</span><span style="color: #c4a000;">simdjson</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #c4a000;">case_style</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #c4a000;">camel_case</span><span style="color: #a40000; border: 1px solid #EF2929;">&gt;</span><span style="color: #000; font-weight: bold;">]]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">User</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">string</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">first_name</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">int64_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">user_id</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000; font-weight: bold;">[[</span><span style="color: #a40000; border: 1px solid #EF2929;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #c4a000;">simdjson</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #c4a000;">rename</span><span style="color: #a40000; border: 1px solid #EF2929;">&lt;"</span><span style="color: #c4a000;">KEY</span><span style="color: #a40000; border: 1px solid #EF2929;">"&gt;</span><span style="color: #000; font-weight: bold;">]]</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">api_key</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">};</span>
<span style="color: #8f5902; font-style: italic;">// {"firstName":"Ann","userId":7,"KEY":8}</span>
</code></pre>
</div>
<p>We often get many JSON documents in one file or one network message. simdjson has long supported streams of documents separated by white space (NDJSON). In 5.0 we added:</p>
<ul>
<li><a href="https://www.rfc-editor.org/rfc/rfc7464">RFC 7464</a> JSON text sequences (each document preceded by the record separator character) and comma-separated documents;</li>
<li><code>stream_format::newline_delimited</code>: you promise that each document sits on its own line. When you only read part of a document, simdjson jumps to the next line instead of walking over the rest of the document.</li>
<li><code>simdjson::slice_at</code>, which cuts a stream into blocks at document boundaries so you can parse the blocks on as many threads as you like. The built-in threaded mode uses at most two threads.</li>
</ul>
<p>We also fixed several bugs in <code>document_stream</code>, found in an audit by Francisco Geiman Thiesen.</p>
<p>There are many other smaller features.</p>
<ul>
<li>NaN and infinity: JSON does not allow <code>NaN</code> or <code>Infinity</code>, but many systems produce them anyway. If you define <code>SIMDJSON_ENABLE_NAN_INF</code>, simdjson parses them, and serializes them.</li>
<li>Narrow types: <code>get_uint8()</code>, <code>get_int8()</code>, <code>get_uint16()</code>, <code>get_int16()</code> check the range for you. With C++23, <code>get_float32()</code> and <code>get_float64()</code> return <code>std::float32_t</code> and <code>std::float64_t</code>. The binary32 value is rounded once, directly from the decimal string, not through a double.</li>
<li>The DOM API can parse a buffer that has no padding (<code>parser.parse_unpadded(...)</code>). It is slower than the regular function, but it never reads past the end of your buffer and it does not copy your data.</li>
<li>With C++17, <code>simdjson::padded_input</code> adds padding only when it is needed: when your string ends near a page boundary.</li>
<li>C++20 ranges: you can pipe On-Demand arrays and objects into <code>std::views::transform</code> and other adaptors.</li>
<li>DOM arrays support reverse iteration (<code>rbegin()</code>, <code>rend()</code>) with no allocation.</li>
<li>On-Demand objects offer <code>get_current_position()</code> and <code>revert_position()</code>: if you miss an optional field, you can go back to where you were instead of rescanning the whole object.</li>
<li><code>char8_t</code> (<code>u8</code>) variants of the string accessors in C++20.</li>
<li>Better pretty printing with the <a href="https://github.com/j-brooke/FracturedJson">FracturedJson</a> style, including tables.</li>
<li>Memory-mapped files under Windows.</li>
<li>We support the memory-safe compiler Fil-C.</li>
</ul>
<p>The simdjson 5.0 release improved performance compared to simdjson 4.0 in some key cases. Let me review some of them.</p>
<p>I built both versions with GCC 16.1 (<code>-O3</code>, CMake Release) and ran them on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core.</p>
<p>Let us start with DOM parsing of our standard files (GB/s):</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">file</th>
<th style="text-align: right;">4.0</th>
<th style="text-align: right;">5.0</th>
<th style="text-align: right;">speedup</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">twitter</td>
<td style="text-align: right;">4.78</td>
<td style="text-align: right;">4.82</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">citm_catalog</td>
<td style="text-align: right;">4.82</td>
<td style="text-align: right;">4.79</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">github_events</td>
<td style="text-align: right;">5.33</td>
<td style="text-align: right;">5.30</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">canada</td>
<td style="text-align: right;">1.10</td>
<td style="text-align: right;">1.21</td>
<td style="text-align: right;">1.1</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">marine_ik</td>
<td style="text-align: right;">1.25</td>
<td style="text-align: right;">1.38</td>
<td style="text-align: right;">1.1</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">mesh</td>
<td style="text-align: right;">1.17</td>
<td style="text-align: right;">1.26</td>
<td style="text-align: right;">1.1</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">numbers</td>
<td style="text-align: right;">1.13</td>
<td style="text-align: right;">1.41</td>
<td style="text-align: right;">1.3</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">twitterescaped</td>
<td style="text-align: right;">1.59</td>
<td style="text-align: right;">2.85</td>
<td style="text-align: right;">1.8</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">update-center</td>
<td style="text-align: right;">3.96</td>
<td style="text-align: right;">3.84</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">apache_builds</td>
<td style="text-align: right;">4.95</td>
<td style="text-align: right;">4.78</td>
<td style="text-align: right;">1.0</td>
</tr>
</tbody>
</table>
<p>Files full of numbers (canada, marine_ik, mesh, numbers) are 8% to 25% faster. And a file full of escaped Unicode characters (twitterescaped) is almost twice as fast: among other changes, we now decode consecutive <code>uXXXX</code> sequences without going back to the string scanner between them.</p>
<p>We also serialize faster. Printing floating-point numbers used to be a bottleneck. We replaced the ancient Grisu2 by Dragonbox, and removed calls to <code>memcpy</code> and <code>memmove</code> from the hot path.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">file</th>
<th style="text-align: right;">4.0</th>
<th style="text-align: right;">5.0</th>
<th style="text-align: right;">speedup</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">twitter</td>
<td style="text-align: right;">0.94</td>
<td style="text-align: right;">0.96</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">citm_catalog</td>
<td style="text-align: right;">1.07</td>
<td style="text-align: right;">1.08</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">gsoc-2018</td>
<td style="text-align: right;">1.10</td>
<td style="text-align: right;">1.25</td>
<td style="text-align: right;">1.1</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">canada</td>
<td style="text-align: right;">0.31</td>
<td style="text-align: right;">0.52</td>
<td style="text-align: right;">1.7</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">marine_ik</td>
<td style="text-align: right;">0.28</td>
<td style="text-align: right;">0.37</td>
<td style="text-align: right;">1.3</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">mesh</td>
<td style="text-align: right;">0.34</td>
<td style="text-align: right;">0.47</td>
<td style="text-align: right;">1.4</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">numbers</td>
<td style="text-align: right;">0.32</td>
<td style="text-align: right;">0.49</td>
<td style="text-align: right;">1.6</td>
</tr>
</tbody>
</table>
<p>The simdjson library is a community project. Since version 4.6, contributions came from fior512, 吴杨帆, Alecto Irene Perez, Francisco Geiman Thiesen, Max Bachmann, MoonFlowww, Advit Arora, Jaël Champagne Gareau, Taimoor Kiani, Vasily Pelikh, jmestwa-coder, liyinlong, AlbertoFVisconti, Aylin Dmello, Cuda Chen, Ezra Li, Madhurendra Purbay, Makkar, Pastoray, Paul Dreik, Pavel Kruglov, Piotr Kubaj, Yusuf İhsan Görgel, metsw24-max, neil, pratap singh, Vladimir Saraikin, wankun, xaldarof, Riyane El Qoqui, Justin Li and others. Thank you!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/28/simdjson-5-0-is-out/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>How fast can you fix a UTF-16 string in C#?</title>
		<link>https://lemire.me/blog/2026/09/26/how-fast-can-you-fix-a-utf-16-string-in-c/</link>
					<comments>https://lemire.me/blog/2026/09/26/how-fast-can-you-fix-a-utf-16-string-in-c/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 16:53:53 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22970</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="How fast can you fix a UTF-16 string in C#" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a high surrogate (U+D800 to U+DBFF) followed by a low surrogate (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed. You should never send an ill-formed string to disk &#8230; <a href="https://lemire.me/blog/2026/09/26/how-fast-can-you-fix-a-utf-16-string-in-c/" class="more-link">Continue reading <span class="screen-reader-text">How fast can you fix a UTF-16 string in C#?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-cover-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="How fast can you fix a UTF-16 string in C#" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a <em>high surrogate</em> (U+D800 to U+DBFF) followed by a <em>low surrogate</em> (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed.</p>
<p>You should never send an ill-formed string to disk or to the network. It is a bad practice.</p>
<p>In JavaScript, we have fast functions to fix strings or check whether they need fixing:</p>
<ul>
<li><code>String.prototype.toWellFormed()</code> replaces every lone surrogate with U+FFFD.</li>
<li><code>isWellFormed()</code> reports whether any replacement is needed.</li>
</ul>
<p>I added both functions to my C# library <a href="https://github.com/simdutf/SimdUnicode">SimdUnicode</a>, in <a href="https://github.com/simdutf/SimdUnicode/pull/54">pull request 54</a>. The algorithm is the same that we contributed to the JavaScript engine V8, so Chrome already fixes strings this way.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">string</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">UTF16</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">ToWellFormed</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">input</span><span style="color: #000; font-weight: bold;">);</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">// same instance, when the input is already well formed</span>
<span style="color: #204a87; font-weight: bold;">bool</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">ok</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">UTF16</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">IsWellFormed</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">span</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>When the input is well formed, <code>ToWellFormed</code> returns it as is. No allocation.</p>
<p>How are strings fixed? Basically, you replace bad inputs by the replacement character <code>U+FFFD</code>.</p>
<p>Our processors have special instructions called SIMD that allow data parallelism: you can compare multiple values at once. Recent x64 processors from AMD and Intel have better data parallelism than ARM chips, although both have powerful instructions.</p>
<p>The conventional approach in C# to repair a string is a function such as the following.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">static</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Repair</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">ReadOnlySpan</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">input</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Span</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">output</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">input</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">CopyTo</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">output</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">NextError</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">output</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">while</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&gt;=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">output</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000;">uFFFD</span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">NextError</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">output</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
<span style="color: #8f5902; font-style: italic;">// Index of the next lone surrogate at or after 'start', or -1 if none.</span>
<span style="color: #204a87; font-weight: bold;">static</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">NextError</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">ReadOnlySpan</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">start</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">start</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">while</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">true</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">k</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Slice</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">).</span><span style="color: #000;">IndexOfAnyInRange</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000;">uD800</span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000;">uDFFF</span><span style="color: #a40000; border: 1px solid #EF2929;">'</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">k</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-</span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">k</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">IsHighSurrogate</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">])</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Length</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">IsLowSurrogate</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">]))</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">2</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">// valid pair, skip it</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">else</span>
<span style="color: #f8f8f8;">            </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In SimdUnicode, I also use data parallelism.</p>
<p>Let me measure.</p>
<p><img decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-validate-xeon.webp" alt="Is the UTF-16 string well formed? Intel Xeon Gold 6548N" class="alignnone size-large" /></p>
<p>On the Xeon, with AVX-512, Latin validates at 69 GB/s against 33 GB/s for <code>IndexOfAnyInRange</code>. The Emoji input is well formed, and it is nothing but surrogate pairs. The runtime search drops to 0.4 GB/s. Our check holds 53 GB/s.</p>
<p><img decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-validate-m4.webp" alt="Is the UTF-16 string well formed? Apple M4 Max" class="alignnone size-large" /></p>
<p>Our results are similar on the M4 Max, although a bit less impressive compared to the Intel results.</p>
<p>Validation can return at the first lone surrogate. The buffer form of <code>ToWellFormed</code> writes every code unit, a copy of the input or U+FFFD. When the input is well formed, it is effectively a memory copy. Thus we can compare the performance against a copy.</p>
<p><img decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-buffer-xeon.webp" alt="Copy the string, replace lone surrogates. Intel Xeon Gold 6548N" class="alignnone size-large" /></p>
<p><img decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/utf16-buffer-m4.webp" alt="Copy the string, replace lone surrogates. Apple M4 Max" class="alignnone size-large" /></p>
<p>Roughly speaking, we are consistently about as fast as a copy.</p>
<p><strong>Versions used</strong>: .NET SDK 10.0.400 on Linux, 10.0.103 on macOS. Intel Xeon Gold 6548N (Emerald Rapids). Apple M4 Max.</p>
<p>Clausecker, R., &amp; Lemire, D. (2026). <a href="https://doi.org/10.1002/spe.70105">Fixing ill-formed UTF-16 strings with SIMD instructions</a>. Software: Practice and Experience. (<a href="https://arxiv.org/abs/2601.06349">arXiv</a>)</p>
<p><a href="https://github.com/simdutf/SimdUnicode/pull/54">Source code</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/26/how-fast-can-you-fix-a-utf-16-string-in-c/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>A thesis isn&#8217;t enough for a PhD</title>
		<link>https://lemire.me/blog/2026/09/26/a-thesis-isnt-enough-for-a-phd/</link>
					<comments>https://lemire.me/blog/2026/09/26/a-thesis-isnt-enough-for-a-phd/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 02:48:13 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22961</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_ehvp0eehvp0eehvp-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />To get a PhD, you typically have to enroll in a graduate program. Then you complete a few relatively easy courses.   You might need to pass a comprehensive examination that checks whether you have a basic understanding of the field. A few students fail at this point, but not many.Then you write a thesis &#8230; <a href="https://lemire.me/blog/2026/09/26/a-thesis-isnt-enough-for-a-phd/" class="more-link">Continue reading <span class="screen-reader-text">A thesis isn&#8217;t enough for a PhD</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_ehvp0eehvp0eehvp-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><div class="" data-block="true" data-editor="f3cv" data-offset-key="98c3d-0-0">
<div data-offset-key="98c3d-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="98c3d-0-0">To get a PhD, you typically have to enroll in a graduate program. Then you complete a few relatively easy courses.</span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="2niva-0-0">
<div data-offset-key="2niva-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="2niva-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="49c49-0-0">
<div data-offset-key="49c49-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="49c49-0-0">You might need to pass a comprehensive examination that checks whether you have a basic understanding of the field. A few students fail at this point, but not many.Then you write a thesis and defend it.</span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="crf9l-0-0">
<div data-offset-key="crf9l-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="crf9l-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="5bcv3-0-0">
<div data-offset-key="5bcv3-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="5bcv3-0-0">In theory, you could write a strong thesis and still fail the oral defense. That is uncommon. The defense is typically a public event, and there may be guests. Failing a student at that stage would be a public humiliation. I have seen students who could not answer basic questions still receive their PhDs because the thesis itself was good enough.</span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="e7303-0-0">
<div data-offset-key="e7303-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="e7303-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="1d8fi-0-0">
<div data-offset-key="1d8fi-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="1d8fi-0-0">Thus, to a very good approximation, getting a PhD amounted to writing a thesis.But we have a problem. AI can write something that looks quite a bit like a thesis. With a bit of prompting, you can produce something that looks like a PhD thesis.I have been telling everyone I can that we have a big problem.Apparently other people have realized this as well. People at Harvard are now saying that a thesis is insufficient for a PhD.</span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="4o4ed-0-0">
<div data-offset-key="4o4ed-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="4o4ed-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="fogll-0-0">
<div data-offset-key="fogll-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><a href="https://cmsa.fas.harvard.edu/media/2026/09/Summit-on-PhD-Math-Education-in-the-Age-of-AI.pdf"><span data-offset-key="fogll-0-0">Here is the key passage:</span></a></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="7u1rh-0-0">
<div data-offset-key="7u1rh-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="7u1rh-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="bodfc-0-0">
<div data-offset-key="bodfc-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="bodfc-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="f3cv" data-offset-key="cc61t-0-0">
<div data-offset-key="cc61t-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="cc61t-0-0">« The dissertation has in the past frequently been used as a proxy for the kind of mathematical development of a student that we expect: acquiring mathematical knowledge and demonstrating independent achievement. However, with modern AI systems, dissertations are not (&#8230;) a reliable tool for evaluating students, and PhDs should not be awarded primarily on the basis of the text of the dissertation. We thus recommend regular, multi-faceted evaluations of students in person, by multiple faculty, as the central component of assessment. These evaluations must be rigorous, not pro-forma, and should produce reports that might become part of the student’s dossier for future employment. » </span></div>
</div>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/26/a-thesis-isnt-enough-for-a-phd/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>How many strings can you create per second?</title>
		<link>https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/</link>
					<comments>https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 01:30:49 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22958</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/strings-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="How many strings can you create per second?" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />We create new strings all the time. How quickly can you produce short strings in your programming languge? To create a meaningful string, we convert an integer to a string. In Python, that is str(i). The loop stores each new string in a small ring buffer of 1024 slots, so that the engine cannot simply &#8230; <a href="https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/" class="more-link">Continue reading <span class="screen-reader-text">How many strings can you create per second?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/strings-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="How many strings can you create per second?" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>We create new strings all the time. How quickly can you produce short strings in your programming languge? To create a meaningful string, we convert an integer to a string. In Python, that is <code>str(i)</code>. The loop stores each new string in a small ring buffer of 1024 slots, so that the engine cannot simply discard the work.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">def</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">from_int</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">):</span>
    <span style="color: #000;">b</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #000;">buf</span>
    <span style="color: #204a87; font-weight: bold;">for</span> <span style="color: #000;">i</span> <span style="color: #204a87; font-weight: bold;">in</span> <span style="color: #204a87;">range</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">):</span>
        <span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span> <span style="color: #ce5c00; font-weight: bold;">&amp;</span> <span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">]</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #204a87;">str</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">)</span>
</code></pre>
</div>
<p>In JavaScript (Node.js and Bun), I use <code>String(i)</code>:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">function</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">from_int</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">let</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">String</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In C++, I use <code>std::to_string</code>:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">from_int</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">to_string</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In Rust, I use the standard <code>to_string()</code> and, as an alternative, the popular <a href="https://crates.io/crates/itoa">itoa</a> crate, which is designed for fast integer formatting:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">fn</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">std_to_string</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">&amp;</span><span style="color: #000;">mut</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">[</span><span style="color: #204a87;">String</span><span style="color: #000; font-weight: bold;">],</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">u64</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">in</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #ce5c00; font-weight: bold;">..</span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">[(</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">as</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">usize</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">to_string</span><span style="color: #000; font-weight: bold;">();</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
<span style="color: #204a87; font-weight: bold;">fn</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">itoa_to_string</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">&amp;</span><span style="color: #000;">mut</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">[</span><span style="color: #204a87;">String</span><span style="color: #000; font-weight: bold;">],</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">u64</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">let</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">mut</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">itoa</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #000;">Buffer</span><span style="color: #000; font-weight: bold;">::</span><span style="color: #000;">new</span><span style="color: #000; font-weight: bold;">();</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">in</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #ce5c00; font-weight: bold;">..</span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">[(</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">as</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">usize</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">format</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">).</span><span style="color: #000;">to_owned</span><span style="color: #000; font-weight: bold;">();</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In Go, I use <code>strconv.Itoa</code>:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">func</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">fromInt</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">buf</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">strconv</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Itoa</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In Nim, I use the <code>$</code> operator:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">proc</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">fromInt</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">:</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">int</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">in</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">..</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">:</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">buf</span><span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">and</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1023</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">$</span><span style="color: #000;">i</span>
</code></pre>
</div>
<p>The integers go from 0 to 100 million (10 million in Python), so each string has up to eight digits. I report the best of five runs on an Apple M4 Max. C++ is compiled with <code>-O3</code>, Rust in release mode, Nim with <code>-d:danger</code>.</p>
<p><img decoding="async" alt="Time to convert an integer to a new string on an Apple M4 Max" src="https://lemire.me/blog/wp-content/uploads/2026/09/strings.webp" class="alignnone size-large" /></p>
<p>Millions of strings per second:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;"></th>
<th style="text-align: right; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">million strings per second</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">C++ <code>std::to_string</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">183.8</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Nim <code>$i</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">85.9</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Go <code>strconv.Itoa</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">84.3</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Node.js <code>String(i)</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">73.1</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Rust <code>itoa</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">71.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Bun <code>String(i)</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">68.5</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Rust <code>to_string()</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">63.6</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Python <code>str(i)</code></td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">22.9</td>
</tr>
</tbody>
</table>
<p>Python is the slowest at 44 ns per string, about three times slower than JavaScript and eight times slower than C++.</p>
<p>The compiled languages that allocate each string on the heap (Rust, Go, Nim) end up in the same range as JavaScript: 12 to 16 ns per string. Rust with its standard <code>to_string()</code> is even a bit slower than Node.js and Bun. Garbage-collected runtimes like Go and JavaScript are very good at allocating many small, short-lived objects.</p>
<p>C++ wins by a wide margin at 5.4 ns per string. The trick is the <em>small string optimization</em>: a <code>std::string</code> stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.</p>
<p><strong>Versions used</strong>: macOS 15.7.7, CPython 3.14.5, Node.js 25.9.0, Bun 1.4.2, Apple clang 17.0.0 (clang-1700.6.4.2) with libc++, rustc 1.94.1 with itoa 1.0.18, Go 1.24.3, Nim 2.2.12.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/24">Source code</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>If you don&#8217;t have the factories, you lose the expertise</title>
		<link>https://lemire.me/blog/2026/09/24/if-you-dont-have-the-factories-you-lose-the-expertise/</link>
					<comments>https://lemire.me/blog/2026/09/24/if-you-dont-have-the-factories-you-lose-the-expertise/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 14:09:39 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22952</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/factories-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Mainstream economists have been advocating for the marvellous effects of global trade for decades. Who needs these factory jobs anyway? We&#8217;ll be designing the robots and the nuclear rockets. Except that, no. It does not work like that. If you have the factories, sooner or later, you get the designers. If you don&#8217;t have the &#8230; <a href="https://lemire.me/blog/2026/09/24/if-you-dont-have-the-factories-you-lose-the-expertise/" class="more-link">Continue reading <span class="screen-reader-text">If you don&#8217;t have the factories, you lose the expertise</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/factories-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Mainstream economists have been advocating for the marvellous effects of global trade for decades. Who needs these factory jobs anyway? We&#8217;ll be designing the robots and the nuclear rockets.</p>
<p>Except that, no. It does not work like that.</p>
<p>If you have the factories, sooner or later, you get the designers. If you don&#8217;t have the factories, you lose the expertise. Or you never get it.</p>
<p>There aren&#8217;t that many people left in the Silicon Valley capable of working on silicon. You can design a chip from California. Making one is done where the fab is.</p>
<p>You can get a PhD in robotics in Quebec City, but you won&#8217;t be designing robots unless you fly over to where they are made. Zoom calls won&#8217;t cut it at scale.<a href="https://lemire.me/blog/wp-content/uploads/2026/09/factories.jpg"></a></p>
<p>Jensen Huang&#8217;s account, as Joseph Steinberg <a href="https://x.com/jbsteinberg/status/2102912376852914190">reported it</a>, is that American manufacturing jobs declined because the work was outsourced. Steinberg says this is flatly wrong, and that technology accounts for the vast majority of the decline in manufacturing&#8217;s share of employment.</p>
<p><a href="https://lemire.me/blog/wp-content/uploads/2026/09/jobs.jpg"><img loading="lazy" decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/jobs-1024x570.jpg" alt="U.S. manufacturing employment, millions of jobs" width="660" height="367" class="alignnone size-large wp-image-22950" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/jobs-1024x570.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/jobs-300x167.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/jobs-768x428.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/jobs-1536x855.jpg 1536w, https://lemire.me/blog/wp-content/uploads/2026/09/jobs.jpg 2048w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>In 2000 there were 17.3 million manufacturing jobs in the United States. The peak was 19.6 million, in June 1979. From the early 1980s to 2000 the count stayed high, apart from the recessions. China joined the WTO in December 2001. By 2010 manufacturing employment was 11.5 million. In August 2026 it was 12.6 million.</p>
<p>That is 5.8 million jobs gone in a decade. About a million have come back since the bottom. The rest have not.</p>
<p>If technology were the main driver, output should have kept rising while the jobs fell. The fifteen years before 2000 are what that looks like. Manufacturing output nearly doubled, from an index of 51 to an index of 93 (2017 = 100). Employment went from 17.8 million to 17.3 million. The robots were already here in 1985. Employment did not collapse.</p>
<p><a href="https://lemire.me/blog/wp-content/uploads/2026/09/output.jpg"><img loading="lazy" decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/output-1024x562.jpg" alt="U.S. industrial production, index 2017 = 100" width="660" height="362" class="alignnone size-large wp-image-22951" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/output-1024x562.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/output-300x165.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/output-768x421.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/output-1536x842.jpg 1536w, https://lemire.me/blog/wp-content/uploads/2026/09/output.jpg 2048w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>After 2000, production stalled. Manufacturing output peaked just under 107 in December 2007. In August 2026 the index was 99.1. Total industrial production was 103.1. American factories are not turning out a flood of extra goods. They are turning out roughly what they turned out twenty years ago.</p>
<p>Output per hour did rise while the jobs were disappearing. The BLS index of manufacturing labor productivity went from about 70 in 2000 to about 100 in 2010. Then it stopped. In early 2026 it was still about 100. Flat output, fewer workers, a higher ratio. The ratio has been flat for fifteen years.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden;">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem;">Year</th>
<th style="text-align: right; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem;">Jobs (millions)</th>
<th style="text-align: right; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem;">Manufacturing output (2017 = 100)</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">1985</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">17.8</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">51</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">2000</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">17.3</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">93</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">2007</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">13.9</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">105</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">2010</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">11.5</td>
<td style="text-align: right; padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2;">93</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem;">August 2026</td>
<td style="text-align: right; padding: 0.65rem 0.8rem;">12.6</td>
<td style="text-align: right; padding: 0.65rem 0.8rem;">99.1</td>
</tr>
</tbody>
</table>
<p>The problem after 2000 was not that American factories became too productive. What changed after 2000 was where the goods were made.</p>
<p>And now, often, we don&#8217;t know anymore how to make things. Human expertise matters, and you maintain it by building stuff locally. You are not going to design microprocessors in Maine. It just won&#8217;t happen.</p>
<p>&nbsp;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/24/if-you-dont-have-the-factories-you-lose-the-expertise/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>A summer of AI optimization</title>
		<link>https://lemire.me/blog/2026/09/22/a-summer-of-ai-optimization/</link>
					<comments>https://lemire.me/blog/2026/09/22/a-summer-of-ai-optimization/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 02:27:44 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22946</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="A summer of AI optimization" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC&#8217;s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines. These libraries are mature. They &#8230; <a href="https://lemire.me/blog/2026/09/22/a-summer-of-ai-optimization/" class="more-link">Continue reading <span class="screen-reader-text">A summer of AI optimization</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="A summer of AI optimization" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC&#8217;s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines.</p>
<p>These libraries are mature. They have been optimized for years, by me and by others. For a long time, their performance was flat. Not because nobody cared, but because the remaining gains were expensive: each one required a few days of careful work, and nobody had the days.</p>
<p>Then, in 2026, six of them got much faster, most of it in a few weeks of summer.</p>
<p>To formalize my feeling, I rebuilt every commit of each library from scratch and benchmarked it on one machine (an Intel Xeon Gold 6548N). I track the speedup over time relative to August 2024. Thus the value 1.0 means no speedup. Whereas 2.0 means that the performance doubled. The lines are steps because performance only changes at a commit.</p>
<p>I should say that I cannot know how much AI was involved in each instance. I don&#8217;t ask how people arrived at their code. All I ask is that it be good. As for myself, I code with Claude (Opus 5), Grok and DeepSeek (V4 Pro). I was an early adopter of Grok for coding, and it got really good over time.</p>
<h2>1. roaring (compressed bitmaps, Go)</h2>
<p><img decoding="async" alt="roaring: speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-roaring.webp" class="alignnone size-large" /></p>
<p>The roaring library is the Go version of the Roaring index data structure. Decoding to an array got 2.5 times faster, the multi-way union <code>FastOr</code> got 3.1 times faster on one data set, the many-value iterator got 4.5 to 5.9 times faster, and the intersection cardinality gained 10%.</p>
<p>One of the contributors is an AI, actually. It is <a href="https://www.perfloop.com">perfloop</a>. (Disclosure: I am an advisor for perfloop.)</p>
<p>I did a lot of work. We also got help from Philipp Klose who declared using Claude.</p>
<h2>2. ada (URL parsing)</h2>
<p><img decoding="async" alt="ada: speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-ada.webp" class="alignnone size-large" /></p>
<p>The ada library is a standard compliant URL parser. From August 2024 to July 2026, about 550 commits went in and the throughput on a corpus of 100,000 URLs stayed at 0.54 GB/s. Then, in six weeks, it went to 1.28 GB/s: 2.4 times faster, about 15 million URLs per second on one core.</p>
<p>Most of the optimizations were done by Yagiz Nizipli, my long-time co-author. Yagiz works at SpaceX and uses Cursor (presumably with a grok model). Abdul Rawoof Khan and Dillon Mulroy also contributed an optimization each. I worked at optimizing IP address parsing, but it won&#8217;t show in this particular benchmark.</p>
<h2>3. fast_float (number parsing)</h2>
<p><img decoding="async" alt="fast_float: speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-fast-float.webp" class="alignnone size-large" /></p>
<p>The fast_float library parses floating-point numbers from text. It is part of GCC and most browsers. Performance was flat for fifteen months. Then, from March to July 2026, it gained 43% on one file (<code>canada.txt</code>, long coordinates) and 70% on another (<code>mesh.txt</code>, short coordinates). The optimizations should be credited to Koleman Nix and Filipe Oliveira.</p>
<h2>4. simdjson (JSON serialization and deserialization with C++26 reflection)</h2>
<p><img decoding="async" alt="simdjson: serialization and deserialization speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-simdjson.webp" class="alignnone size-large" /></p>
<p>The simdjson library recently gained support for C++26 static reflection: you serialize and parse your own structs directly, with no glue code. Since February 2026, serialization is 1.6 times faster on <code>twitter.json</code> and 2.1 times faster on <code>citm_catalog.json</code>. Deserialization, JSON straight into a struct, gained a more modest 10% and 14% (the second panel). (The reflection code only exists since early 2026.) The number of instructions per byte fell by almost exactly the same ratio as the throughput rose: from 6.1 to 3.1 instructions per byte on <code>citm_catalog.json</code> serialization.</p>
<p>Francisco Geiman Thiesen (Microsoft) did most of the work on the serialization side while I mostly helped improve our parsing. Francisco uses Claude.</p>
<h2>5. simdutf (Unicode validation and transcoding)</h2>
<p><img decoding="async" alt="simdutf: speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-simdutf.webp" class="alignnone size-large" /></p>
<p>The simdutf library validates and transcodes UTF-8, UTF-16 and UTF-32, and encodes and decodes base64. ASCII validation went from 83 GB/s to 160 GB/s. UTF-16 validation went from 62 GB/s to 102 GB/s. Base64 decoding gained 17%.</p>
<p>The work was done by Yagiz Nizipli (again) and myself.</p>
<p>The library got other amazing optimizations that do not show up on this benchmark by Gaspard Petit and Shreesh Adiga.</p>
<h2>6. CRoaring (compressed bitmaps, C)</h2>
<p><img decoding="async" alt="CRoaring: speedup over time" src="https://lemire.me/blog/wp-content/uploads/2026/09/summer-ai-croaring.webp" class="alignnone size-large" /></p>
<p>CRoaring implements Roaring bitmaps in C. On the real data sets from the repository, membership tests (<code>contains</code>) got 2.4 times faster, the cardinality of 64-bit bitmaps got 4.9 times faster, iterating over a 64-bit bitmap got 1.9 times faster, decoding a dense bitmap to an array got 2.2 times faster. Unions gained a more modest 13% to 16%.</p>
<p>The authors were Andrei Gudkov and myself.</p>
<h2>What happened</h2>
<p>The techniques used are all well-known. So why all these optimizations all of a sudden? Simply put, in my view, because it got cheap to try new ideas.</p>
<p>There is a lot of talk about the risks of AI in software. Human beings tend to be susceptible to the one-sided bet fallacy: when we see the downsides, we tend to ignore the benefits. Cars kill people, but ambulances save them.</p>
<p>In this instance, the benefits are concrete. Millions of people run these libraries, and this summer, they got faster.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/22/a-summer-of-ai-optimization/feed/</wfw:commentRss>
			<slash:comments>5</slash:comments>
		
		
			</item>
		<item>
		<title>More than a taken branch per cycle?</title>
		<link>https://lemire.me/blog/2026/09/21/more-than-a-taken-branch-per-cycle/</link>
					<comments>https://lemire.me/blog/2026/09/21/more-than-a-taken-branch-per-cycle/#comments</comments>
		
		<dc:creator><![CDATA[Antonio Badia]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 18:36:05 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22921</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/branching_branch_art-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Our processors can execute many instructions per cycle; they are superscalar. But not all instructions are equal. Branches are particularly tricky. A branch occurs often when you use an if-then clause or a loop. We distinguish between a taken branch and a not taken branch. A not-taken branch is often cheap. The processor just keeps &#8230; <a href="https://lemire.me/blog/2026/09/21/more-than-a-taken-branch-per-cycle/" class="more-link">Continue reading <span class="screen-reader-text">More than a taken branch per cycle?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/branching_branch_art-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Our processors can execute many instructions per cycle; they are <em>superscalar</em>. But not all instructions are equal.</p>
<p>Branches are particularly tricky. A branch occurs often when you use an if-then clause or a loop. We distinguish between a taken branch and a not taken branch.</p>
<p>A not-taken branch is often cheap. The processor just keeps going.</p>
<p>A taken branch jumps to a new location. A taken branch can be more expensive.</p>
<p>You will often hear that processors are limited to one taken branch per cycle.</p>
<p>I decided to test it out with a loop with an <code>if</code> inside it. Here is the function I tested in Go.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">func</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">lastHit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">[]</span><span style="color: #204a87; font-weight: bold;">byte</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">thresh</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">byte</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">last</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #204a87; font-weight: bold;">byte</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">len</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">v</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">p</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">]</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">v</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">thresh</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">            </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">last</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">v</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">n</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">            </span><span style="color: #204a87; font-weight: bold;">break</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>Look at the main loop. We load a value from an array, we compare it with a threshold. If it is greater than the threshold, then we assign it to the <code>last</code> pointer. So the function effectively records the last seen value that is greater than the threshold. That&#8217;s pretty reasonable code.</p>
<p>Consider the case where you always miss. The values are always smaller than or equal to the threshold. In these cases, we get two taken branches in close proximity, but no store. (It is a bit confusing but that&#8217;s how the Go compiler does it.)</p>
<p><a href="https://lemire.me/blog/wp-content/uploads/2026/09/misses.webp"><img decoding="async" src="https://lemire.me/blog/wp-content/uploads/2026/09/misses.webp" alt="Cycles per iteration when the if always misses" class="alignnone size-large" /></a></p>
<p>The processor that struggles the most is the AMD Zen 4 processor. But AMD Zen 5 is much better.</p>
<p>So two processors are able to take two branches in less than 2 cycles on average in this test: the Apple processor (M4 Max) and the Granite Rapids processor.</p>
<p>This means that, yes, modern processors can execute more than one taken branch per cycle under some conditions.</p>
<p>The code is under <code>benchmark/experiments/ifloop</code> in the <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/21">GitHub repo</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/21/more-than-a-taken-branch-per-cycle/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>The data does not show mass unemployment</title>
		<link>https://lemire.me/blog/2026/09/21/nomassunemployment/</link>
					<comments>https://lemire.me/blog/2026/09/21/nomassunemployment/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 15:04:08 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22933</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/P9wow-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />If AI is causing mass unemployment among software developers, it is not showing up in the data yet.   The USA had more software developers as a percentage of the population in 2025 than in 2021.   On average salaries are slightly up (average of 148k$ a year). The 90th percentile is up from 2021 &#8230; <a href="https://lemire.me/blog/2026/09/21/nomassunemployment/" class="more-link">Continue reading <span class="screen-reader-text">The data does not show mass unemployment</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/P9wow-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><div class="" data-block="true" data-editor="1lfk9" data-offset-key="2ql3n-0-0">
<div data-offset-key="2ql3n-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="2ql3n-0-0">If AI is causing mass unemployment among software developers, it is not showing up in the data yet.</span></div>
</div>
<div class="" data-block="true" data-editor="1lfk9" data-offset-key="8rfph-0-0">
<div data-offset-key="8rfph-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="8rfph-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="1lfk9" data-offset-key="bsf6a-0-0">
<div data-offset-key="bsf6a-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="bsf6a-0-0">The USA had more software developers as a percentage of the population in 2025 than in 2021.</span></div>
</div>
<div class="" data-block="true" data-editor="1lfk9" data-offset-key="1algn-0-0">
<div data-offset-key="1algn-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="1algn-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="1lfk9" data-offset-key="2pvg6-0-0">
<div data-offset-key="2pvg6-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="2pvg6-0-0">On average salaries are slightly up (average of 148k$ a year). The 90th percentile is up from 2021 (215k$US a year).</span><a href="http://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-scaled.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-1024x475.png" alt="" width="660" height="306" class="alignnone size-large wp-image-22934" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-1024x475.png 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-300x139.png 300w, https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-768x356.png 768w, https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-1536x712.png 1536w, https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-21-a-10.59.39-2048x950.png 2048w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></div>
</div>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/21/nomassunemployment/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>AI is breaking the academic sorting machine</title>
		<link>https://lemire.me/blog/2026/09/20/ai-is-breaking-the-academic-sorting-machine/</link>
					<comments>https://lemire.me/blog/2026/09/20/ai-is-breaking-the-academic-sorting-machine/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:47:46 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22928</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/INeo8-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />I see the mathematicians panicking at what we can do with agentic AI. Let us be clear on what we are talking about. In a few weeks, someone who is not a high-level mathematician can, with AI, produce the equivalent of a good PhD thesis in math. I could see a top 1% high school &#8230; <a href="https://lemire.me/blog/2026/09/20/ai-is-breaking-the-academic-sorting-machine/" class="more-link">Continue reading <span class="screen-reader-text">AI is breaking the academic sorting machine</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/INeo8-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p><span>I see the mathematicians panicking at what we can do with agentic AI. </span></p>
<p>Let us be clear on what we are talking about. In a few weeks, someone who is not a high-level mathematician can, with AI, produce the equivalent of a good PhD thesis in math. I could see a top 1% high school student, given enough of an AI token budget, just write the equivalent of a PhD thesis as a hobby.</p>
<p>If Joe Smith completed a PhD thesis in math back in 2022, it is now possible for a really smart high schooler to generate the same output while playing video games. Joe must not be too happy about it, especially if Joe is still looking for a prestigious faculty position.</p>
<p>Notice how we are not so excited. I mean, why aren’t we celebrating this incredible breakthrough? Finally, all the math problems we have can be solved faster! It is worth reflecting on why we don’t care.</p>
<p>Part of the issue is that most of academic research lost its customers years ago.</p>
<p>In part thanks to the arrival of “peer review” in the 1970s, we have long ago closed the research world into siloed communities. The purpose of the academic output is to sort people out for jobs. Write papers that are impressive and you may get a good professorship. If you don’t, then you will be flipping burgers. There is an incredible glut of people with academic credentials, but only so many jobs at the top. The glut has been continuously, and somewhat deliberately, increasing.</p>
<p>If you hold a PhD today, your chance of having a tenure-track or tenured position is about 10% and falling. Let me be clear. Any academic who shows worry for what the PhDs might do now should look in a mirror because you have been training too many for decades. And, also, if fewer people decide to go for the PhD that might be more than fine.<br />
<a href="http://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq.jpeg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq-1024x822.jpeg" alt="" width="660" height="530" class="alignnone size-large wp-image-22931" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq-1024x822.jpeg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq-300x241.jpeg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq-768x616.jpeg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/HSqvM7IXQAAL7iq.jpeg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>AI disrupts the sorting mechanism in mathematics… But do you think for a minute that it does not apply to mechanical engineering, sociology, etc.? Is that bad news? No. I think that it is excellent news and I have been saying so for years.</p>
<p>Here is what I wrote in 2024…</p>
<p>« AI’s ability to generate vast amounts of text raises concerns about a potential flood of irrelevant theoretical papers, further straining the evaluation system. Stonebraker’s (2018) call for rewarding problem-solving over publication needs revisiting. Perhaps the emphasis should be on the impact and significance of research, not just its passage through peer review—a skill replicable by AI. AI can pave the way for a “golden age” of scientific progress if we can develop new evaluation methods focused on problem-solving and real-world impact. The scientific community must adapt to the evolving landscape. By recognizing the limitations of peer review and prioritizing the pursuit of meaningful solutions, we can ensure that AI becomes a catalyst for scientific advancement, not a detriment. »</p>
<p>I predict that, over time, the focus will move away from “papers as the final output.”</p>
<p>Nobody wants a paper about cancer, we want to eradicate cancer. We should reward people who get results, no matter which tools they use. And solving a problem because it is impressive to do so is not enough.</p>
<p>« But Daniel, Mathematicians can&#8217;t cure cancer or give us antigravity. »</p>
<p>Maybe it is time they try. Frankly, the AI disruption might be precisely what we needed.</p>
<p>« But Daniel, won’t AI just replace all of us? »</p>
<p>I wish. But thus far, it is not happening. I have more AI accounts than most people have pencils and I am working 50 hours a week. Intelligence is not a scalar quantity. Once an AI can do something, I somehow always find more work to do.</p>
<p><strong>Further reading</strong>. Daniel Lemire, <a href="https://dl.acm.org/doi/full/10.1145/3673649">Will AI Flood Us with Irrelevant Papers?</a> Communications of the ACM, Vol. 67, No. 9 (September 2024).</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/20/ai-is-breaking-the-academic-sorting-machine/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>How did Apple Silicon get 50% faster in three years?</title>
		<link>https://lemire.me/blog/2026/09/19/how-did-apple-silicon-get-50-faster-in-three-years/</link>
					<comments>https://lemire.me/blog/2026/09/19/how-did-apple-silicon-get-50-faster-in-three-years/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 16:13:41 +0000</pubDate>
				<category><![CDATA[Science and Technology]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22914</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/apple-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Four processor packages in a row on a desk" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Recently, I looked at how AMD made its chips 50% faster within a span of two years, in How did AMD Ryzen get 50% faster in two years? Many people asked me to cover Apple processors. So let us go. Consider the base Apple Silicon chips: M2 (2022), M3 (2023), M4 (2024) and M5 (2025). &#8230; <a href="https://lemire.me/blog/2026/09/19/how-did-apple-silicon-get-50-faster-in-three-years/" class="more-link">Continue reading <span class="screen-reader-text">How did Apple Silicon get 50% faster in three years?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/apple-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Four processor packages in a row on a desk" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Recently, I looked at how AMD made its chips 50% faster within a span of two years, in <a href="https://lemire.me/blog/2026/09/18/how-did-amd-ryzen-get-50-faster-in-two-years/">How did AMD Ryzen get 50% faster in two years?</a> Many people asked me to cover Apple processors. So let us go.</p>
<p>Consider the base Apple Silicon chips: M2 (2022), M3 (2023), M4 (2024) and M5 (2025). I am skipping the original one (M1) for simplicity and I am not considering the Pro and Max chips. The M6 is out there too, just released, but I am going to cover it at the end.</p>
<p>On Geekbench 6, performance went up by about 50% in three years. The increase is rather gradual. No big jump.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/geekbench.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/geekbench-1024x424.jpg" alt="Geekbench 6 single-core and multi-core scores for Apple M2, M3, M4 and M5" width="660" height="273" class="alignnone size-large wp-image-22906" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/geekbench-1024x424.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/geekbench-300x124.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/geekbench-768x318.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/geekbench.jpg 1199w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
M2<br />
4P+4E</td>
<td>2023<br />
M3<br />
4P+4E</td>
<td>2024<br />
M4<br />
4P+6E</td>
<td>2025<br />
M5<br />
4P+6E</td>
</tr>
<tr>
<td>Single-core</td>
<td>2,401</td>
<td>2,767</td>
<td>3,278</td>
<td>3,642</td>
</tr>
<tr>
<td>Multi-core</td>
<td>9,814</td>
<td>11,544</td>
<td>15,345</td>
<td>17,955</td>
</tr>
</tbody>
</table>
<p>The 2025 chip is 52% faster on a single core than the 2022 chip, and 83% faster with all cores.</p>
<p>Unlike AMD, Apple improved the CPU frequency quite a bit. The P-core went from 3.5 GHz to 4.6 GHz, a 30% increase. On its own, it can explain the bulk of the single-core performance increase.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/frequency.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/frequency-1024x753.jpg" alt="P-core and E-core max frequencies for Apple M2, M3, M4 and M5" width="660" height="485" class="alignnone size-large wp-image-22907" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/frequency-1024x753.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/frequency-300x221.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/frequency-768x564.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/frequency.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
M2</td>
<td>2023<br />
M3</td>
<td>2024<br />
M4</td>
<td>2025<br />
M5</td>
</tr>
<tr>
<td>P-core max</td>
<td>3.5 GHz</td>
<td>4.0 GHz</td>
<td>4.4 GHz</td>
<td>4.6 GHz</td>
</tr>
<tr>
<td>E-core max</td>
<td>2.4 GHz</td>
<td>2.8 GHz</td>
<td>2.9 GHz</td>
<td>3.0 GHz</td>
</tr>
</tbody>
</table>
<p>But still, even if we normalized the clock speed, an Apple M5 would be noticeably faster than an Apple M2.</p>
<p>Let us look at the transistor count. For the M5, I could not find any information. But there is otherwise a nice gradual increase. From the M2 to the M4, the number went up by about 40%, from 20 billion to 28 billion.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/transistors.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/transistors-1024x753.jpg" alt="Transistor counts for Apple M2, M3 and M4; M5 not disclosed" width="660" height="485" class="alignnone size-large wp-image-22908" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/transistors-1024x753.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/transistors-300x221.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/transistors-768x564.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/transistors.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>Where did they go?</p>
<p>The base M4 processor has two extra efficiency cores. This will not help with single-core performance, but it will help multicore performance.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/cores.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/cores-1024x753.jpg" alt="Performance and efficiency core counts for Apple M2, M3, M4 and M5" width="660" height="485" class="alignnone size-large wp-image-22909" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/cores-1024x753.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/cores-300x221.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/cores-768x564.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/cores.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
M2</td>
<td>2023<br />
M3</td>
<td>2024<br />
M4</td>
<td>2025<br />
M5</td>
</tr>
<tr>
<td>Performance cores</td>
<td>4</td>
<td>4</td>
<td>4</td>
<td>4</td>
</tr>
<tr>
<td>Efficiency cores</td>
<td>4</td>
<td>4</td>
<td>6</td>
<td>6</td>
</tr>
</tbody>
</table>
<p>But what about each individual core?</p>
<p>An important variable at the core level is the number of instructions per cycle. And this went up. Decode width went from 8 instructions per cycle on the M2 to 10 on the M4 and the M5, which is higher than what most competitors can do. So you can get more compute per cycle.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ipc.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ipc-1024x753.jpg" alt="Decode width in instructions per cycle for Apple M2, M3, M4 and M5" width="660" height="485" class="alignnone size-large wp-image-22910" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ipc-1024x753.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ipc-300x221.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ipc-768x564.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ipc.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>Thus far, we do not see much difference between the M4 and the M5, and yet the M5 can be significantly faster. Let us look next at memory bandwidth.</p>
<p>And yeah, the M5 has much more memory bandwidth: 154 GB/s against 100 GB/s on the M2 and the M3. That helps on multicore processing tasks, as it is when you are most likely to run out of bandwidth.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/bandwidth.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/bandwidth-1024x751.jpg" alt="Memory bandwidth for Apple M2, M3, M4 and M5" width="660" height="484" class="alignnone size-large wp-image-22911" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/bandwidth-1024x751.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/bandwidth-300x220.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/bandwidth-768x563.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/bandwidth.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
M2</td>
<td>2023<br />
M3</td>
<td>2024<br />
M4</td>
<td>2025<br />
M5</td>
</tr>
<tr>
<td>Bandwidth</td>
<td>100 GB/s</td>
<td>100 GB/s</td>
<td>120 GB/s</td>
<td>154 GB/s</td>
</tr>
<tr>
<td>Memory</td>
<td>LPDDR5-6400</td>
<td>LPDDR5-6400</td>
<td>LPDDR5X-7500</td>
<td>LPDDR5X-9600</td>
</tr>
</tbody>
</table>
<p>What about data parallelism? AMD improved data parallelism (SIMD) considerably. The story is much less interesting with Apple processors. They have four 128-bit execution units, all of them. Meanwhile the Zen 5 has four 512-bit units. It is four times more. What is new with the M4 and the M5 is the new 512-bit SME matrix unit.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/simd.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/simd-1024x445.jpg" alt="Apple M-series four 128-bit NEON units versus AMD Zen 5 four 512-bit AVX-512 units" width="660" height="287" class="alignnone size-large wp-image-22912" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/simd-1024x445.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/simd-300x131.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/simd-768x334.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/simd.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>But there is a new iteration. The M6.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/m6.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/m6-1024x834.jpg" alt="Apple M6 processor" width="660" height="538" class="alignnone size-large wp-image-22913" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/m6-1024x834.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/m6-300x244.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/m6-768x625.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/m6.jpg 1086w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>The number of cores goes from 10 to 12. They added two &#8216;super&#8217; cores. The bandwidth is up a bit. My expectation is that Apple will still again deliver a boost in performance.</p>
<p>In effect, what I am demonstrating is that CPUs are not boring. They are improving fast.</p>
<p>I already made the point that <a href="https://lemire.me/blog/2025/09/01/processors-are-getting-wider/">processors are getting wider</a>. See also <a href="https://lemire.me/blog/2025/07/09/memory-level-parallelism-apple-m2-vs-apple-m4/">Memory-level parallelism: Apple M2 vs Apple M4</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/19/how-did-apple-silicon-get-50-faster-in-three-years/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Faster JSON parsing with SVE2 on ARM processors</title>
		<link>https://lemire.me/blog/2026/09/18/faster-json-parsing-with-sve2-on-arm-processors/</link>
					<comments>https://lemire.me/blog/2026/09/18/faster-json-parsing-with-sve2-on-arm-processors/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 21:38:09 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22899</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_co7959co7959co79-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />ARM processors, like those in your phone, have instructions capable of processing several elements at once (SIMD). These instructions are called NEON. But many newer processors have a different SIMD extension called SVE. The latest ARM processors have SVE2. Unfortunately, Apple has not yet adopted SVE, but SVE processors are available in the cloud. In &#8230; <a href="https://lemire.me/blog/2026/09/18/faster-json-parsing-with-sve2-on-arm-processors/" class="more-link">Continue reading <span class="screen-reader-text">Faster JSON parsing with SVE2 on ARM processors</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_co7959co7959co79-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>ARM processors, like those in your phone, have instructions capable of processing several elements at once (SIMD). These instructions are called NEON. But many newer processors have a different SIMD extension called SVE. The latest ARM processors have SVE2. Unfortunately, Apple has not yet adopted SVE, but SVE processors are available in the cloud.</p>
<p>In April, I wrote that the SVE2 <code>match</code> instruction might be <a href="https://lemire.me/blog/2026/04/19/the-fastest-way-to-match-characters-on-arm-processors/">the fastest way to match characters on ARM processors</a>. At the time, my benchmark was a toy. The question was whether the idea survives contact with a real parser. Madhurendra Purbay, an engineer at ARM, answered the question with a <a href="https://github.com/simdjson/simdjson/pull/2863">pull request to the simdjson library</a>. Let me go through what it does and what it buys us.</p>
<p>The simdjson library includes a fast JSON parser. JSON is a ubiquitous data format online; everyone uses it. It is made of strings, numbers, arrays (<code>[1,2,3]</code>) and <em>objects</em>. An object is a key-value map where keys are strings, written as <code>{"key1": 1, "key2": 2}</code>. You can combine arrays and objects (e.g., an object can be in an array).</p>
<p>When the simdjson library indexes a JSON document, it first computes, for each block of 64 bytes, a few 64-bit masks. One of them marks the JSON <em>structural</em> characters (<code>,</code>, <code>:</code>, <code>[</code>, <code>]</code>, <code>{</code>, <code>}</code>). From these masks and a few others, we derive the positions of all the JSON tokens.</p>
<p>The ARM NEON version of the classifier, which I designed with Geoff Langdale years ago, uses a table lookup (<code>tbl</code>). Take the byte, add 3, keep the high nibble, and look it up in a 16-byte table that returns the one structural character with that nibble (or <code>0xff</code>). If the looked-up byte is equal to the input, the input is structural. With the <code>vaddq</code>/<code>vshrq</code>/<code>vqtbl1q</code>/<code>vceqq</code> NEON intrinsics, it is four instructions per 16 bytes:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">op_table</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">simd8</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">uint8_t</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #000; font-weight: bold;">(</span>
<span style="color: #f8f8f8;">  </span><span style="color: #0000cf; font-weight: bold;">0xff</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">','</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">':'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'['</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">']'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'{'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'}'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span>
<span style="color: #000; font-weight: bold;">);</span>
<span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">match_op_0</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vceqq_u8</span><span style="color: #000; font-weight: bold;">(</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">vqtbl1q_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">op_table</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vshrq_n_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">vaddq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">d0_0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vdupq_n_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #0000cf; font-weight: bold;">3</span><span style="color: #000; font-weight: bold;">)),</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">)),</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">d0_0</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>We get a vector of 16 bytes that are either <code>0x00</code> or <code>0xff</code>. To turn four such vectors into one 64-bit mask, we AND each byte with a bit weight (1, 2, 4, &#8230;, 128) and sum adjacent bytes three times with <code>addp</code>. That is another eight instructions or so per 64-byte block, and it is shared with the white-space mask.</p>
<p>SVE2 has an instruction, <code>match</code>, that takes a vector of bytes and a second vector that acts as a small set: it produces a predicate (a mask) with a bit set at each position where the input byte belongs to the set. Because it works within 128-bit segments, the set is at most 16 bytes, which is plenty for our six structural characters. NEON has nothing like it; on x64, the closest thing is the SSE4.2 string-comparison instructions (<code>pcmpistrm</code>), which are slow.</p>
<p>In C++, using intrinsics, the code might look as follows:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">// input is a set of 16 ASCII bytes we want to classify</span>
<span style="color: #8f5902; font-style: italic;">// the whole thing compiles to little more than the match instruction</span>
<span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">match_operators_sve2</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">input</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// The characters we care about. We use `0xff` as a</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// filler (it is an impossible byte value within a JSON document)</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">operators</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #0000cf; font-weight: bold;">0xff</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">','</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">':'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'['</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">']'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'{'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'}'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0xff</span><span style="color: #000; font-weight: bold;">,</span>
<span style="color: #f8f8f8;">    </span><span style="color: #4e9a06;">','</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">':'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'['</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">']'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'{'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">'}'</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">','</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">':'</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// pg is a mask over the first 16 values</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">pg</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svptrue_pat_b8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">SV_VL16</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// 'move' the NEON register to SVE</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svuint8_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">data</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svset_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svundef_u8</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">input</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// 'move' the table to SVE</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svuint8_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">table</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svset_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svundef_u8</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">operators</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// call the match instruction</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svmatch_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">pg</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">data</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">table</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>This function <em>classifies</em> 16 ASCII bytes with maybe just one instruction (<code>match</code>).</p>
<p>We use <code>svset_neonq_u8</code>, which is part of the NEON-SVE bridge. It allows you to mix and match NEON and SVE. The <code>uint8x16_t</code> type is a NEON type (16 8-bit integers). The type <code>svuint8_t</code> is an SVE type (a vector of 8-bit integers). As a convention, the first 16 bytes of SVE types are shared with NEON. Thus I expect that <code>svuint8_t data = svset_neonq_u8(svundef_u8(), input)</code> might compile to nothing.</p>
<p>The catch, as I explained in April, is that a predicate lives in a predicate register. My function returns <code>svbool_t</code>. SVE gives you no cheap way to move it to a general-purpose register: the architecture does not want to assume that a mask fits in 16 bits, since the registers might be wider. So we materialize the predicate as bytes instead, with a predicated select (<code>svsel</code>). The whole thing is a bit complicated (see <a href="https://doi.org/10.1002/spe.3420">Lemire (2025)</a> for an explanation of the trick).</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">// We have four masks, p0, p1, p2, p3</span>
<span style="color: #8f5902; font-style: italic;">// and we want to convert them each to a 16-bit value and then combine them to</span>
<span style="color: #8f5902; font-style: italic;">// form a 64-bit mask.</span>
<span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">operator_predicates_to_bytes</span><span style="color: #000; font-weight: bold;">(</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">p0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">p1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">p2</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svbool_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">p3</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">bit_mask</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">0x01</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x02</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x4</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x8</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x10</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x20</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x40</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x80</span><span style="color: #000; font-weight: bold;">,</span>
<span style="color: #f8f8f8;">                         </span><span style="color: #0000cf; font-weight: bold;">0x01</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x02</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x4</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x8</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x10</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x20</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x40</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0x80</span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// map the NEON register bit_mask to an SVE register</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svuint8_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">weights</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svset_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svundef_u8</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">bit_mask</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// create a zero register</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svuint8_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svdup_n_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// where p0 is set, put the value from bit_mask, otherwise zero</span>
<span style="color: #f8f8f8;">  </span><span style="color: #8f5902; font-style: italic;">// The `svget_neonq_u8` function is part of the NEON-SVE bridge.</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b0</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svget_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svsel_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">weights</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b1</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svget_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svsel_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">weights</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b2</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svget_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svsel_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p2</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">weights</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b3</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">svget_neonq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">svsel_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">p3</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">weights</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">const</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">uint8x16_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">sum</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vpaddq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">vpaddq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">b0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b1</span><span style="color: #000; font-weight: bold;">),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vpaddq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">b2</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b3</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">vgetq_lane_u64</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">vreinterpretq_u64_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">vpaddq_u8</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">sum</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">sum</span><span style="color: #000; font-weight: bold;">)),</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>Does it help?</p>
<p>I benchmarked the <code>match</code> classifier against the NEON classifier, using the <code>parse</code> benchmark that comes with simdjson, over the 22 JSON files that we use as our standard corpus (about 24 MB in total). The benchmark parses each file 300 times and keeps the best time. I ran each binary three times, interleaved with its counterpart, pinned to one core, and I kept the best. The run-to-run variation is under 1%. I used GCC 15 and LLVM clang 21 on Ubuntu 26.04, with <code>-mcpu=native</code>.</p>
<p>The <code>match</code> instruction is part of SVE2, not the original SVE. Among the AWS Graviton processors, the Graviton 3 (Neoverse V1) has SVE but not SVE2: it cannot run this code. So I used the two processors that can:</p>
<ul>
<li>Graviton 4 (Neoverse V2) on a <code>c8g.2xlarge</code> instance,</li>
<li>Graviton 5 (Neoverse V3) on a <code>c9g.2xlarge</code> instance.</li>
</ul>
<p>Both have 128-bit SVE registers.</p>
<p>Here is the gain in the indexing stage (stage 1), file by file, as the percentage increase in throughput with <code>match</code> over NEON. The dashed line in each panel is the geometric mean over the 22 files. First with GCC:<br />
<a href="http://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc.webp"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-1024x768.webp" alt="" width="660" height="495" class="alignnone size-large wp-image-22902" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-1024x768.webp 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-300x225.webp 300w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-768x576.webp 768w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-1536x1152.webp 1536w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_gcc-2048x1536.webp 2048w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>And with clang:</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang.webp"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-1024x768.webp" alt="" width="660" height="495" class="alignnone size-large wp-image-22903" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-1024x768.webp 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-300x225.webp 300w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-768x576.webp 768w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-1536x1152.webp 1536w, https://lemire.me/blog/wp-content/uploads/2026/09/speedup_clang-2048x1536.webp 2048w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>The files that gain the least (<code>canada</code>, <code>mesh</code>, <code>marine_ik</code>) are mostly numbers, where the indexing stage is cheap to begin with. The files that gain the most (<code>gsoc-2018</code>, <code>random</code>, <code>github_events</code>) are the ones with a lot of structure. No file gets slower, except <code>canada</code> and <code>mesh</code> on the Graviton 5 with GCC (by 2% to 3%, at the edge of what I can measure).</p>
<p>In absolute terms, the indexing stage goes from 4.8 GB/s to 5.3 GB/s on the Graviton 4 with clang (5.5 GB/s to 5.8 GB/s with GCC), and from 6.3 GB/s to 6.6 GB/s on the Graviton 5 with clang (7.1 GB/s to 7.3 GB/s with GCC). The Graviton 4 benefits more than the Graviton 5.</p>
<p>The second stage of the parser is untouched, so the gain on the whole parse is smaller: 2% to 4% on the Graviton 4 and 1% to 2% on the Graviton 5. It is a modest gain, but it comes from replacing four NEON instructions with one, in a routine that we had already tuned carefully.</p>
<p>In a real parser, <code>match</code> gives 3% to 9% faster indexing on Graviton 4 and Graviton 5, with a handful of intrinsics and no assembly. The code is in <a href="https://github.com/simdjson/simdjson/pull/2866">simdjson pull request 2866</a>. It requires SVE2, which Apple processors and the older Graviton processors do not have, so simdjson falls back on NEON when SVE2 is not available at compile time.</p>
<p>The limitation today is that the code is compiled in only if you build with <code>-mcpu=native</code> or the equivalent: a default build gets the NEON code everywhere. The next step for simdjson is to select the SVE2 code at runtime, as we do with the various x64 instruction sets, so that a default build uses it on processors that have the instruction. I am working on it.</p>
<p><em>Credit</em>: The <code>match</code> classifier is the work of <a href="https://github.com/mpurbay-arm">Madhurendra Purbay</a> (ARM), from his <a href="https://github.com/simdjson/simdjson/pull/2863">pull request 2863</a>, where he used a different (and slightly faster) technique, with inline assembly, to extract the predicates. <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/svehelp">My benchmark results and scripts are available</a>.</p>
<p><em>References</em></p>
<p>Keiser, J., &amp; Lemire, D. (2024). <a href="https://doi.org/10.1002/spe.3313">On-demand JSON: A better way to parse documents?</a>. Software: Practice and Experience, 54(6), 1074-1086.</p>
<p>Langdale, G., &amp; Lemire, D. (2019). <a href="https://doi.org/10.1007/s00778-019-00578-5">Parsing gigabytes of JSON per second</a>. The VLDB Journal, 28(6), 941-960. (<a href="https://arxiv.org/abs/1902.08318">arXiv</a>)</p>
<p>Lemire, D. (2025). <a href="https://lemire.me/blog/2025/03/29/mixing-arm-neon-with-sve-code-for-fun-and-profit/">Mixing ARM NEON with SVE code for fun and profit</a>.</p>
<p>Lemire, D. (2025). <a href="https://doi.org/10.1002/spe.3420">Scanning HTML at tens of gigabytes per second on ARM processors</a>. Software: Practice and Experience, 55(7), 1256-1265.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/18/faster-json-parsing-with-sve2-on-arm-processors/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>How did AMD Ryzen get 50% faster in two years?</title>
		<link>https://lemire.me/blog/2026/09/18/how-did-amd-ryzen-get-50-faster-in-two-years/</link>
					<comments>https://lemire.me/blog/2026/09/18/how-did-amd-ryzen-get-50-faster-in-two-years/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 17:26:04 +0000</pubDate>
				<category><![CDATA[Science and Technology]]></category>
		<guid isPermaLink="false">http://lemire.me/blog/?p=22896</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Three desktop processors in a row, from nearest to farthest" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />People still tell me that CPUs are boring. That nothing much happens anymore. Let us look at AMD Ryzen 7 processors from 2022 to 2024: the 5800X3D (Zen 3), the 7800X3D (Zen 4) and the 9800X3D (Zen 5). They are comparable 8-core chips with 3D V-Cache. I have written about them before, in How stagnant &#8230; <a href="https://lemire.me/blog/2026/09/18/how-did-amd-ryzen-get-50-faster-in-two-years/" class="more-link">Continue reading <span class="screen-reader-text">How did AMD Ryzen get 50% faster in two years?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-featured-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="Three desktop processors in a row, from nearest to farthest" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>People still tell me that CPUs are boring. That nothing much happens anymore.</p>
<p>Let us look at AMD Ryzen 7 processors from 2022 to 2024: the 5800X3D (Zen 3), the 7800X3D (Zen 4) and the 9800X3D (Zen 5). They are comparable 8-core chips with 3D V-Cache. I have written about them before, in <a href="https://lemire.me/blog/2026/01/14/how-stagnant-is-cpu-technology/">How stagnant is CPU technology?</a></p>
<p>On Geekbench 6, performance went up by about 50% in two years.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6-1024x424.jpg" alt="Geekbench 6 single-core and multi-core scores for Ryzen 7 5800X3D, 7800X3D and 9800X3D" width="660" height="273" class="alignnone size-large wp-image-22891" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6-1024x424.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6-300x124.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6-768x318.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-geekbench6.jpg 1199w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
5800X3D<br />
Zen 3</td>
<td>2023<br />
7800X3D<br />
Zen 4</td>
<td>2024<br />
9800X3D<br />
Zen 5</td>
</tr>
<tr>
<td>Single-core</td>
<td>2,016</td>
<td>2,426</td>
<td>2,969</td>
</tr>
<tr>
<td>Multi-core</td>
<td>11,832</td>
<td>15,508</td>
<td>18,751</td>
</tr>
</tbody>
</table>
<p>The 2024 chip is 47% faster on a single core than the 2022 chip, and 58% faster with all cores.</p>
<p>The clock did not do that. Max boost went from 4.5 GHz to 5.2 GHz, a 15% increase.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency-1024x753.jpg" alt="Base and max boost frequencies for Ryzen 7 5800X3D, 7800X3D and 9800X3D" width="660" height="485" class="alignnone size-large wp-image-22892" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency-1024x753.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency-300x221.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency-768x564.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-frequency.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>2022<br />
Zen 3</td>
<td>2023<br />
Zen 4</td>
<td>2024<br />
Zen 5</td>
</tr>
<tr>
<td>Base</td>
<td>3.4 GHz</td>
<td>4.2 GHz</td>
<td>4.7 GHz</td>
</tr>
<tr>
<td>Max boost</td>
<td>4.5 GHz</td>
<td>5.0 GHz</td>
<td>5.2 GHz</td>
</tr>
</tbody>
</table>
<p>The number of transistors is way up, by about 50%, from roughly 11 billion to 16 billion. Most of the extra transistors went into the core, not the cache.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors-1024x564.jpg" alt="Transistor counts for Ryzen 7 5800X3D, 7800X3D and 9800X3D, split into core, I/O and cache" width="660" height="364" class="alignnone size-large wp-image-22893" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors-1024x564.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors-300x165.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors-768x423.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-transistors.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>How do you turn extra transistors into extra performance?</p>
<p>You make the core wider, and you give it more to work with. Dispatch width went from a maximum of 6 instructions per cycle to 8. The L2 cache per core doubled, from 512 KB to 1 MB. The L1 data cache went from 32 KB to 48 KB. Integer ALUs went from 4 to 6. The reorder buffer grew from 256 to 448 entries, so the processor can keep more instructions in flight and schedule them better.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes-1024x596.jpg" alt="Cache sizes, dispatch width, integer ALUs and reorder buffer from Zen 3 to Zen 5" width="660" height="384" class="alignnone size-large wp-image-22894" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes-1024x596.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes-300x175.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes-768x447.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-core-changes.jpg 1200w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>Zen 3<br />
2022</td>
<td>Zen 4<br />
2023</td>
<td>Zen 5<br />
2024</td>
</tr>
<tr>
<td>L2 cache per core</td>
<td>512 KB</td>
<td>1 MB</td>
<td>1 MB</td>
</tr>
<tr>
<td>L1 data cache</td>
<td>32 KB</td>
<td>32 KB</td>
<td>48 KB</td>
</tr>
<tr>
<td>Dispatch width</td>
<td>6</td>
<td>6</td>
<td>8</td>
</tr>
<tr>
<td>Integer ALUs</td>
<td>4</td>
<td>4</td>
<td>6</td>
</tr>
<tr>
<td>Reorder buffer</td>
<td>256</td>
<td>320</td>
<td>448</td>
</tr>
</tbody>
</table>
<p>For data parallelism (SIMD), Zen 5 is a different machine. Zen 3 and Zen 4 had four 256-bit SIMD arithmetic units. Zen 5 has four 512-bit units. Loads and stores widened the same way: two 512-bit loads per cycle, one 512-bit store.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd.jpg"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd-1024x341.jpg" alt="SIMD arithmetic units, loads per cycle and stores per cycle from Zen 3 to Zen 5" width="660" height="220" class="alignnone size-large wp-image-22895" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd-1024x341.jpg 1024w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd-300x100.jpg 300w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd-768x256.jpg 768w, https://lemire.me/blog/wp-content/uploads/2026/09/ryzen-simd.jpg 1199w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<table>
<tbody>
<tr>
<td></td>
<td>Zen 3<br />
2022</td>
<td>Zen 4<br />
2023</td>
<td>Zen 5<br />
2024</td>
</tr>
<tr>
<td>SIMD arithmetic units</td>
<td>4 × 256-bit</td>
<td>4 × 256-bit</td>
<td>4 × 512-bit</td>
</tr>
<tr>
<td>Loads per cycle</td>
<td>2 × 256-bit</td>
<td>2 × 256-bit</td>
<td>2 × 512-bit</td>
</tr>
<tr>
<td>Stores per cycle</td>
<td>1 × 256-bit</td>
<td>1 × 256-bit</td>
<td>1 × 512-bit</td>
</tr>
</tbody>
</table>
<p>I already made the point that <a href="https://lemire.me/blog/2025/09/01/processors-are-getting-wider/">processors are getting wider</a>. This is what that looks like on a desktop chip you can buy.</p>
<p>What about the next step? Zen 6 is arriving. AMD is talking about a 256-core Epyc part (Venice) with a gigabyte of L3 cache. We do not yet know what the desktop cores will look like. It could be wild.</p>
<p>Further reading: <a href="https://www.tomshardware.com/pc-components/cpus/amds-256-core-epyc-9996-venice-claims-up-to-a-3-4x-jump-over-intel-xeon-competition-20-percent-over-nvidia-vera-zen-6-comes-with-up-to-1024mb-of-l3-16-channel-memory-and-5ghz-clock-speeds">AMD’s 256-core Epyc 9996 ‘Venice’</a> (Tom&#8217;s Hardware).</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/18/how-did-amd-ryzen-get-50-faster-in-two-years/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>How fast is C++23&#8217;s std::flat_map?</title>
		<link>https://lemire.me/blog/2026/09/16/how-fast-is-c23s-stdflat_map/</link>
					<comments>https://lemire.me/blog/2026/09/16/how-fast-is-c23s-stdflat_map/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 20:26:36 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22881</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-16-a-16.20.19-e1789590380973-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />C++23 added a new type to the standard library: std::flat_map. There is also a std::flat_set and other variants, but let me focus on std::flat_map. A flat map is a sorted vector of keys next to a vector of values. A query is a binary search over the sorted keys. You need a recent standard library: &#8230; <a href="https://lemire.me/blog/2026/09/16/how-fast-is-c23s-stdflat_map/" class="more-link">Continue reading <span class="screen-reader-text">How fast is C++23&#8217;s std::flat_map?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Capture-decran-le-2026-09-16-a-16.20.19-e1789590380973-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>C++23 added a new type to the standard library: <code>std::flat_map</code>. There is also a <code>std::flat_set</code> and other variants, but let me focus on <code>std::flat_map</code>.</p>
<p>A flat map is a sorted vector of keys next to a vector of values. A query is a binary search over the sorted keys.</p>
<p>You need a recent standard library: GCC 15&#8217;s libstdc++ has <code>std::flat_map</code>, as does LLVM&#8217;s libc++ since LLVM 20 (Apple&#8217;s clang 17 has it too).</p>
<p>Because the container is just two arrays, we can save it by copying the two arrays. Take a map from 64-bit keys to a fixed-size value:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">struct</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">point</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">double</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">x</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">double</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">y</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">};</span>
<span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">flat_map</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">point</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">42</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">1.0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">2.0</span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">7</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">3.0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">4.0</span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">1000</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">5.0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">6.0</span><span style="color: #000; font-weight: bold;">};</span>
<span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">vector</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">keys</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">keys</span><span style="color: #000; font-weight: bold;">().</span><span style="color: #000;">begin</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">keys</span><span style="color: #000; font-weight: bold;">().</span><span style="color: #000;">end</span><span style="color: #000; font-weight: bold;">());</span>
<span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">vector</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #000;">point</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">().</span><span style="color: #000;">begin</span><span style="color: #000; font-weight: bold;">(),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">().</span><span style="color: #000;">end</span><span style="color: #000; font-weight: bold;">());</span>
<span style="color: #8f5902; font-style: italic;">// keys   = {7, 42, 1000}</span>
<span style="color: #8f5902; font-style: italic;">// values = {{3, 4}, {1, 2}, {5, 6}}</span>
</code></pre>
</div>
<p>The <code>keys</code> array is <code>8 * keys.size()</code> bytes and the <code>values</code> array is <code>16 * values.size()</code> bytes: you can write them to disk with two <code>memcpy</code> or <code>write</code> calls, as they are.</p>
<p>If you loaded both arrays from a network or the disk, you can then move them straight into your <code>std::flat_map</code> like so.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">flat_map</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">point</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">back</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">sorted_unique</span><span style="color: #000; font-weight: bold;">,</span>
<span style="color: #f8f8f8;">                                    </span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">move</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">keys</span><span style="color: #000; font-weight: bold;">),</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">move</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #8f5902; font-style: italic;">// back == m</span>
</code></pre>
</div>
<p>With the <code>std::sorted_unique</code> tag, you promise that the keys are sorted and distinct. In practice, if the data comes from the network or some untrusted source, you should do some sanity testing.</p>
<p>Like the good old <code>std::map</code>, a flat map keeps its keys in sorted order, so you can iterate over it in key order. But the <code>std::map</code> is a red-black tree so you have significant storage overhead and possibly poor memory locality.</p>
<p>The downside of a <code>std::flat_map</code> is that it might be slower if you need to mutate it.</p>
<p>Let us examine the speed.</p>
<p>I use random 64-bit keys mapped to 64-bit values, GCC 16.1 with <code>-O3<br />
-march=native</code>, on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core. I report nanoseconds per operation</p>
<p>Let us start with inserting elements one at a time, in random order, into an initially empty container.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">1K</th>
<th style="text-align: right;">100K</th>
<th style="text-align: right;">1M</th>
<th style="text-align: right;">10M</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::map</code></td>
<td style="text-align: right;">93</td>
<td style="text-align: right;">253</td>
<td style="text-align: right;">465</td>
<td style="text-align: right;">1118</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::flat_map</code></td>
<td style="text-align: right;">83</td>
<td style="text-align: right;">7149</td>
<td style="text-align: right;"></td>
<td style="text-align: right;"></td>
</tr>
</tbody>
</table>
<p>Up to maybe a thousand elements or so, the <code>std::flat_map</code> is fine and maybe faster than the <code>std::map</code>. But as the size grows, the time goes up quadratically. Thus, do not use an <code>std::flat_map</code> if you need to insert millions of keys in random order. It is bad.</p>
<p>If the keys arrive in increasing order, then it is entirely different. The <code>std::flat_map</code> is much faster.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">1K</th>
<th style="text-align: right;">100K</th>
<th style="text-align: right;">1M</th>
<th style="text-align: right;">10M</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::map</code></td>
<td style="text-align: right;">31</td>
<td style="text-align: right;">67</td>
<td style="text-align: right;">125</td>
<td style="text-align: right;">202</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::flat_map</code></td>
<td style="text-align: right;">6.8</td>
<td style="text-align: right;">12</td>
<td style="text-align: right;">15</td>
<td style="text-align: right;">19</td>
</tr>
</tbody>
</table>
<p>You can also do bulk inserts. Given a batch of new pairs, in any order, the map sorts the batch and merges it with its arrays in one pass, instead of shifting the arrays once per element:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">vector</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #000;">std</span><span style="color: #ce5c00; font-weight: bold;">::</span><span style="color: #000;">pair</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">uint64_t</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">point</span><span style="color: #ce5c00; font-weight: bold;">&gt;&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">more</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{{</span><span style="color: #0000cf; font-weight: bold;">500</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">7.0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">8.0</span><span style="color: #000; font-weight: bold;">}},</span>
<span style="color: #f8f8f8;">                                                </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">3</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #0000cf; font-weight: bold;">9.0</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">10.0</span><span style="color: #000; font-weight: bold;">}}};</span>
<span style="color: #000;">m</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">insert_range</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">more</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #8f5902; font-style: italic;">// keys = {3, 7, 42, 500, 1000}</span>
</code></pre>
</div>
<p>Constructing a map from a range of random key-value pairs works the same way, and it is much faster with a <code>std::flat_map</code>:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">1K</th>
<th style="text-align: right;">100K</th>
<th style="text-align: right;">1M</th>
<th style="text-align: right;">10M</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::map</code></td>
<td style="text-align: right;">69</td>
<td style="text-align: right;">176</td>
<td style="text-align: right;">316</td>
<td style="text-align: right;">825</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::flat_map</code></td>
<td style="text-align: right;">23</td>
<td style="text-align: right;">62</td>
<td style="text-align: right;">72</td>
<td style="text-align: right;">87</td>
</tr>
</tbody>
</table>
<p>Random lookups are also much faster for large maps in part because the <code>std::flat_map</code> uses less memory.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">1K</th>
<th style="text-align: right;">100K</th>
<th style="text-align: right;">1M</th>
<th style="text-align: right;">10M</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::map</code></td>
<td style="text-align: right;">52</td>
<td style="text-align: right;">192</td>
<td style="text-align: right;">386</td>
<td style="text-align: right;">980</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::flat_map</code></td>
<td style="text-align: right;">54</td>
<td style="text-align: right;">96</td>
<td style="text-align: right;">152</td>
<td style="text-align: right;">239</td>
</tr>
</tbody>
</table>
<p>In many practical cases, the new <code>std::flat_map</code> is a better alternative to the <code>std::map</code>. It is somewhat amusing considering that you are replacing a fancy textbook data structure (red-black tree) with a trivial one.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/17">My source code is available</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/16/how-fast-is-c23s-stdflat_map/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>Subnormal floating-point numbers are expensive&#8230; on Intel processors</title>
		<link>https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/</link>
					<comments>https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 12:54:32 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22876</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_gh1asmgh1asmgh1a-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />We represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special subnormal numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance. How slow are they? Let me measure. I wrote a small C++ &#8230; <a href="https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/" class="more-link">Continue reading <span class="screen-reader-text">Subnormal floating-point numbers are expensive&#8230; on Intel processors</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_gh1asmgh1asmgh1a-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>We represent floating-point numbers using the IEEE standard. For very small numbers, the standard uses special <em>subnormal</em> numbers. Unfortunately, they have a reputation of making operations slow. Thus video game programmers and machine learning specialists sometimes avoid computing with subnormal numbers for performance.</p>
<p>How slow are they? Let me measure. I wrote a small C++ benchmark with a few kernels over arrays of 16384 values (small enough to fit in cache):</p>
<ul>
<li>multiply each value by 0.75,</li>
<li>add two arrays,</li>
<li>divide each value by 3,</li>
<li>multiply normal values by a tiny constant (2<sup>-1030</sup>) so that the <em>inputs</em> are normal but the <em>outputs</em> are subnormal,</li>
<li>a dependent chain <code>x *= 0.9999</code> repeated 16384 times.</li>
</ul>
<p>For each kernel, I feed either normal values (in <code>[0.5, 1)</code>), subnormal values, or normal values where one value in a hundred is subnormal. The compiler is allowed to autovectorize the array computations. I use GCC 15 with <code>-O3 -march=native</code> on Linux and Apple clang 17 with the same flags on macOS. I also checked with clang 21 on Linux to make sure.</p>
<p>I ran the benchmark on five processors:</p>
<ul>
<li>Intel Xeon 6975P-C (Granite Rapids), on an AWS <code>c8i.xlarge</code> instance,</li>
<li>Intel Xeon Gold 6548N (Emerald Rapids), a server in my lab,</li>
<li>AMD EPYC 9R45 (Zen 5), on an AWS <code>c8a.xlarge</code> instance,</li>
<li>AWS Graviton 5 (Arm Neoverse V3), on a <code>c9g.xlarge</code> instance,</li>
<li>Apple M4 Max.</li>
</ul>
<p>Here are the results for <code>double</code> values, in nanoseconds per element (or per step).</p>
<p><em>Intel Granite Rapids</em></p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">kernel</th>
<th style="text-align: right;">normal</th>
<th style="text-align: right;">1% subnormal</th>
<th style="text-align: right;">subnormal</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">multiply by 0.75</td>
<td style="text-align: right;">0.17</td>
<td style="text-align: right;">0.44</td>
<td style="text-align: right;">8.35</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">add two arrays</td>
<td style="text-align: right;">0.20</td>
<td style="text-align: right;">0.18</td>
<td style="text-align: right;">0.18</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">divide by 3</td>
<td style="text-align: right;">0.51</td>
<td style="text-align: right;">0.85</td>
<td style="text-align: right;">9.38</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">normal in, subnormal out</td>
<td style="text-align: right;">0.17</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">8.58</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">dependent chain</td>
<td style="text-align: right;">0.77</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">32.69</td>
</tr>
</tbody>
</table>
<p><em>Intel Emerald Rapids</em></p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">kernel</th>
<th style="text-align: right;">normal</th>
<th style="text-align: right;">1% subnormal</th>
<th style="text-align: right;">subnormal</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">multiply by 0.75</td>
<td style="text-align: right;">0.21</td>
<td style="text-align: right;">0.49</td>
<td style="text-align: right;">9.25</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">add two arrays</td>
<td style="text-align: right;">0.23</td>
<td style="text-align: right;">0.25</td>
<td style="text-align: right;">0.25</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">divide by 3</td>
<td style="text-align: right;">0.57</td>
<td style="text-align: right;">0.94</td>
<td style="text-align: right;">10.40</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">normal in, subnormal out</td>
<td style="text-align: right;">0.21</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">9.27</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">dependent chain</td>
<td style="text-align: right;">1.14</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">36.51</td>
</tr>
</tbody>
</table>
<p><em>AMD Zen 5</em></p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">kernel</th>
<th style="text-align: right;">normal</th>
<th style="text-align: right;">1% subnormal</th>
<th style="text-align: right;">subnormal</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">multiply by 0.75</td>
<td style="text-align: right;">0.07</td>
<td style="text-align: right;">0.10</td>
<td style="text-align: right;">0.08</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">add two arrays</td>
<td style="text-align: right;">0.09</td>
<td style="text-align: right;">0.09</td>
<td style="text-align: right;">0.09</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">divide by 3</td>
<td style="text-align: right;">0.11</td>
<td style="text-align: right;">0.24</td>
<td style="text-align: right;">0.25</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">normal in, subnormal out</td>
<td style="text-align: right;">0.07</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.07</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">dependent chain</td>
<td style="text-align: right;">0.66</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.88</td>
</tr>
</tbody>
</table>
<p><em>AWS Graviton 5</em></p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">kernel</th>
<th style="text-align: right;">normal</th>
<th style="text-align: right;">1% subnormal</th>
<th style="text-align: right;">subnormal</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">multiply by 0.75</td>
<td style="text-align: right;">0.17</td>
<td style="text-align: right;">0.17</td>
<td style="text-align: right;">0.16</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">add two arrays</td>
<td style="text-align: right;">0.19</td>
<td style="text-align: right;">0.20</td>
<td style="text-align: right;">0.20</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">divide by 3</td>
<td style="text-align: right;">0.30</td>
<td style="text-align: right;">0.30</td>
<td style="text-align: right;">0.30</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">normal in, subnormal out</td>
<td style="text-align: right;">0.17</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.17</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">dependent chain</td>
<td style="text-align: right;">0.91</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.91</td>
</tr>
</tbody>
</table>
<p><em>Apple M4 Max</em></p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">kernel</th>
<th style="text-align: right;">normal</th>
<th style="text-align: right;">1% subnormal</th>
<th style="text-align: right;">subnormal</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">multiply by 0.75</td>
<td style="text-align: right;">0.06</td>
<td style="text-align: right;">0.06</td>
<td style="text-align: right;">0.06</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">add two arrays</td>
<td style="text-align: right;">0.12</td>
<td style="text-align: right;">0.12</td>
<td style="text-align: right;">0.12</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">divide by 3</td>
<td style="text-align: right;">0.11</td>
<td style="text-align: right;">0.11</td>
<td style="text-align: right;">0.11</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">normal in, subnormal out</td>
<td style="text-align: right;">0.06</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.06</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">dependent chain</td>
<td style="text-align: right;">0.72</td>
<td style="text-align: right;"></td>
<td style="text-align: right;">0.75</td>
</tr>
</tbody>
</table>
<p>On Intel processors, a multiplication involving a subnormal number is about 45 to 50 times slower than a multiplication over normal numbers. A division is 18 times slower. The dependent chain, where each multiplication waits for the previous one, goes from about 1 ns to over 30 ns per step. A normal multiplication in the dependent chain has a latency of 4 cycles. With a subnormal, it has a latency of 128 cycles. It does not matter whether the subnormal is an input or an output: multiplying normal numbers into a subnormal result is just as slow as multiplying subnormal numbers. The exception is additions and subtractions: they run at full speed. Even if subnormals are rare (1%), the cost on Intel can be significant because when the compiler vectorizes the computation, a single subnormal can slow down a whole block of computations.</p>
<p>AMD does much better. On Zen 5, multiplications and additions run at full speed regardless of the inputs. The dependent multiplication chain is a third slower (0.66 ns to 0.88 ns per step): the multiplier needs an extra cycle or so to handle a subnormal. Divisions are about twice as slow. Interestingly, with divisions, having one subnormal in a hundred is almost as slow as having all subnormals. The two Arm processors, the Graviton 5 and the Apple M4 Max, do not care at all. Subnormal numbers are handled at full speed.</p>
<p>Thus it appears that on the latest AMD and ARM processors subnormals might not be a concern. But they remain very much a performance issue under Intel processors.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/15">My source code is available</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/15/subnormal-floating-point-numbers-are-expensive-on-intel-processors/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>The four-colour theorem was only the start</title>
		<link>https://lemire.me/blog/2026/09/11/the-four-colour-theorem-was-only-the-start/</link>
					<comments>https://lemire.me/blog/2026/09/11/the-four-colour-theorem-was-only-the-start/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Fri, 11 Sep 2026 19:53:31 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22869</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/kALnB-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Mathematicians are unhappy about OpenAI. Several influential mathematicians wrote an open letter. The gist of their argument is that they form a community that trains young people. When AI started producing breakthroughs on hard mathematical problems, I asked what a very smart 17-year-old would feel. Do you still choose a math major and train yourself &#8230; <a href="https://lemire.me/blog/2026/09/11/the-four-colour-theorem-was-only-the-start/" class="more-link">Continue reading <span class="screen-reader-text">The four-colour theorem was only the start</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/kALnB-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p><iframe loading="lazy" width="560" height="315" src="https://www.youtube.com/embed/exa7EX5YYeA?si=n58hi6V7oVusOB27" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"></iframe></p>
<p><span class="css-1jxf684 r-bcqeeo r-1ttztb7 r-qvutc0 r-poiln3"><br />
Mathematicians are unhappy about OpenAI. <a href="https://mathandai.org">Several influential mathematicians wrote an open letter.</a> The gist of their argument is that they form a community that trains young people. </span></p>
<p>When AI started producing breakthroughs on hard mathematical problems, I asked what a very smart 17-year-old would feel. Do you still choose a math major and train yourself to prove difficult results by hand?</p>
<p>This crisis has been coming for a long time. When I was a kid, the four-colour theorem was proved by a computer, in 1976. An intense debate followed. Does it count as a proof?</p>
<p>I wrote my PhD thesis using symbolic algebra software. To my knowledge, nobody then would publish their scripts. I tried to include mine. I was told it would make me look bad.</p>
<p>The letter says:</p>
<blockquote><p>&#8220;In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas.&#8221;</p></blockquote>
<p>They worry that kids will not choose to become old-school mathematicians. That is a reasonable fear. Yet they do not seem to consider that some kids might still do mathematics, just in a very different way.</p>
<p><a href="https://sites.math.rutgers.edu/~zeilberg/Opinion94.html">The famous mathematician Doron Zeilberger announced the problem in 2009</a>:</p>
<blockquote><p>&#8220;Teaching computers how to discover and prove mathematical results is certainly the way to go, and I believe that mathematicians who continue to do pure human, pencil-and-paper, computer-less, research, are wasting their time.&#8221;</p></blockquote>
<p>The letter implies that people, because of AI, will stop having ideas, or will stop taking the time to understand the issues. If that were true, mathematics would continue only on computers, with no humans interested, or it would stop. Both scenarios assume that people care about mathematics only when they can claim credit. </p>
<p>I am not sure why I cannot study a proof generated by AI if I want to. I can give talks about it. What becomes less likely is the reward of having been the one who proved the result. </p>
<p>The rest of the letter makes a big deal of credit. What if the AI builds on what it read and does not give proper credit? Where is the evidence that AI is worse than human beings at citing sources? </p>
<p>The letter notes that mathematics contributes to society. It never considers that faster progress might increase those contributions. </p>
<p>My stance is simple. </p>
<p>Mathematicians, software developers, engineers, lawyers, physicians will all learn to work with AI. You cannot put the genie back in the bottle. Only a totalitarian world government could try, and some of us would rather avoid that outcome. </p>
<p>Difficult proofs will now be built with AI, just as most code will be written with AI. People who insist on pen and paper should view themselves as artists. </p>
<p>Other mathematicians will work with AI. They will have no trouble finding interested kids. They will contribute to society.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/11/the-four-colour-theorem-was-only-the-start/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Fear Is Not an Argument</title>
		<link>https://lemire.me/blog/2026/09/10/fear-is-not-an-argument/</link>
					<comments>https://lemire.me/blog/2026/09/10/fear-is-not-an-argument/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 18:23:42 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22845</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/wm0AF-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />We are told that AI entities much like ChatGPT might soon kill us all. The statement is vague and unfalsifiable. It might be true, it might be false. People with credentials (e.g., Turing Award recipient Yoshua Bengio) believe it. Many still remember the Year-2000 bug. Our computers used two-digit coding for dates, and some software &#8230; <a href="https://lemire.me/blog/2026/09/10/fear-is-not-an-argument/" class="more-link">Continue reading <span class="screen-reader-text">Fear Is Not an Argument</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/wm0AF-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>We are told that AI entities much like ChatGPT might soon kill us all. The statement is vague and unfalsifiable. It might be true, it might be false. People with credentials (e.g., Turing Award recipient Yoshua Bengio) believe it.</p>
<p>Many still remember the Year-2000 bug. Our computers used two-digit coding for dates, and some software could get confused. At the time, experts worried that a bug in dates might trigger nuclear Armageddon or an infrastructure collapse. At the very least, planes could fall.</p>
<p dir="auto">Americans had a moral panic over alcohol that lasted more than a century. It began with pledges of moderation. From 1920 the Volstead Act banned nearly all legal drink. It collapsed in 1933, at great expense.</p>
<p id="ember371" class="ember-view reader-text-block__paragraph">The Club of Rome predicted mass starvation. As an answer, we sterilized by force millions of Indians, and introduced the devastating one-child policy in China. The authors were never held accountable. The projections were purely mathematical, unescapable. They said. But also totally wrong and silly. Yet, we listened to them and caused great harm.</p>
<p>End-of-the-world scenarios are nothing new. Pretty much all civilizations have lived with various such predictions.</p>
<p>Some people are offended by my comparisons. I truly do not mean to offend. But the fact that disagreeing can lead to deep offense is, by itself, a sign that we face a moral issue. There is a sense in which you must agree that these intense fears are warranted.</p>
<p>Many will remember that when OpenAI first developed GPT-2, they told the world that it was too dangerous to release. Year after year, we were warned that the next iteration of it would doom us all.</p>
<p>A large language model takes tokens (words) in and outputs tokens (words). The big models can take many, many tokens in. And they do much compute. And they are based on clever ideas like vector embeddings. But, ultimately, no large language model can do anything but output tokens. You can build a better model, but the model itself does not &#8216;learn&#8217;. It is a fixed set of weights. If you take a model that has been used for months, and always feed the same tokens, you will get the same results (up to some randomness).</p>
<p>I fear that some people exploit the fact that people cannot grasp how conceptually simple a language model is. In any case, most people don&#8217;t understand how things work.</p>
<p>Things become interesting because these models can be hooked up to tools. So you can tell your language model that whenever it outputs &#8216;boom&#8217;, then a nuclear weapon will be launched. And if you hook up a nuclear weapon, then, certainly, you may start a nuclear war. So don&#8217;t do that.</p>
<p>One unproven thesis is that the models (that take in tokens and output tokens) will &#8216;decide&#8217; to acquire access to these nuclear weapons, maybe through a subterfuge. What does that mean? It is always conveniently vague. In the movie WarGames (1983), a teenager uses his computer to access a computer in charge of nuclear weapons. He almost wipes out humanity. Could this happen by accident with a kid using a language model hooked up to the Internet? But if it happens, we should blame whoever hooked up a deadly computer to the Internet.</p>
<p>Of course, any technology is inherently dangerous. Invent the bow to go hunting, and someone might soon turn the bow against you. Invent the engine, and one might soon build tanks and destroy nations. Develop nuclear technology, and one might soon raze your cities.<br />
 <br />
Yet that is not what is at stake in these discussions. The concrete threats are not ascertained and addressed. No doubt, there are some people doing this work, hopefully in the US military. What if an adversary can take control of the economy or military installations? What if an AI agent goes rogue? It is worth investing time in designing defenses.</p>
<p>What we have instead is something of the sort:</p>
<ol>
<li>A vague but global threat. It could be a fatal virus engineered in a lab, a climate catastrophe, a fatal bug affecting all our software, an alien invasion, a rogue AI, Jews taking over our institutions.</li>
<li>A few people come forward and they offer to save us. Importantly we must give them resources and influence. Ultimately, they seek a totalitarian solution: everyone must be made to agree so that we can be saved.</li>
<li>As the process unfolds, people with an opposing viewpoint are described as a danger. They must be silenced and discredited. Eventually, it can become moral to exaggerate the threat or to rewrite counterpoints. People must be made to understand one way or another.</li>
</ol>
<p>In this instance, I refer to people who advocate that AI will doom us as AI Doomers. These people tend to carry a totalitarian ideology. Their ideas will only work if everyone is made to agree. And it would severely restrict the freedom of billions of people, although they usually present it differently.</p>
<p>Doomers do not have bad intentions. On the contrary, they are often really out there to save the world. But good intentions do not, in any way, justify the means nor guarantee a good outcome.</p>
<p><span>Human beings reason based on cultural knowledge. For centuries or more, totalitarian ideas have led to ruin or pain. We ignore the warnings at our peril. It is one after the other: alcohol, the need to restrict the number of children, and so on.</span></p>
<p>But shouldn’t we just be prudent and adopt their views, just in case? It is a fallacious argument. Members of the intellectual elite have a tendency to fall for the kind of hubris where they think that, if only they were given more power, the world would be better off. It is rarely true. Thomas Sowell has an excellent book on the topic, Intellectuals and Society. He makes the case that intellectuals often promote harmful ideas, at no cost to themselves. Rationally, we should therefore be cautious.</p>
<p><span>You are not safer without technology. In fact, the risk of human extinction is assuredly higher if we are poorer and have less technology.</span></p>
<p>Is this unprecedented? The printing press was unprecedented. Arabic numbers were unprecedented. Maybe we should go back to Roman numerals, to be safe. Fear of what is without precedent soon becomes indistinguishable from an anti-innovation stance.</p>
<p>What if you do not like the people who lead the big AI companies like OpenAI and Anthropic. Maybe you think that these billionaires are a danger. And you might be right. But consider the history of humanity. Wealthy people have primarily caused harm through the promotion of bad ideas. The mass murders are almost invariably derived from politics. Stalin, Hitler, Mao.</p>
<div class="" data-block="true" data-editor="3bk51" data-offset-key="8gbvc-0-0">
<div data-offset-key="8gbvc-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="8gbvc-0-0">Are the fears grounded in reason or is some of it signaling? We have been deploying AI-enhanced drones in Ukraine for two years. Once we designate the target, they engage, autonomously. </span>At a strategic level, Palantir’s Maven Smart System is used to pick targets. It has been deployed against Iran. I have not seen much opposition to drone attacks by Ukraine against Russia, at least in the West. I cannot recall any AI Doomer denouncing Ukraine&#8217;s drones and some even endorsed them. Yet it is largely the West that is funding these drones. If you fear rogue AIs, it seems that drones able to engage a target on their own would cause enormous worry&#8230; But it would be morally inconvenient in the West to criticize the use of AI against Russia&#8230; and so, the AI Doomers are largely silent. They do not lobby their governments to require Ukraine to abstain from building AI-driven weapons. That is another sign that they do not act on reason, but, rather, on moral grounds. You might argue that these drones are not entirely autonomous, since, as far as we know, they do not pick their own targets. But ChatGPT also does not pick the prompts. The hypocrisy is par for the course <span>for many</span>. It is akin to the governor of California dining at a fancy restaurant while a stay-at-home order is in effect. Or the prime minister of Canada ranking in the top air travelers of all time, while advocating for a carbon-neutral lifestyle. You can be quite sure that many of the AI Doomers are heavy users of AI services and, in some cases, investors. It is telling you that their stance is primarily moral. Expressing fear of AI can become a form of virtue signaling. It is a convenient stance, but you would not go so far as to stop using AI, and shut down Ukraine&#8217;s drones. </p>
<p><span> In some sense, there is also a form of luxury beliefs involved. A luxury belief is a belief that makes you look good and cost you nothing, while it might harm people who are not so well-off. AI in the form of ChatGPT is proving to be a great equalizer. My plumber has access to AI that is comparable to that of a billionaire. It has the potential to serve as a superior tutor to all these kids who are left out. Many of the people engaging in the promotion of fear are either upper middle class or better. Many of them pay little attention to the fact that many of the beneficiaries of the huge investments in AI have been the men building the data centres. Thus far, AI has been great at creating jobs for blue collars. That&#8217;s not nothing.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/HRjc6QSXYAIlS2k.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/HRjc6QSXYAIlS2k.png" alt="" width="600" height="731" class="alignnone size-full wp-image-22863" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/HRjc6QSXYAIlS2k.png 600w, https://lemire.me/blog/wp-content/uploads/2026/09/HRjc6QSXYAIlS2k-246x300.png 246w" sizes="auto, (max-width: 600px) 100vw, 600px" /></a><br />
</span></div>
</div>
<p>
Throughout much of the world, we are facing demographic collapse and an inverted age pyramid. Soon there may be just one worker per retiree. Choosing to have fewer kids has consequences and we are about to face them. AI might be a way out of significant problems. If you are otherwise wealthy, that might not be a significant concern to you. But for the least fortunate, it might turn out to be quite a problem. Who will take care of you when you are sick? AI and robotics might prove useful to reduce suffering, don&#8217;t we think?</p>
<p>Further, we need to consider how powerful people might use the fear of AI for their own purposes. It is entirely credible that the owners of large companies could promote fear so that they get to write the regulations that will keep out their competitors, or merely as a form of cheap marketing.</p>
<p>How do I know that it is moral? Because there cannot be reasoned debate about a moral question. The facts are obvious or you are a bad person. Whenever there is a complex question, one that involves predicting the future, that cannot be discussed, unless it is in agreement with the side of fear, then you are very likely in a moral question. « Don&#8217;t you see, AI will soon kill all of us, it is obvious. » No explanation can be demanded.</p>
<p><span>You might accuse me of, in turn, promoting fear. But it should be obvious that I am doing no such thing. What I am encouraging rather is the use of reason. I am forced to give examples where inciting fear has led to disastrous effects, but my hope is that it will lead my reader to sit and reflect.</span></p>
<p><span>To my friends who fear AI, I urge you. Use reason. Do the work. Do not rely on hasty thought experiments. Work out the details. Think. Think about the countermeasures. And, please, do not include abstract thinking machines. A language model is a box that takes in token and produces tokens. Nothing more.</span></p>
<p>And for the rest of us. Let us build. Let us bring prosperity. Let us hasten the cure for cancer. Let us dream of exploring our solar system.</p>
<p><strong>Further reading</strong>.</p>
<ul>
<li style="list-style-type: none;">
<ul>
<li><span>Nirit Weiss-Blatt, “<a href="https://www.aipanic.news/p/the-weak-foundations-of-ai-doomsday">The Weak Foundations of AI Doomsday”</a> — AI Panic, May 15, 2026.</span></li>
<li><span></span>Cal Newport, “<a href="https://www.nytimes.com/2026/06/17/opinion/ai-dangerous-openai-anthropic.html">Dear A.I. Companies, the Doom Trolling Needs to Stop</a>” — New York Times, June 17, 2026</li>
<li>Andrew Orlowski, “<a href="https://www.telegraph.co.uk/news/2026/06/05/anthropics-ai-doom-predictions-hype-share-price/">Anthropic’s doom predictions are merely hype intended to make AI look important</a>” — The Telegraph, June 5, 2026.</li>
</ul>
</li>
</ul>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/10/fear-is-not-an-argument/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>A quick overview of atomics in C</title>
		<link>https://lemire.me/blog/2026/09/09/a-quick-overview-of-atomics-in-c/</link>
					<comments>https://lemire.me/blog/2026/09/09/a-quick-overview-of-atomics-in-c/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 20:41:53 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22835</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/image-4-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />If you write in C, by default, you use a single thread. Extra cores do not help until you create more threads. However, if you include the header &#60;threads.h&#62;, you can pass a function to thrd_create, and wait for it with thrd_join. #include &#60;threads.h&#62; #include &#60;stdio.h&#62; int worker(void *arg) { printf("hello from thread %d\n", *(int &#8230; <a href="https://lemire.me/blog/2026/09/09/a-quick-overview-of-atomics-in-c/" class="more-link">Continue reading <span class="screen-reader-text">A quick overview of atomics in C</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/image-4-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>If you write in C, by default, you use a single thread. Extra cores do not help until you create more threads. However, if you include the header <code>&lt;threads.h&gt;</code>, you can pass a function to <code>thrd_create</code>, and wait for it with <code>thrd_join</code>.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">#include</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">&lt;threads.h&gt;</span>
<span style="color: #8f5902; font-style: italic;">#include</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">&lt;stdio.h&gt;</span>
<span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">worker</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">arg</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">printf</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #4e9a06;">"hello from thread %d\n"</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #000;">arg</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
<span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">main</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">thrd_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">t</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">id</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">thrd_create</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">t</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">worker</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">id</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">thrd_join</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">t</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">NULL</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>Be warned that C11 threads are an optional feature. If the macro <code>__STDC_NO_THREADS__</code> is defined, you do not have them. Apple&#8217;s C library has never shipped <code>&lt;threads.h&gt;</code>, so the program above does not compile on macOS, and glibc only added it in version 2.28 (2018). On such systems you fall back on POSIX threads (<code>pthread_create</code>, <code>pthread_join</code>).</p>
<p>Once you have more than one thread, they may share memory. If two threads access the same non-atomic variable with no ordering between them, and at least one of them writes, the C language calls that a data race. In other words, it is unsafe.</p>
<p>If you have a variable and it is effectively constant, then it is fine to share it. But as soon as anyone changes it, then it might get corrupted. If it is not guarded somewhat, you are in trouble.</p>
<p>To be clear, that is what the C programming language says. I don&#8217;t mean that it will happen on your machine.</p>
<p>To get a better behavior, we can use atomic variables. In C, you have the <code>&lt;stdatomic.h&gt;</code> header.</p>
<p>An atomic integer is never garbage. You always read a value that was once written.</p>
<p>In practice, on most computers you might use today, aligned 8-, 16-, 32- and 64-bit loads and stores are atomic. The C language does not care about that, so if you don&#8217;t specifically require atomicity, you might get in trouble with your C compiler.</p>
<p>The next funny problem is that instructions can be reordered. When you write:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>x = 2
y = 3
</code></pre>
</div>
<p>This may not happen in this sequence. The variable <code>y</code> might be set before the variable <code>x</code>. You may wonder why this is allowed at all. The fundamental reason is that our processors are quite complex. They have layers of buffers and they can execute multiple instructions at once. They can issue several memory loads or stores at once.</p>
<p>By default, in C, atomic accesses are all ordered. It is as if there is an oracle that watches all threads and comes up with a consistent story where everything is in order. This can be expensive, so we prefer not to do it that way.</p>
<p>At the other extreme is the relaxed model: your reads and stores are not garbage, and a given atomic still has one modification order (you will not see 1 and then 0 if the counter only went from 0 to 1), but there is no ordering with respect to other memory.</p>
<p>So we use something intermediate, the release and acquire semantics. They are ordering barriers. A strict barrier would be &#8216;everything before me really happens before me, and everything after me really happens after me&#8217;. (Where &#8216;really happens&#8217; refers to visible effects, the hardware and compiler are allowed to cheat as long as you don&#8217;t catch them.) It is a bit too strong. So we split it in two parts: release and acquire. Intuitively, release means &#8216;if you see me, you see all the stuff before me&#8217;. Acquire means &#8216;I take that package, and everything I do after this load really happens after it&#8217;.</p>
<p>Consider the case where you have a resource (such as a block of allocated memory). You share this resource, but count how many people have access to it. When the counter goes to zero, you free the resource.</p>
<p>One thread could do&#8230;</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>access resource
decrement counter // I won't need it anymore
</code></pre>
</div>
<p>You see these operations happening one after the other, but they may not execute this way. It is possible that they overlap, or even that the decrement occurs before the access. It is entirely safe in a single threaded context.</p>
<p>Anyhow, so the following could happen</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>decrement counter // I won't need it anymore
access resource
</code></pre>
</div>
<p>But what if you have a second thread that does:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
<span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">//</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">I</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">won' need it anymore</span>
<span style="color: #a40000; border: 1px solid #EF2929;">if (counter is zero)</span>
<span style="color: #a40000; border: 1px solid #EF2929;">  free(resource)</span>
</code></pre>
</div>
<p>You could have this interplay:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">resource</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
</code></pre>
</div>
<p>That would be a bug.</p>
<p>So what you first do is make the decrement a &#8216;release access&#8217; which means that operations that come before it cannot be reordered after it. So if you do it this way&#8230;</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>access resource
decrement counter using release
</code></pre>
</div>
<p>Then it is not possible that we &#8216;see&#8217; the operations as if they happened in the reverse order.</p>
<p>But then we have a second problem. Release is enough for this thread: we cannot still be using the resource after we drop it. It does not tell the last owner that everyone else is finished. The last decrement is itself a release, so it does not observe the other threads&#8217; releases. Without an acquire, that last thread can call <code>free</code> while another thread&#8217;s earlier access is not yet done.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">using</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">release</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">using</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">release</span>
<span style="color: #ce5c00; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #ce5c00; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">resource</span><span style="color: #000; font-weight: bold;">)</span>
</code></pre>
</div>
<p>Thread 2 did its access before its release decrement, but thread 1 never acquired, so it is not required to see that access as finished before <code>free</code>.</p>
<p>So we need the counterpart to a release, an acquire. The last owner acquires before it frees, and that pairs with everyone else&#8217;s release:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
<span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">with</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">release</span>
<span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">(</span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">is</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #4e9a06;">)</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">acquire</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">barrier</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">//</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">see</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">that</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">everyone</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">else</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">is</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">done</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">free</span><span style="color: #4e9a06;">(</span><span style="color: #000;">resource</span><span style="color: #4e9a06;">)</span>
</code></pre>
</div>
<p>Alternatively, you could do this.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">access</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">resource</span>
<span style="color: #000;">decrement</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">with</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">release</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">and</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">acquire</span>
<span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">(</span><span style="color: #000;">counter</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">is</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">zero</span><span style="color: #4e9a06;">)</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">free</span><span style="color: #4e9a06;">(</span><span style="color: #000;">resource</span><span style="color: #4e9a06;">)</span>
</code></pre>
</div>
<p>The two are equivalent, but they are not necessarily equally cheap.</p>
<p>So let us consider a nice example. Let us build a small array that several threads can share. If you are the only owner, you overwrite an element in place. If not, you copy, then you update the copy. That is called copy-on-write. It is a really nice idea that you will find in many important systems.</p>
<p>We start with the type.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">#include</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">&lt;assert.h&gt;</span>
<span style="color: #8f5902; font-style: italic;">#include</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">&lt;stdatomic.h&gt;</span>
<span style="color: #8f5902; font-style: italic;">#include</span><span style="color: #f8f8f8;"> </span><span style="color: #8f5902; font-style: italic;">&lt;stdlib.h&gt;</span>
<span style="color: #8f5902; font-style: italic;">#define STR_SIZE 16</span>
<span style="color: #204a87; font-weight: bold;">typedef</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">struct</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">atomic_int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">STR_SIZE</span><span style="color: #000; font-weight: bold;">];</span>
<span style="color: #000; font-weight: bold;">}</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">shared_array</span><span style="color: #000; font-weight: bold;">;</span>
</code></pre>
</div>
<p>The payload is a plain <code>int</code> array. Only <code>refs</code> is atomic. That is deliberate. We never write <code>values</code> while another thread might be reading them.</p>
<p>We create an instance like so.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">str_new</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">malloc</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">sizeof</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">NULL</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">NULL</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">atomic_init</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">STR_SIZE</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>The <code>atomic_init</code> is not an atomic access in the memory-model sense. Nobody else has the pointer yet, so there is no other thread to race with. The caller owns one reference. It is just how we initialize an <code>atomic_int</code>.</p>
<p>Here is how we might naively release an instance.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">// not real code</span>
<span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">obj_release</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">auto</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">ref</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">ref</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">!=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #8f5902; font-style: italic;">// we are the last copy</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>What is the problem with this code?</p>
<p>The load and the decrement are two operations. Two threads can both read 2, both subtract, the counter hits zero, and nobody frees: the resource leaks. Write the check the other way around, decrementing first and then testing whether the counter is zero, as in the pseudocode above, and you get the mirror-image bug instead: with <code>refs</code> at 2, one thread decrements to 1, the other decrements to 0, both then read 0, and both call <code>free</code>. You need one atomic subtract that hands you the previous value: only the thread that saw 1 was last.</p>
<p>So you could try</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">obj_release</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">atomic_fetch_sub_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_relaxed</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">!=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>But suppose you have two owners, so that <code>refs</code> is 2. And you have two threads doing</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">];</span>
<span style="color: #000;">obj_release</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>One of them will call <code>free</code>, but the order could be</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000; font-weight: bold;">...
[</span><span style="color: #000;">thread1</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">];</span>
<span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">atomic_fetch_sub_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_relaxed</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">atomic_fetch_sub_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_relaxed</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread1</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">];</span>
</code></pre>
</div>
<p>It is a bit confusing because things are not happening in order within thread 2:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">atomic_fetch_sub_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_relaxed</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">thread2</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #0000cf; font-weight: bold;">4</span><span style="color: #000; font-weight: bold;">];</span>
</code></pre>
</div>
<p>But this is allowed.</p>
<p>So what we can do is put a release on the <code>atomic_fetch_sub_explicit</code> and then an acquire right before the <code>free</code>.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">void</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">obj_release</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">atomic_fetch_sub_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_release</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">!=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">atomic_thread_fence</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">memory_order_acquire</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">free</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>The release on every decrement means &#8220;I am done with the payload.&#8221; The acquire fence, only on the last owner, means &#8220;I have seen that everyone else is done.&#8221; Then <code>free</code> is safe.</p>
<p>That release does double duty, as we are about to see. It is also what lets the last remaining owner write to the payload in place.</p>
<p>If a thread wants another reference to the same instance, it only needs a relaxed access.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">str_retain</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">atomic_fetch_add_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_relaxed</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>Why relaxed? Because the caller already holds a reference, so the object cannot be freed under us: the last owner would need our reference to be gone first.</p>
<p>We can now write <code>update</code>. It consumes the caller&#8217;s reference and returns a reference to the array that contains the new value, which may or may not be the same object. After you call it, you must not touch the pointer you passed in. There is one exception: if a copy was needed and the allocation failed, it returns <code>NULL</code> and leaves the caller&#8217;s reference to <code>o</code> untouched, so you still own it and must still release it.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">update</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">size_t</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">idx</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">value</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">assert</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">idx</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">STR_SIZE</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">atomic_load_explicit</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">refs</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">memory_order_acquire</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">idx</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">value</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">shared_array</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">new_o</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">str_new</span><span style="color: #000; font-weight: bold;">();</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">new_o</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">NULL</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">NULL</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">STR_SIZE</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">new_o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">];</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">new_o</span><span style="color: #ce5c00; font-weight: bold;">-&gt;</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">idx</span><span style="color: #000; font-weight: bold;">]</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">value</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">obj_release</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">o</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">new_o</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>If the load reads 1, we are the only owner. No other thread holds a reference, so we can write <code>values[idx]</code> in place.</p>
<p>The load is an acquire. When the load reads 1, it may read the value written by the release decrement of the last <em>other</em> owner to drop out. Everything that thread did with <code>values</code> happens before our write. Nothing in our code appearing after such as <code>o-&gt;values[idx] = value</code> may move before it. No other thread still holds a reference, so the write does not race with a concurrent reader. Later, after a retain, other threads can see it.</p>
<p>On x64, acquire and release are effectively free at the CPU: ordinary loads already behave like acquire, ordinary stores like release. You still have to write them in C, or the compiler may reorder the payload accesses. ARM has a weaker memory model so the acquire/release require different instructions (<code>ldapr</code>, <code>ldaddl</code>) which may incur a small perforamnce hit.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/09">The code is available.</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/09/a-quick-overview-of-atomics-in-c/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>AI programming: a layered model</title>
		<link>https://lemire.me/blog/2026/09/05/ai-programming-a-layered-model/</link>
					<comments>https://lemire.me/blog/2026/09/05/ai-programming-a-layered-model/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 05 Sep 2026 14:02:12 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22831</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_km32u4km32u4km32-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />In the late 1960s and 1970s, people like David Parnas faced a problem. A decade earlier there were almost no programmers. Suddenly there were hordes of inexperienced ones. What could have been a golden era was turning into a mess: far more software, much of it falling apart.   It sent Edsger Dijkstra into a &#8230; <a href="https://lemire.me/blog/2026/09/05/ai-programming-a-layered-model/" class="more-link">Continue reading <span class="screen-reader-text">AI programming: a layered model</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/Gemini_Generated_Image_km32u4km32u4km32-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><div class="" data-block="true" data-editor="egmet" data-offset-key="9b314-0-0">
<div data-offset-key="9b314-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9b314-0-0">In the late 1960s and 1970s, people like David Parnas faced a problem. A decade earlier there were almost no programmers. Suddenly there were hordes of inexperienced ones. What could have been a golden era was turning into a mess: far more software, much of it falling apart. </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="d2c0s-0-0">
<div data-offset-key="d2c0s-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="d2c0s-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="9e8td-0-0">
<div data-offset-key="9e8td-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9e8td-0-0">It sent Edsger Dijkstra into a depression. Does this sound familiar? </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="fqomm-0-0">
<div data-offset-key="fqomm-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="fqomm-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="8rvod-0-0">
<div data-offset-key="8rvod-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="8rvod-0-0">AI-assisted coding is producing far more code. Whether the projects will work or crumble remains to be seen. There is a danger. </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="djbp6-0-0">
<div data-offset-key="djbp6-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="djbp6-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="4tm2v-0-0">
<div data-offset-key="4tm2v-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="4tm2v-0-0">I&#8217;d like to propose the layered model. </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="lqkp-0-0">
<div data-offset-key="lqkp-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="lqkp-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="beh5r-0-0">
<div data-offset-key="beh5r-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="beh5r-0-0">Keep a small core that changes slowly and on purpose. For that part you actually read the code. You insist on tests. You can use AI assistance, but there is no vibe coding allowed. </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="e0m4-0-0">
<div data-offset-key="e0m4-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="e0m4-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="9umb1-0-0">
<div data-offset-key="9umb1-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="9umb1-0-0">Everything else can move fast. There will be bugs, but the AI fixes them quickly. </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="elnou-0-0">
<div data-offset-key="elnou-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="elnou-0-0"> </span></div>
</div>
<div class="" data-block="true" data-editor="egmet" data-offset-key="1u4p7-0-0">
<div data-offset-key="1u4p7-0-0" class="public-DraftStyleDefault-block public-DraftStyleDefault-ltr"><span data-offset-key="1u4p7-0-0">Dependencies should be one way: the outer layers depend on the core. The core cannot depend on the outer layers.</span></div>
</div>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/05/ai-programming-a-layered-model/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Python sets and dictionaries can have quadratic-time performance</title>
		<link>https://lemire.me/blog/2026/09/03/python-sets-and-dictionaries-can-have-quadratic-time-performance/</link>
					<comments>https://lemire.me/blog/2026/09/03/python-sets-and-dictionaries-can-have-quadratic-time-performance/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 14:01:45 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22821</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/quadratic-150x150.webp" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />In Python, the dict data structure is the conventional key-value structure. E.g., you might store a list of names as keys and have their phone numbers as values. Valentin Ignatev wrote this amusing post on X: It is indeed widely believed that, in the strict sense, the dict data structure and its companion, the set &#8230; <a href="https://lemire.me/blog/2026/09/03/python-sets-and-dictionaries-can-have-quadratic-time-performance/" class="more-link">Continue reading <span class="screen-reader-text">Python sets and dictionaries can have quadratic-time performance</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/09/quadratic-150x150.webp" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>In Python, the <code>dict</code> data structure is the conventional key-value structure. E.g., you might store a list of names as keys and have their phone numbers as values. Valentin Ignatev wrote this amusing post on X:</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/09/xpost.webp"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/09/xpost-612x1024.webp" alt="" width="612" height="1024" class="alignnone size-large wp-image-22824" srcset="https://lemire.me/blog/wp-content/uploads/2026/09/xpost-612x1024.webp 612w, https://lemire.me/blog/wp-content/uploads/2026/09/xpost-179x300.webp 179w, https://lemire.me/blog/wp-content/uploads/2026/09/xpost-768x1286.webp 768w, https://lemire.me/blog/wp-content/uploads/2026/09/xpost.webp 908w" sizes="auto, (max-width: 612px) 100vw, 612px" /></a><img decoding="async" alt="" src="xpost.webp" /></p>
<p>It is indeed widely believed that, in the strict sense, the <code>dict</code> data structure and its companion, the set data structure, are O(1), meaning that as you increase the size of the data structure, the time to insert or query a key remains constant.</p>
<p>Let us examine the claim.</p>
<p>A hash function is a function from objects (like strings, integers, etc.) to integer values. We typically expect hash functions to be random-like, although they should always map the same object to the same integer within the current program execution. From hash functions, we construct hash tables:</p>
<ol>
<li>Create an array of buckets.</li>
<li>Given an object, apply the hash function to map it to a bucket.</li>
<li>Store the object in the bucket. When the bucket is already occupied, use some other trick (such as using a nearby bucket).</li>
</ol>
<p>If everything goes well, access and insertion in a hash table take nearly constant time, meaning that the time they take is independent of the size of the hash table.</p>
<p>This can be almost true in many instances. However, it is not formally true. There are many reasons why it is false. For example, if your data structure grows, it might be necessary to reallocate, which will typically take time proportional to the size of the data structure. But we also have the issue of collisions. A collision is what happens when two objects have the same hash value. When we use hash tables, we assume that collisions are uncommon. But it is not difficult to create many of them by picking our objects carefully.</p>
<p>In Python, <code>set</code> and <code>dict</code> are hash tables. I can &#8216;easily&#8217; make my version of Python crumble:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">M</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #000; font-weight: bold;">(</span><span style="color: #0000cf; font-weight: bold;">1</span> <span style="color: #ce5c00; font-weight: bold;">&lt;&lt;</span> <span style="color: #0000cf; font-weight: bold;">61</span><span style="color: #000; font-weight: bold;">)</span> <span style="color: #ce5c00; font-weight: bold;">-</span> <span style="color: #0000cf; font-weight: bold;">1</span>
<span style="color: #000;">values</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">i</span> <span style="color: #ce5c00; font-weight: bold;">*</span> <span style="color: #000;">M</span> <span style="color: #204a87; font-weight: bold;">for</span> <span style="color: #000;">i</span> <span style="color: #204a87; font-weight: bold;">in</span> <span style="color: #204a87;">range</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">,</span> <span style="color: #000;">n</span> <span style="color: #ce5c00; font-weight: bold;">+</span> <span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)]</span>
<span style="color: #000;">s</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #204a87;">set</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">)</span>                       <span style="color: #8f5902; font-style: italic;"># insertions</span>
<span style="color: #000;">count</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #204a87;">sum</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">v</span> <span style="color: #204a87; font-weight: bold;">in</span> <span style="color: #000;">s</span> <span style="color: #204a87; font-weight: bold;">for</span> <span style="color: #000;">v</span> <span style="color: #204a87; font-weight: bold;">in</span> <span style="color: #000;">values</span><span style="color: #000; font-weight: bold;">)</span>   <span style="color: #8f5902; font-style: italic;"># checks</span>
</code></pre>
</div>
<p>If the insertions and the checks are constant-time operations, then the whole construction and the entire check should take linear time. I ran this on an Apple M4 Max with Python 3.14, reporting the median of three runs.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;">n</th>
<th style="text-align: right;">time</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">1000</td>
<td style="text-align: right;">4.8 ms</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">2000</td>
<td style="text-align: right;">15.5 ms</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">4000</td>
<td style="text-align: right;">65.5 ms</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">8000</td>
<td style="text-align: right;">257 ms</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">16000</td>
<td style="text-align: right;">1072 ms</td>
</tr>
</tbody>
</table>
<p>The time roughly quadruples each time <em>n</em> doubles. That is quadratic time, not linear time. The membership checks behave the same way: 1066 ms at <em>n</em> = 16000. At a hundred thousand elements, building the set takes 45 seconds.</p>
<p>But could we create a hash table that would be truly constant-time? No. As the size of your data structure grows, it requires progressively slower memory. If you have a small hash table, it can reside in the CPU cache and be fast. Once it reaches megabytes in size, the data structure tends to live in RAM, which is much slower. And then, eventually, you have to store it on disk, which is even slower. And so forth.</p>
<p>To put it differently, saying that a hash table is O(1) or constant time is a model. It can be true, maybe even often, but it is not reality. Models are great teaching tools: they present a simplified model that you can quickly learn. But models can also introduce biases in how we think.</p>
<p>For example, even though you have read my paragraph that says that the dict data structure gets slower, you may not believe it. You may also believe that it is typically going to be the fastest approach you can use.</p>
<p>Let us consider another practical case. Suppose that you have a large map from strings to integers, that you build once and then only query. That is a common situation: a dictionary of words to identifiers, a lookup table of country codes, a table of feature names.</p>
<p>The <a href="https://pypi.org/project/fastconstmap/">fastconstmap</a> library builds an immutable map from a <code>dict[str, int]</code>. It is suitable when your keys are known in advance.</p>
<p>I build a map from a million random sixteen-character strings to integers, and then look up every key in a shuffled order. With a <code>dict</code>, I write the obvious loop:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">total</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #0000cf; font-weight: bold;">0</span>
<span style="color: #204a87; font-weight: bold;">for</span> <span style="color: #000;">k</span> <span style="color: #204a87; font-weight: bold;">in</span> <span style="color: #000;">probes</span><span style="color: #000; font-weight: bold;">:</span>
    <span style="color: #000;">total</span> <span style="color: #ce5c00; font-weight: bold;">+=</span> <span style="color: #000;">d</span><span style="color: #000; font-weight: bold;">[</span><span style="color: #000;">k</span><span style="color: #000; font-weight: bold;">]</span>
</code></pre>
</div>
<p>With fastconstmap, I ask for all the keys at once, writing the values into a buffer that I own, so that no Python object is allocated per key:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">out</span> <span style="color: #ce5c00; font-weight: bold;">=</span> <span style="color: #000;">array</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #4e9a06;">"Q"</span><span style="color: #000; font-weight: bold;">,</span> <span style="color: #204a87;">bytes</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #0000cf; font-weight: bold;">8</span> <span style="color: #ce5c00; font-weight: bold;">*</span> <span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">))</span>
<span style="color: #000;">cm</span><span style="color: #ce5c00; font-weight: bold;">.</span><span style="color: #000;">get_many_into</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">probes</span><span style="color: #000; font-weight: bold;">,</span> <span style="color: #000;">out</span><span style="color: #000; font-weight: bold;">)</span>
</code></pre>
</div>
<p>I am being generous to the <code>dict</code>. I reuse the same string objects for the lookups, and a Python string caches its hash value the first time it is computed. So the <code>dict</code> does not pay for hashing at all, while fastconstmap hashes every key every time. Here are the results, in nanoseconds per key.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;">n</th>
<th style="text-align: right;">dict</th>
<th style="text-align: right;">get_many_into</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">1000</td>
<td style="text-align: right;">21.8</td>
<td style="text-align: right;">4.3</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">10000</td>
<td style="text-align: right;">31.9</td>
<td style="text-align: right;">4.8</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">100000</td>
<td style="text-align: right;">48.1</td>
<td style="text-align: right;">5.2</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">1000000</td>
<td style="text-align: right;">201.9</td>
<td style="text-align: right;">11.8</td>
</tr>
</tbody>
</table>
<p>The <code>dict</code> is not constant time. It goes from 22 ns to 202 ns per key as the map grows, a factor of nine, and it is not because the algorithm changed or because of collisions. It is because a million keys, their string objects, and their integer objects occupy about 116 bytes per key, so the lookups miss in the cache. The fastconstmap version needs 9 bytes per key: it stays in the cache much longer. Pay attention to how the numbers scale: the dict becomes 10 times slower as the size grows.</p>
<p>The lesson is always the same. Some models are useful but none of them is reality. Be mindful of cognitive biases.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/upcoming">The code is available.</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/09/03/python-sets-and-dictionaries-can-have-quadratic-time-performance/feed/</wfw:commentRss>
			<slash:comments>8</slash:comments>
		
		
			</item>
		<item>
		<title>The new Go JSON API: twice as fast, or 1.5x slower?</title>
		<link>https://lemire.me/blog/2026/08/29/the-new-go-json-api-twice-as-fast-or-1-5x-slower/</link>
					<comments>https://lemire.me/blog/2026/08/29/the-new-go-json-api-twice-as-fast-or-1-5x-slower/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 29 Aug 2026 18:33:33 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22792</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/xZAWW-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />JSON is a standard format for data interchange. It is effectively a tiny subset of JavaScript made of objects and arrays. It looks as follows {"key":1, "text":[1.0,2.0]}. Many programming languages include a JSON library in their standard libraries: C#, Go, Java (soon), Python, JavaScript, etc. The Go implementation is convenient, but not especially fast. Go &#8230; <a href="https://lemire.me/blog/2026/08/29/the-new-go-json-api-twice-as-fast-or-1-5x-slower/" class="more-link">Continue reading <span class="screen-reader-text">The new Go JSON API: twice as fast, or 1.5x slower?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/xZAWW-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>JSON is a standard format for data interchange. It is effectively a tiny subset of JavaScript made of objects and arrays. It looks as follows <code>{"key":1, "text":[1.0,2.0]}</code>.</p>
<p>Many programming languages include a JSON library in their standard libraries: C#, Go, Java (soon), Python, JavaScript, etc. The Go implementation is convenient, but not especially fast.</p>
<p>Go 1.27 makes a new JSON package (<code>encoding/json/v2</code>) available by default in its standard library. The two APIs look almost the same:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">import</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">json</span><span style="color: #f8f8f8;">    </span><span style="color: #4e9a06;">"encoding/json"</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">jsonv2</span><span style="color: #f8f8f8;">  </span><span style="color: #4e9a06;">"encoding/json/v2"</span>
<span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">err</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">json</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Marshal</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">v</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">err</span><span style="color: #f8f8f8;">  </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">jsonv2</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Marshal</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">v</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000;">err</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">json</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Unmarshal</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">v</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000;">err</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">jsonv2</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Unmarshal</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">v</span><span style="color: #000; font-weight: bold;">)</span>
</code></pre>
</div>
<p>The two are not directly comparable as they differ with respect to Unicode validation, case sensitivity, etc. So it is not a drop-in replacement.</p>
<p>However, the legacy API (<code>encoding/json</code>) has also been reimplemented on top of the new engine. You can use the legacy API with either the new engine or the old one (<code>GOEXPERIMENT=nojsonv2</code>) through a flag. So we have three possibilities.</p>
<ol>
<li><em>json (legacy)</em> — <code>encoding/json</code> built with <code>GOEXPERIMENT=nojsonv2</code>, the original implementation</li>
<li><em>json (Go 1.27)</em> — <code>encoding/json</code> as of 1.27, v1 API on the v2 backend</li>
<li><em>json/v2</em> — <code>encoding/json/v2</code></li>
</ol>
<p>I used the usual <a href="https://simdjson.org/">simdjson documents</a>: <code>twitter.json</code> (632 kB, nested objects with short string keys), <code>canada.json</code> (2.25 MB, one large array of coordinates), and <code>citm_catalog.json</code> (1.73 MB, nested objects with numeric keys). I parse them into <code>any</code> (<code>interface{}</code>), which is the general-purpose path.</p>
<p>I ran this on an Apple M4 Max and on an Intel Xeon Gold 6548N (Emerald Rapids) using Go 1.27.0, on a single core (<code>GOMAXPROCS=1</code>), reporting the median of eight runs.</p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2-1024x790.png" alt="" width="660" height="509" class="alignnone size-large wp-image-22795" srcset="https://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2-1024x790.png 1024w, https://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2-300x232.png 300w, https://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2-768x593.png 768w, https://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2-1536x1186.png 1536w, https://lemire.me/blog/wp-content/uploads/2026/08/json-v1-v2.png 1938w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>When unmarshalling, the legacy API on the new backend is faster than the original on <code>twitter.json</code> (172 MB/s to 203 MB/s) and on <code>citm_catalog.json</code> (186 MB/s to 241 MB/s), but slower on <code>canada.json</code> (128 MB/s down to 106 MB/s). When marshalling, it is up to twice as fast: 198 MB/s to 374 MB/s on <code>twitter.json</code>. So merely upgrading to Go 1.27, without changing a line of code, should make marshalling faster.</p>
<p>Switching to the new API helps more. Compared to the original implementation, <code>encoding/json/v2</code> unmarshals 1.5x to 2.3x faster and marshals 1.2x to 3x faster. Compared to the Go 1.27 legacy API, unmarshalling gains another 1.8x to 2x, while marshalling gains much less (1.0x to 1.7x): part of the remaining difference is that <code>json/v2</code> does less work during marshalling.</p>
<p>Thus far, I was unmarshalling into <code>any</code>, meaning that I assumed that I did not know the structure of the document. I also round-trip a slice of 10,000 small structs:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">type</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Record</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">struct</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">ID</span><span style="color: #f8f8f8;">     </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;">      </span><span style="color: #4e9a06;">`json:"id"`</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">Name</span><span style="color: #f8f8f8;">   </span><span style="color: #204a87; font-weight: bold;">string</span><span style="color: #f8f8f8;">   </span><span style="color: #4e9a06;">`json:"name"`</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">Email</span><span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">string</span><span style="color: #f8f8f8;">   </span><span style="color: #4e9a06;">`json:"email"`</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">Active</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">bool</span><span style="color: #f8f8f8;">     </span><span style="color: #4e9a06;">`json:"active"`</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">Score</span><span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">float64</span><span style="color: #f8f8f8;">  </span><span style="color: #4e9a06;">`json:"score"`</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">Tags</span><span style="color: #f8f8f8;">   </span><span style="color: #000; font-weight: bold;">[]</span><span style="color: #204a87; font-weight: bold;">string</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">`json:"tags"`</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>The schema is specified: the JSON must be <code>[{"id":..., "name":...}, {"id":..., "name":...}...]</code>. I still get faster unmarshalling with the new API, but the legacy API with the legacy engine is faster when marshalling.<a href="http://lemire.me/blog/wp-content/uploads/2026/08/json-records.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/08/json-records-1024x419.png" alt="" width="660" height="270" class="alignnone size-large wp-image-22796" srcset="https://lemire.me/blog/wp-content/uploads/2026/08/json-records-1024x419.png 1024w, https://lemire.me/blog/wp-content/uploads/2026/08/json-records-300x123.png 300w, https://lemire.me/blog/wp-content/uploads/2026/08/json-records-768x314.png 768w, https://lemire.me/blog/wp-content/uploads/2026/08/json-records-1536x628.png 1536w, https://lemire.me/blog/wp-content/uploads/2026/08/json-records.png 1870w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>The original implementation is faster. The Go 1.27 release notes said that marshal performance is broadly at parity with the previous implementation. For my test, it is not the case.</p>
<p>So unmarshalling gets faster across the board with <code>encoding/json/v2</code>, and marshalling gets faster for <code>any</code>, but it is about 1.5x slower for typed structs in my tests.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/29">The code is available.</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/29/the-new-go-json-api-twice-as-fast-or-1-5x-slower/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Java&#8217;s String.indexOf can be slow (quadratic)</title>
		<link>https://lemire.me/blog/2026/08/22/javas-string-indexof-can-be-slow-quadratic/</link>
					<comments>https://lemire.me/blog/2026/08/22/javas-string-indexof-can-be-slow-quadratic/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 22 Aug 2026 14:56:16 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22778</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/lYgqA-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />In Java, you find the location of a substring using indexOf. String haystack = "The quick brown fox jumps over the lazy dog"; String needle = "fox"; int index = haystack.indexOf(needle); Naively, you might implement indexOf by a loop inside a loop, like so. int naiveIndexOf(String haystack, String needle) { for (int i = 0; &#8230; <a href="https://lemire.me/blog/2026/08/22/javas-string-indexof-can-be-slow-quadratic/" class="more-link">Continue reading <span class="screen-reader-text">Java&#8217;s String.indexOf can be slow (quadratic)</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/lYgqA-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>In Java, you find the location of a substring using <code>indexOf</code>.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"The quick brown fox jumps over the lazy dog"</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"fox"</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">index</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">indexOf</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>Naively, you might implement <code>indexOf</code> by a loop inside a loop, like so.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">naiveIndexOf</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">length</span><span style="color: #000; font-weight: bold;">()</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">length</span><span style="color: #000; font-weight: bold;">();</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">j</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">j</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">length</span><span style="color: #000; font-weight: bold;">()</span><span style="color: #f8f8f8;"> </span>
<span style="color: #f8f8f8;">          </span><span style="color: #ce5c00; font-weight: bold;">&amp;&amp;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">charAt</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">i</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">j</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">charAt</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">j</span><span style="color: #000; font-weight: bold;">);</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">j</span><span style="color: #ce5c00; font-weight: bold;">++</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{}</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">j</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">==</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">length</span><span style="color: #000; font-weight: bold;">())</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">i</span><span style="color: #000; font-weight: bold;">;</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-</span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>The Java implementation is much more sophisticated, and it is highly accelerated.</p>
<p>However, there are pathological cases where the Java implementation can be slow. What do I mean? Well, you do expect that the search will be more and more expensive as the size of the string grows. Right? So if you search through a 1 kilobyte string and then search through a 10 kilobyte search, you would not be surprised if the latter takes ten times slower.</p>
<p>But what of the substring? If you search for short substrings (<code>fox</code> in my example), the everything is fine. But what if you search longer and longer substrings (<code>fox jumps</code> or <code>fox jumps over</code>)? If it gets more expensive when both the string and the substring get longer, then you have what we call a quadratic complexity. In other words, it is slow.</p>
<p>In Java, if <em>n</em> is the length of your string and <em>m</em> is the length of the substring, then the complexity of <code>indexOf</code> is O(<em>n</em>·<em>m</em>). And if you look at my naive implementation (<code>naiveIndexOf</code>) then you see that in the worst case, it might do up to close to <code>haystack.length() * needle.length()</code> comparisons, that is, it is O(<em>n</em>·<em>m</em>).</p>
<p>The exact implementation of the <code>indexOf</code> function depends on your CPU and Java version. I am using OpenJDK 25 on Apple Silicon (ARM). For my purposes, I will use as a haystack of <em>n</em> copies of <code>a</code> and for the needle, the same thing, but ending with a different letter.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #8f5902; font-style: italic;">// n &gt; m</span>
<span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">haystack</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"a"</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">repeat</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">n</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #000;">String</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">needle</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"a"</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #c4a000;">repeat</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">m</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">-</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">+</span><span style="color: #f8f8f8;"> </span><span style="color: #4e9a06;">"b"</span><span style="color: #000; font-weight: bold;">;</span>
</code></pre>
</div>
<p>I measured OpenJDK 25 on an Apple M4 Max. The haystack is one megabyte. Numbers are nanoseconds per haystack character.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;"><em>m</em></th>
<th style="text-align: right;"><code>indexOf</code></th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">512</td>
<td style="text-align: right;">140</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">1024</td>
<td style="text-align: right;">273</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">2048</td>
<td style="text-align: right;">543</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">4096</td>
<td style="text-align: right;">1076</td>
</tr>
</tbody>
</table>
<p>At <em>m</em> = 4096, a single <code>indexOf</code> over one megabyte takes 1.1 seconds.</p>
<p>Can you do better against such adversarial inputs? The textbook solution is the Two-Way algorithm of Crochemore and Perrin (1991). The implementation is simple and your favourite AI can code it for you in any programming language.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;"><em>m</em></th>
<th style="text-align: right;"><code>indexOf</code></th>
<th style="text-align: right;">Two-Way</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">8</td>
<td style="text-align: right;">0.44</td>
<td style="text-align: right;">0.29</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">32</td>
<td style="text-align: right;">0.48</td>
<td style="text-align: right;">0.29</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">128</td>
<td style="text-align: right;">0.45</td>
<td style="text-align: right;">0.30</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">256</td>
<td style="text-align: right;">73.9</td>
<td style="text-align: right;">0.32</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">1024</td>
<td style="text-align: right;">273</td>
<td style="text-align: right;">0.32</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">4096</td>
<td style="text-align: right;">1076</td>
<td style="text-align: right;">0.31</td>
</tr>
</tbody>
</table>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic-1024x675.png" alt="" width="660" height="435" class="alignnone size-large wp-image-22780" srcset="https://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic-1024x675.png 1024w, https://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic-300x198.png 300w, https://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic-768x506.png 768w, https://lemire.me/blog/wp-content/uploads/2026/08/indexof-quadratic.png 1393w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>Two-Way stays at about 0.3 ns/character no matter how long the needle is. At <em>m</em> = 4096 it is about 3500 times faster than <code>indexOf</code> on the first-character adversary.</p>
<p>So, should you switch to Two-Way for everything? No. On random text, the <code>indexOf</code> function is much faster than Two-Way.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;"><em>m</em></th>
<th style="text-align: right;"><code>indexOf</code></th>
<th style="text-align: right;">Two-Way</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">8</td>
<td style="text-align: right;">0.30</td>
<td style="text-align: right;">0.55</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">64</td>
<td style="text-align: right;">0.10</td>
<td style="text-align: right;">0.56</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">256</td>
<td style="text-align: right;">0.24</td>
<td style="text-align: right;">0.53</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">4096</td>
<td style="text-align: right;">0.22</td>
<td style="text-align: right;">0.55</td>
</tr>
</tbody>
</table>
<p>And Two-Way has to do non-trivial work before the search begins. So it has additional fixed overhead. It would lose most of the time in the real world, sometimes by a wide margin.</p>
<p>Should you worry about this? No. The <code>indexOf</code> function in Java is fine.</p>
<p>If an adversary can control the needle (substring), then make sure to reject long needles. Most of the time, we search for short sequences (say, less than 80 characters). If you are worried about your system crashing, you will put bounds on inputs in any case.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/22">The Java source is available</a>.</p>
<p><em>Further reading</em>: Crochemore, M., &amp; Perrin, D. (1991). <a href="http://monge.univ-mlv.fr/~mac/Articles-PDF/CP-1991-jacm.pdf">Two-way string-matching</a>. <em>Journal of the ACM</em>, 38(3), 650–674.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/22/javas-string-indexof-can-be-slow-quadratic/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Parsing IP addresses in C# at crazy speeds</title>
		<link>https://lemire.me/blog/2026/08/19/parsing-ip-addresses-in-c-at-crazy-speeds/</link>
					<comments>https://lemire.me/blog/2026/08/19/parsing-ip-addresses-in-c-at-crazy-speeds/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 19:07:48 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22772</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/Cgic2-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />We are all familiar with IP addresses such as 192.168.0.1. They are typically written as four numbers in the range 0 to 255 inclusive, separated by dots. In C#, you can parse them with the standard library using IPAddress.TryParse. Pedantic people are quick to point out that IP addresses can take different forms: they can &#8230; <a href="https://lemire.me/blog/2026/08/19/parsing-ip-addresses-in-c-at-crazy-speeds/" class="more-link">Continue reading <span class="screen-reader-text">Parsing IP addresses in C# at crazy speeds</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/Cgic2-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>We are all familiar with IP addresses such as <code>192.168.0.1</code>. They are typically written as four numbers in the range 0 to 255 inclusive, separated by dots. In C#, you can parse them with the standard library using <code>IPAddress.TryParse</code>.</p>
<p>Pedantic people are quick to point out that IP addresses can take different forms: they can be IPv6 or IPv4 and there are many weird ways to write an IPv4 address. But for the purpose of performance optimization, we care about the common case. The common case is strings such as <code>192.168.0.1</code> or <code>12.121.244.111</code>.</p>
<p>Our processors are capable of data parallelism, meaning that they have instructions (called SIMD) that can process several bytes at once, at least 16 bytes, sometimes more. <a href="https://lemire.me/blog/2023/06/08/parsing-ip-addresses-crazily-fast/">A few years ago, I showed that you can parse IPv4 addresses with SIMD</a>. I have been <a href="https://github.com/lemire/simdip">revisiting this idea with AVX-512</a>, the instruction set that recent x64 (AMD/Intel) processors support. I expect that all Intel and AMD processors made in the near future will have great support for AVX-512, and it is already the case for server processors and recent AMD processors.</p>
<p>So I wondered, could we do it in C#? People are sometimes surprised that I care about C#. Isn&#8217;t that more Microsoft slop? No. Not at all. C# and .NET are very reasonable, portable systems.</p>
<p>Plus you can write fast code in C#. I have two optimized libraries that I hope the Microsoft .NET team will one day adopt in the standard .NET library: an optimized <code>Utf8Utility.GetPointerToFirstInvalidByte</code> function used internally to validate Unicode strings (in the <a href="https://github.com/simdutf/SimdUnicode">SimdUnicode library</a>) and a <a href="https://github.com/simdutf/SimdBase64">fast base64 decoding library</a>. I love working with .NET C#.</p>
<p>As of .NET 10, we have AVX-512 support, including masked loads. What are masked loads and why do they matter? Suppose that I give you a string that is no longer than 16 bytes, but could be shorter. If you load data in a SIMD register, you normally have to load the full register width (so 8, 16, 32, 64 bytes). So what do you do when it is not possible? You can pad the input string or pull other tricks, but it gets dirty. A nice approach is to have masked loads where you, say, load the full register (say 16 bytes), but you indicate which bytes you want to be loaded from memory with a mask. So if you use <code>0b10011</code> as a mask, then only the first, second, and fifth bytes are loaded from memory. This makes it possible to initialize a 16-byte register with a string that has between 0 and 16 bytes, while never reading beyond the string. I have an article entitled <a href="https://lemire.me/blog/2022/11/08/modern-vector-programming-with-masked-loads-and-stores/">Modern vector programming with masked loads and stores</a> if you want to know more.</p>
<p>To make things trickier, C#, like Java and JavaScript, defaults to UTF-16, meaning that each character, even if it is an ASCII character like <code>A</code> or <code>1</code>, uses two bytes. The ASCII codepoint value occupies the least significant bits of a 16-bit word.</p>
<p>So what we need to do is to selectively load from a 32-byte input, and then drop the unnecessary zero bytes. The gist of it looks as follows in C#.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">unsafe</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">bool</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">TryParseAvx512</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">ReadOnlySpan</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">out</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">uint</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">ip</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">len</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Length</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">        </span><span style="color: #204a87; font-weight: bold;">fixed</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #204a87; font-weight: bold;">char</span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">cp</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">            </span><span style="color: #8f5902; font-style: italic;">// next two lines are a trick to load just the first len characters</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000;">Vector256</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">charMask</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Vector256</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">LessThan</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">CharLaneIndex</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Vector256</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Create</span><span style="color: #000; font-weight: bold;">((</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #000;">len</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000;">Vector256</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">chars</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Avx512BW</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">VL</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">MaskLoad</span><span style="color: #000; font-weight: bold;">((</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #000;">cp</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">charMask</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Vector256</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Create</span><span style="color: #000; font-weight: bold;">((</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #4e9a06;">'0'</span><span style="color: #000; font-weight: bold;">));</span>
<span style="color: #f8f8f8;">            </span><span style="color: #8f5902; font-style: italic;">// check that everything is ASCII otherwise, it is not an IP!</span>
<span style="color: #f8f8f8;">            </span><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">Avx512BW</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">VL</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">CompareGreaterThan</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">chars</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Vector256</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Create</span><span style="color: #000; font-weight: bold;">((</span><span style="color: #204a87; font-weight: bold;">ushort</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #0000cf; font-weight: bold;">0x7F</span><span style="color: #000; font-weight: bold;">)).</span><span style="color: #000;">ExtractMostSignificantBits</span><span style="color: #000; font-weight: bold;">()</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">!=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">0</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">                </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">false</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #f8f8f8;">            </span><span style="color: #8f5902; font-style: italic;">// There we go, we have the address as ASCII</span>
<span style="color: #f8f8f8;">            </span><span style="color: #8f5902; font-style: italic;">// in a 16-byte register.</span>
<span style="color: #f8f8f8;">            </span><span style="color: #000;">Vector128</span><span style="color: #ce5c00; font-weight: bold;">&lt;</span><span style="color: #204a87; font-weight: bold;">byte</span><span style="color: #ce5c00; font-weight: bold;">&gt;</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">str</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Avx512BW</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">VL</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">ConvertToVector128Byte</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">chars</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">            </span><span style="color: #8f5902; font-style: italic;">// ...</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>This looks a bit difficult to read, but that&#8217;s fine. Most people never need to worry about such code.</p>
<p>Then we use a somewhat fancy trick where we locate the dots, and use the fact that there are only 81 ways to position the dots. We then move the bytes, do a dot product and validate. It is the same routine as the C++ code. It is not trivial, but I am working on a formal paper to document the tricks used.</p>
<p>The pedantic people will say: wait, there are other ways to write IP addresses !!! Ok fine. We handle them with a fallback, like so.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">TryParseAvx512</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">out</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">uint</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">ip</span><span style="color: #000; font-weight: bold;">))</span>
<span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">address</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">new</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">IPAddress</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">ip</span><span style="color: #000; font-weight: bold;">);</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">true</span><span style="color: #000; font-weight: bold;">;</span>
<span style="color: #000; font-weight: bold;">}</span>
<span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">IPAddress</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">TryParse</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">s</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">out</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">address</span><span style="color: #000; font-weight: bold;">);</span>
</code></pre>
</div>
<p>What about the cases where your processor does not support AVX-512? C# makes this dead easy. You can just guard it with one if:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">if</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">Avx512BW</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">VL</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">IsSupported</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">...</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>To benchmark this, I generated 10,000 random 32-bit addresses and parsed the resulting strings 20 million times, constructing an <code>IPAddress</code> each time. On a relatively recent Intel processor (Intel Xeon Gold 6548N, Emerald Rapids) running .NET 10, I get the following.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">function</th>
<th style="text-align: right;">ns/addr</th>
<th style="text-align: right;">million addr/s</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>IPAddress.TryParse</code></td>
<td style="text-align: right;">45.3</td>
<td style="text-align: right;">22.1</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">AVX-512 + fallback</td>
<td style="text-align: right;">14.1</td>
<td style="text-align: right;">71.1</td>
</tr>
</tbody>
</table>
<p>So the AVX-512 approach is about three times faster than the standard library. My routine itself does not take fourteen nanoseconds; there is other overhead.</p>
<p>As usual, <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/19">the C# source is available</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/19/parsing-ip-addresses-in-c-at-crazy-speeds/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Go 1.27 will make some allocations cheaper</title>
		<link>https://lemire.me/blog/2026/08/15/go-1-27-will-make-some-allocations-cheaper/</link>
					<comments>https://lemire.me/blog/2026/08/15/go-1-27-will-make-some-allocations-cheaper/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 15 Aug 2026 20:59:19 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22769</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/2N3RJ-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Like most programming languages, Go has both stack allocations, whose lifetime is limited to the current function, and dynamic (or heap) allocations. The name stack comes from the fact that the memory management is somewhat trivial. There is typically one stack per thread (or goroutine in Go). When a function needs memory, it simply appends &#8230; <a href="https://lemire.me/blog/2026/08/15/go-1-27-will-make-some-allocations-cheaper/" class="more-link">Continue reading <span class="screen-reader-text">Go 1.27 will make some allocations cheaper</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/2N3RJ-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Like most programming languages, Go has both stack allocations, whose lifetime is limited to the current function, and dynamic (or heap) allocations.</p>
<p>The name stack comes from the fact that the memory management is somewhat trivial. There is typically one stack per thread (or goroutine in Go). When a function needs memory, it simply appends data to the stack. When the function returns, the memory is dropped from the end of the stack. So the memory last allocated is deallocated first.</p>
<p>Heap memory is potentially considerably more complex. For one thing, it is meant to be accessible by several threads (or goroutines). An object can be allocated by one function and later reclaimed after an entirely different function, possibly running on a different thread (or goroutine), has dropped the last reference to it. Unlike the stack, there is no prescribed order for allocating and reclaiming heap memory. In Go, the garbage collector does the reclaiming.</p>
<p>Typically, stack allocations have a size known at compile time. Many systems give each thread a fixed-size stack, although Go grows goroutine stacks as needed.</p>
<p>There are many ways in Go to do a heap allocation. A common one is when you allocate a slice, as in this instance where you allocate memory for 100 integers:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">x</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87;">make</span><span style="color: #000; font-weight: bold;">([]</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">100</span><span style="color: #000; font-weight: bold;">)</span>
</code></pre>
</div>
<p>If the slice <code>x</code> is not entirely local to a function, Go will typically just allocate it on the heap. It will do so similarly when a function returns a pointer. For example, in the following instance, I assign the value 1 to a local integer variable, but I return a pointer to it.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">func</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">f</span><span style="color: #000; font-weight: bold;">()</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #204a87; font-weight: bold;">int</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">  </span><span style="color: #000;">x</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #0000cf; font-weight: bold;">1</span>
<span style="color: #f8f8f8;">  </span><span style="color: #204a87; font-weight: bold;">return</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">x</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>In C/C++, this would be quite bad. You should get a warning such as <code>address of local variable 'x' returned</code>. In Go, the variable <code>x</code> will typically get allocated on the heap.</p>
<p>In many Go programs, we end up doing a lot of heap allocations of small objects. It can become a bottleneck in some cases. Think about when you are maintaining a tree or a linked list where each value (node) is an object that must live on the heap. If the data structure is highly dynamic, you will be constantly allocating these small objects.</p>
<p>Memory allocation on the heap is usually not done at arbitrary sizes. You often cannot get exactly, say, 13 bytes. In Go, small allocations are rounded up to a size class: 8 bytes, 16 bytes, 24 bytes, 32 bytes, and so forth. There is also some overhead to each heap allocation, from rounding and from allocator metadata.</p>
<p>The compiler knows the size of the object, but prior to Go 1.27, Go would call a generic function when doing a heap allocation. This generic function would then look up the size class and take the corresponding path. Starting with 1.27, for small objects (under 80 bytes), Go relies on dedicated functions.</p>
<p>It is easy to benchmark in Go. A basic benchmark might look as follows.</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #204a87; font-weight: bold;">type</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">Node</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">struct</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">value</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">int64</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000;">next</span><span style="color: #f8f8f8;">  </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">Node</span>
<span style="color: #000; font-weight: bold;">}</span>
<span style="color: #204a87; font-weight: bold;">var</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">sink</span><span style="color: #f8f8f8;"> </span><span style="color: #204a87; font-weight: bold;">any</span>
<span style="color: #204a87; font-weight: bold;">func</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">BenchmarkAllocNode16</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">b</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">*</span><span style="color: #000;">testing</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">B</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">    </span><span style="color: #204a87; font-weight: bold;">for</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">b</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Loop</span><span style="color: #000; font-weight: bold;">()</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">{</span>
<span style="color: #f8f8f8;">        </span><span style="color: #000;">sink</span><span style="color: #f8f8f8;"> </span><span style="color: #000; font-weight: bold;">=</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">&amp;</span><span style="color: #000;">Node</span><span style="color: #000; font-weight: bold;">{}</span>
<span style="color: #f8f8f8;">    </span><span style="color: #000; font-weight: bold;">}</span>
<span style="color: #000; font-weight: bold;">}</span>
</code></pre>
</div>
<p>On my MacBook, the results are quite telling. Go 1.27 is nearly twice as fast!</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">allocation</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Go 1.26</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Go 1.27</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">speedup</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">16 B, has pointer</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">9.5 ns</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">5.5 ns</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">1.8x</td>
</tr>
</tbody>
</table>
<p>This will not help all software, just the components that do many small allocations.</p>
<p><em>The code is <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/15">available</a>.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/15/go-1-27-will-make-some-allocations-cheaper/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>AI programming : are you angry yet?</title>
		<link>https://lemire.me/blog/2026/08/12/ai-programming-are-you-angry-yet/</link>
					<comments>https://lemire.me/blog/2026/08/12/ai-programming-are-you-angry-yet/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 12 Aug 2026 15:50:48 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22764</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/o1lot-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />AI-assisted programming is fast evolving and there is a tension between &#8216;we no longer need to understand the code&#8217; and &#8216;what is my purpose as a programmer&#8217;. I recorded a short video on this topic with how I think the tension can result in conflicts.]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/o1lot-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>AI-assisted programming is fast evolving and there is a tension between &#8216;we no longer need to understand the code&#8217; and &#8216;what is my purpose as a programmer&#8217;. I recorded a short video on this topic with how I think the tension can result in conflicts.</p>
<p><iframe loading="lazy" width="560" height="315" src="https://www.youtube.com/embed/WweTpPgJPhg?si=_wB-VqPyAmufrElY" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"></iframe></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/12/ai-programming-are-you-angry-yet/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Profile-guided optimization in Go</title>
		<link>https://lemire.me/blog/2026/08/09/profile-guided-optimization-in-go/</link>
					<comments>https://lemire.me/blog/2026/08/09/profile-guided-optimization-in-go/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 09 Aug 2026 23:17:36 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22758</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/ltISq-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />When a compiler optimizes your program, it has to guess. Which functions are worth inlining? Which side of a branch is the common one? Which method does this interface call actually reach? At compile time it cannot know, so it uses heuristics. Profile-guided optimization (PGO) replaces the guessing with measurement: you run your program, record &#8230; <a href="https://lemire.me/blog/2026/08/09/profile-guided-optimization-in-go/" class="more-link">Continue reading <span class="screen-reader-text">Profile-guided optimization in Go</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/ltISq-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>When a compiler optimizes your program, it has to guess. Which functions are worth inlining? Which side of a branch is the common one? Which method does this interface call actually reach? At compile time it cannot know, so it uses heuristics. Profile-guided optimization (PGO) replaces the guessing with measurement: you run your program, record where it spends its time, and hand that recording back to the compiler for a second build.</p>
<p>PGO is a common feature of compiler systems. Google applied PGO to Chrome under Windows in 2016, <a href="https://blog.chromium.org/2016/10/making-chrome-on-windows-faster-with-pgo.html">reporting gains of up to 15%</a>. I expect <a href="https://webkit.org/blog/15249/">all mainstream Web browsers to be built with PGO</a>. </p>
<p>There are now fancier techniques than mere heuristics with PGO. You can use AI to recognize patterns and so forth. But they are not always widely available.</p>
<p>Go has supported PGO since version 1.20. You collect a profile, and pass it to the compiler.</p>
<p>A CPU profile is a statistical record of where a program spends its time. While the program runs, the Go runtime interrupts it about a hundred times a second and writes down the call stack at that instant. After a few seconds you have thousands of such samples, and counting them tells you which functions were executing and who called them. In Go you produce one by wrapping the work you care about:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code><span style="color: #000;">f</span><span style="color: #000; font-weight: bold;">,</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">_</span><span style="color: #f8f8f8;"> </span><span style="color: #ce5c00; font-weight: bold;">:=</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">os</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">Create</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #4e9a06;">"cpu.pprof"</span><span style="color: #000; font-weight: bold;">)</span>
<span style="color: #000;">pprof</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">StartCPUProfile</span><span style="color: #000; font-weight: bold;">(</span><span style="color: #000;">f</span><span style="color: #000; font-weight: bold;">)</span><span style="color: #f8f8f8;">   </span><span style="color: #8f5902; font-style: italic;">// from runtime/pprof</span>
<span style="color: #204a87; font-weight: bold;">defer</span><span style="color: #f8f8f8;"> </span><span style="color: #000;">pprof</span><span style="color: #000; font-weight: bold;">.</span><span style="color: #000;">StopCPUProfile</span><span style="color: #000; font-weight: bold;">()</span>
</code></pre>
</div>
<p>The compiler reads the call-stack counts and uses them for two things above all: inlining call sites that turn out to be hot, and devirtualizing interface calls whose target is nearly always the same concrete type.</p>
<p>I took three JSON documents that I wanted to parse:</p>
<ul>
<li><code>twitter.json</code> (632 kB), a nest of small objects with short string keys</li>
<li><code>canada.json</code> (2.25 MB), essentially one enormous array of floating-point coordinates</li>
<li><code>citm_catalog.json</code> (1.73 MB), deeply nested objects with numeric keys</li>
</ul>
<p>I parse each of them with the standard library&#8217;s <code>encoding/json</code> into an <code>interface{}</code>. The baseline, with no profile, parses at 112 MB/s for <code>twitter.json</code>, 74 MB/s for <code>canada.json</code> and 116 MB/s for <code>citm_catalog.json</code>.</p>
<p>The procedure is three commands:</p>
<div class="highlight" style="background: #f8f8f8;">
<pre style="line-height: 125%;"><span></span><code>go<span style="color: #f8f8f8;"> </span>build<span style="color: #f8f8f8;"> </span>-o<span style="color: #f8f8f8;"> </span>bench<span style="color: #f8f8f8;"> </span>.<span style="color: #f8f8f8;">                        </span><span style="color: #8f5902; font-style: italic;"># ordinary build</span>
./bench<span style="color: #f8f8f8;"> </span>-profile<span style="color: #f8f8f8;"> </span>cpu.pprof<span style="color: #f8f8f8;"> </span>-train<span style="color: #f8f8f8;"> </span>twitter.json<span style="color: #f8f8f8;">   </span><span style="color: #8f5902; font-style: italic;"># collect a CPU profile</span>
go<span style="color: #f8f8f8;"> </span>build<span style="color: #f8f8f8;"> </span>-pgo<span style="color: #ce5c00; font-weight: bold;">=</span>cpu.pprof<span style="color: #f8f8f8;"> </span>-o<span style="color: #f8f8f8;"> </span>bench_pgo<span style="color: #f8f8f8;"> </span>.<span style="color: #f8f8f8;">     </span><span style="color: #8f5902; font-style: italic;"># build again, with the profile</span>
</code></pre>
</div>
<p>I did it three times, profiling each document on its own, and then measured all three documents against each of the three builds.</p>
<p><span>Each panel of the figure is one document being parsed, and the three bars inside it are the three PGO builds: the binary trained on twitter.json, the one trained on canada.json, and the one trained on citm_catalog.json. Bar height is the speed gain over the ordinary, profile-free build of that same document, in percent, so zero means PGO changed nothing and a bar below the axis means the PGO build was slower. The green bar in each panel is the matched case, where the profile was collected on the very document being measured.</span></p>
<p><a href="http://lemire.me/blog/wp-content/uploads/2026/08/go-pgo.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/08/go-pgo-707x1024.png" alt="" width="660" height="956" class="alignnone size-large wp-image-22759" srcset="https://lemire.me/blog/wp-content/uploads/2026/08/go-pgo-707x1024.png 707w, https://lemire.me/blog/wp-content/uploads/2026/08/go-pgo-207x300.png 207w, https://lemire.me/blog/wp-content/uploads/2026/08/go-pgo-768x1112.png 768w, https://lemire.me/blog/wp-content/uploads/2026/08/go-pgo-1061x1536.png 1061w, https://lemire.me/blog/wp-content/uploads/2026/08/go-pgo.png 1292w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>The gains are modest. The best result is <code>canada.json</code> at +4.7%, and most differences are in the 2–3% range. Profiling one document usually helps the others, but not reliably. Profiling <code>twitter.json</code> gave a decent improvement everywhere: +3.1%, +2.0%, +2.8%. But profiling <code>canada.json</code> bought 4.7% on <code>canada.json</code> and essentially nothing anywhere else. Interestingly, profiling <code>citm_catalog.json</code> produced a mere +0.8% on its own document while helping <code>twitter.json</code> more.</p>
<p>A 3% speedup is not exciting in isolation, but it may come nearly for free. Observe how you may get slightly negative results for cases you did not train for. That&#8217;s expected generally, but the effect is modest in the case of Go because its optimizations are themselves modest in the first pace. That is, you are not getting a much an effect, but the process is less likely to backfire for other workloads.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/09">The code is available.</a><br />
<a href="http://lemire.me/blog/wp-content/uploads/2026/08/go-pgo.png"></a><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/09"></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/09/profile-guided-optimization-in-go/feed/</wfw:commentRss>
			<slash:comments>5</slash:comments>
		
		
			</item>
		<item>
		<title>How fast is C++26’s std::hive?</title>
		<link>https://lemire.me/blog/2026/08/02/how-fast-is-c26s-stdhive/</link>
					<comments>https://lemire.me/blog/2026/08/02/how-fast-is-c26s-stdhive/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 02 Aug 2026 17:00:10 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22750</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/si3ao-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />C++26 adds a new container to the standard library: std::hive. It is meant to occupy the ground between std::vector and std::list. Like a vector, it keeps its elements in contiguous blocks of memory, so scanning it does not require you to chase a pointer for every element. Like a list, it never moves an element &#8230; <a href="https://lemire.me/blog/2026/08/02/how-fast-is-c26s-stdhive/" class="more-link">Continue reading <span class="screen-reader-text">How fast is C++26’s std::hive?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/08/si3ao-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>C++26 adds a new container to the standard library: <code>std::hive</code>. It is meant to occupy the ground between <code>std::vector</code> and <code>std::list</code>. Like a vector, it keeps its elements in contiguous blocks of memory, so scanning it does not require you to chase a pointer for every element. Like a list, it never moves an element once it has been inserted: your pointers, references and iterators stay valid, and you may erase any element in constant time without disturbing the others.</p>
<p>Internally, a hive is a linked list of blocks. Each block carries a <em>skipfield</em>: a small integer per slot that tells the iterator how many erased slots to jump over.</p>
<p>No standard library ships <code>std::hive</code> yet to my knowledge. Fortunately there is an implementation (<a href="https://github.com/mattreecebentley/plf_hive">plf::hive</a> by Matt Bentley) as a single header file that you can use today.</p>
<p>I use elements of type <code>uint64_t</code>, GCC 16.1 with <code>-O3 -march=native</code>, on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core. Numbers are nanoseconds per element, along with the cycles and instructions retired per element.</p>
<p>We start from an empty container and append a million values. The container is then destroyed.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">ns/element</th>
<th style="text-align: right;">instructions/element</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code> (<code>reserve</code>)</td>
<td style="text-align: right;">0.29</td>
<td style="text-align: right;">8.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">0.81</td>
<td style="text-align: right;">8.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">1.57</td>
<td style="text-align: right;">16.2</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code> (<code>reserve</code>)</td>
<td style="text-align: right;">1.76</td>
<td style="text-align: right;">17.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">14.22</td>
<td style="text-align: right;">220.0</td>
</tr>
</tbody>
</table>
<p>A <code>std::list</code> needs one allocation per element, and glibc&#8217;s malloc and free together cost over 200 instructions per element. It is an order of magnitude behind everyone else. That is not news.</p>
<p>The interesting comparison is vector against hive. A hive is about twice the cost of a vector, and it needs twice the instructions. This is the price of the skipfield: every insertion writes an element <em>and</em> a skipfield entry, and maintains the block bookkeeping. Note that calling <code>reserve</code> on a hive does not help in my experiments.</p>
<p>Next we iterate over the the container and sum the values.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">ns/element</th>
<th style="text-align: right;">cycles/element</th>
<th style="text-align: right;">instructions/element</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">0.22</td>
<td style="text-align: right;">0.78</td>
<td style="text-align: right;">1.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">1.51</td>
<td style="text-align: right;">5.27</td>
<td style="text-align: right;">4.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">1.77</td>
<td style="text-align: right;">6.18</td>
<td style="text-align: right;">9.0</td>
</tr>
</tbody>
</table>
<p>A hive iterates no faster than a linked list here, slightly slower, in fact, and about eight times slower than a vector. (Update: Joseph Garvin points out that I measure the happy case for the <code>std::list</code> in this instance where all the entries were allocated in sequence. The worst case scenario for <code>std::list</code> when the nodes are all over the heap can be much slower.)</p>
<p>The vector loop retires one instruction per element and finishes in 0.78 cycles: the processor is executing several elements at once. This is possible because the <code>std::vector</code> implementation benefits from autovectorization: the compiler recognizes that it can load several words at once in wide (SIMD). Further, it does not have to check the bitfield like the <code>std::hive</code> data structure.</p>
<p>We can check this. Walk the same container with two independent iterators, one starting halfway in, and count the cost per element visited:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">one traversal</th>
<th style="text-align: right;">two interleaved traversals</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">0.78 cycles</td>
<td style="text-align: right;">0.79 cycles</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">5.27 cycles</td>
<td style="text-align: right;">3.02 cycles</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">6.18 cycles</td>
<td style="text-align: right;">3.10 cycles</td>
</tr>
</tbody>
</table>
<p>The vector does not care: it was already throughput-bound. The hive and the list get nearly twice as fast per element, because two independent chains can be in flight at once. Hive iteration is latency-bound, exactly like list iteration. It merely has better locality.</p>
<p>That locality does show up when the data gets big. At ten million elements the list falls apart while the hive holds steady:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">100K</th>
<th style="text-align: right;">1M</th>
<th style="text-align: right;">10M</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">0.08</td>
<td style="text-align: right;">0.22</td>
<td style="text-align: right;">0.32</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">1.48</td>
<td style="text-align: right;">1.51</td>
<td style="text-align: right;">3.51</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">1.76</td>
<td style="text-align: right;">1.77</td>
<td style="text-align: right;">1.96</td>
</tr>
</tbody>
</table>
<p>Erasing is what a hive is for, so it would be unfair not to look. I erase half the elements at scattered positions using <code>std::remove_if</code>:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">ns per original element</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">2.1</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">3.0</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">77.4</td>
</tr>
</tbody>
</table>
<p>The hive wins, but by less than you might expect, and at ten million elements the ordering reverses (1.3 ns for the vector against 2.5 for the hive). <code>std::remove_if</code> is a single streaming pass, and streaming passes are cheap. Of course the vector moved every surviving element and invalidated every pointer into it, which is precisely what a hive promises not to do.</p>
<p>Memory, measured by asking glibc how many bytes it has handed out, per live element:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">container</th>
<th style="text-align: right;">after building</th>
<th style="text-align: right;">after <code>shrink_to_fit</code></th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::vector</code></td>
<td style="text-align: right;">8.4</td>
<td style="text-align: right;">8.0</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::hive</code></td>
<td style="text-align: right;">9.4</td>
<td style="text-align: right;">9.4</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;"><code>std::list</code></td>
<td style="text-align: right;">32.0</td>
<td style="text-align: right;"></td>
</tr>
</tbody>
</table>
<p>A hive costs about a byte per element over a vector, for a payload of eight bytes, when the vector is packed tight. A list costs more due to the overhead of the linked list.</p>
<p>A vector built by <code>push_back</code> has a capacity that typically exceeds its size. Thus even if you have 8-byte entries, you will use, on average, more than 8 bytes per entry even for large vectors. You can recover the excess capacity with the <code>shrink_to_fit</code> method.</p>
<p>What should we conclude?</p>
<p>The <code>std::hive</code> data structure is not a faster vector. But it is a much better <code>std::list</code>. It gives you the same guarantees that make people reach for a list, stable references, cheap erasure anywhere, while using less memory.</p>
<p><a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/08/02">My source code is available</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/08/02/how-fast-is-c26s-stdhive/feed/</wfw:commentRss>
			<slash:comments>10</slash:comments>
		
		
			</item>
		<item>
		<title>Memory-level parallelism: AMD is the king</title>
		<link>https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/</link>
					<comments>https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 25 Jul 2026 15:07:52 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22742</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/Bh4a1-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />When your program asks for memory that is not in cache, the processor has to go to RAM. That trip costs on the order of 100 nanoseconds. On a 3 GHz core, that is about 300 cycles of doing nothing. Memory latency has not improved in ten years. The 2016 Broadwell answers a random access &#8230; <a href="https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/" class="more-link">Continue reading <span class="screen-reader-text">Memory-level parallelism: AMD is the king</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/Bh4a1-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>When your program asks for memory that is not in cache, the processor has to go to RAM. That trip costs on the order of 100 nanoseconds. On a 3 GHz core, that is about 300 cycles of doing nothing.</p>
<p>Memory latency has not improved in ten years. The 2016 Broadwell answers a random access in 100 ns. The 2025 Turin, with DDR5-6400 and every advantage of a decade of progress, takes 140 ns. It got worse.</p>
<p>The good news is that a modern core does not have to sit still. It can issue a second request before the first one comes back, and a third, and a tenth. The number of requests a single core can keep in flight is its memory-level parallelism. It is one of the most important numbers in software performance, and one of the least advertised: you will not find it on a spec sheet.</p>
<p>Thankfully, memory-level parallelism has improved a lot. To measure it, I use my <a href="https://github.com/lemire/testingmlp">testingmlp</a> benchmark. The idea is a pointer chase. We build a 1 GiB array containing a single random cycle covering every element: each element holds the index of the next. Following the cycle is inherently serial. Each load has to complete before you know the address of the next one, so a single chase measures pure memory latency and nothing else. Then we run several such chases at once, from different starting points on the same cycle. We call these <em>lanes</em>. With two lanes, the core has two independent loads to work on. With twenty, twenty. We increase the number of lanes and watch the throughput. When adding a lane stops helping, we have found the limit. As my metric, I use the total estimated bandwidth.</p>
<p>I ran experiments on the Amazon cloud (AWS). The bandwidth shape is the same everywhere: a steep, nearly linear climb as we add lanes, then a knee, then a plateau. <br />
<a href="http://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves.png"><img loading="lazy" decoding="async" src="http://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves-723x1024.png" alt="" width="660" height="935" class="alignnone size-large wp-image-22745" srcset="https://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves-723x1024.png 723w, https://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves-212x300.png 212w, https://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves-768x1087.png 768w, https://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves-1085x1536.png 1085w, https://lemire.me/blog/wp-content/uploads/2026/07/mlp-curves.png 1159w" sizes="auto, (max-width: 660px) 100vw, 660px" /></a></p>
<p>How did it evolve over time? Intel went from 10 to 30, meaning that a single Intel core can sustain 30 memory requests at once in practice. AMD went from 15 to 58. Graviton went from 6 to 19.</p>
<p>Intel was flat for a long time. Broadwell and Cascade Lake both sit at 10 concurrent misses. Ice Lake doubled it to 20. Granite Rapids is at 30. Intel has roughly tripled in a decade, with all the gain arriving in the last two generations.</p>
<p>AMD started ahead and stayed ahead, then jumped. Naples was already at 15 in 2018, when Intel was at 10. Milan reached 22. And then Turin does something different in kind: 58 concurrent cache lines from a single core.</p>
<p>Graviton 1 was a toy: 6 concurrent misses. Graviton 2 doubled it, Graviton 3 went to 17, and then Graviton 4 essentially stood still at 18. Graviton 5 only reaches 19. But look at the latency panel: since 2017, Graviton 5 is the only chip in this entire collection that made a random access <em>faster</em> than its predecessor. <a href="https://www.amazon.science/blog/graviton5s-improved-design-increases-speed-and-energy-efficiency-beyond-moores-law">AWS advertised better DRAM latency for Graviton 5</a>, and that claim holds up.</p>
<p>So who wins? On bandwidth and memory-level parallelism, it is AMD, and it is not close. The Zen 5 core in the <code>m8a</code> instances sustains 58 concurrent cache-line fetches and 24.5 GiB/s of random-access throughput from one core. AMD is roughly twice as fast as Intel.</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Instance</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Year</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Processor</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Memory</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Latency</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Peak BW</th>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Concurrency</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">m8i.large</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">2025</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Xeon 6975P-C, Granite Rapids</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">DDR5-7200</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">133 ns</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">13.3 GiB/s</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">30</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">m8a.large</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">2025</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">EPYC 9R45, Zen 5 (Turin)</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">DDR5-6400</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">142 ns</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">24.5 GiB/s</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">58</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">m9g.large</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">2026</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">Graviton 5, Neoverse V3</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">DDR5-8800</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">96 ns</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">12.0 GiB/s</td>
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">19</td>
</tr>
</tbody>
</table>
<p><em>The raw output, the system information from each machine, and the scripts are <a href="https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/07/25">in the usual place</a>.</p>
<p><span data-offset-key="bfh9p-0-0">Note that Apple Silicon </span></em><span data-offset-key="bfh9p-1-0"><a href="https://lemire.me/blog/2025/07/09/memory-level-parallelism-apple-m2-vs-apple-m4/" rel="noopener noreferrer nofollow" target="_blank" role="link" class="css-1jxf684 r-bcqeeo r-1ttztb7 r-qvutc0 r-poiln3 r-1inkyih r-rjixqe r-1ddef8g r-tjvw6i r-1loqt21">does even better, </a></span><em><span data-offset-key="bfh9p-2-0">but it is another category.</span></em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Does a PhD Pay Off?</title>
		<link>https://lemire.me/blog/2026/07/24/does-a-phd-pay-off/</link>
					<comments>https://lemire.me/blog/2026/07/24/does-a-phd-pay-off/#respond</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Fri, 24 Jul 2026 20:13:57 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22737</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/kfKMS-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />Every week, I discuss with people who want to get a PhD. For years, I have been advising people not to pursue a PhD. It may come as a surprise to some. You would expect people with a PhD to earn more money. Individuals who complete doctorates tend to have higher cognitive abilities and greater &#8230; <a href="https://lemire.me/blog/2026/07/24/does-a-phd-pay-off/" class="more-link">Continue reading <span class="screen-reader-text">Does a PhD Pay Off?</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/kfKMS-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>Every week, I discuss with people who want to get a PhD. For years, I have been advising people not to pursue a PhD. It may come as a surprise to some.</p>
<p>You would expect people with a PhD to earn more money. Individuals who complete doctorates tend to have higher cognitive abilities and greater motivation. But smarter people tend to earn more, period.</p>
<p>So do people with a PhD earn more?</p>
<p>Historically, PhD holders earn more, but the bulk of the observed advantage is concentrated among those who get a professorship after the PhD. And there is no certain path from the PhD to a professorship. We have been producing many more PhDs than we have professorship, for decades. And the disparity is ever growing.</p>
<p>When I entered university at the beginning of the 1990s, about 0.5% of the Canadian population had a PhD. This has nearly tripled today, and it is fast increasing. Something of the order of one person out of 80 has a PhD. Comparatively, there is roughly one professor or university-level instructor per 900 people. With a fast aging population, we simply do not need many more professors and instructors than we already have.</p>
<p>There are specific fields where some jobs are difficult to get without a PhD. Machine learning is one such example. Many people in the industry have a PhD, and they tend to select those who also have a PhD. Further, there is a somewhat direct relationship between the work you might do during your PhD, if you are any good, and the actual work you might do later. It is much less clear in a lot of other disciplines.</p>
<p>The most significant economic cost of a PhD is not tuition but the years of delayed full-time earnings and career progression. In the tech industry, it is typical to award half a year of experience for each year spent on a PhD. This means that even though the individual starting with a PhD might earn more starting out, they are not necessarily getting a higher lifetime income. </p>
<p>Benjamin et al. (2025) find that the early-career benefits a PhD can be effectively zero:</p>
<blockquote>
<p>In the short run, pursuing a PhD entails substantial opportunity costs. Early-career earnings for PhD graduates are significantly lower than those of individuals with master’s or professional degrees, reflecting prolonged enrolment and delayed entry into the labour market. These costs are especially high for non-completers, particularly those who exit the program after several years without earning a credential. Over the lifecycle, earnings do eventually recover (and surpass those of bachelor’s and master’s graduates) but only under specific conditions. The most favourable long-run outcomes are concentrated among those who secure academic employment and remain in full-time work late into life. This “double premium,” combining higher earnings and longer careers, plays a central role in shaping the average return to a PhD. Outside academia, PhD holders resemble master’s graduates in both earnings and employment patterns.</p>
</blockquote>
<p>Thus, the financial case for a PhD is narrower than people assume. If you fail to get a professorship, or you want an early retirement, you may very well end up with a poor outcome. And it is not getting better over time. </p>
<p><em>References</em></p>
<ul>
<li>Altonji, J. G., &amp; Zhu, Z. (2025). Returns to specific graduate degrees: Estimates using Texas administrative records (NBER Working Paper No. 33530). National Bureau of Economic Research. https://www.nber.org/papers/w33530</li>
<li>Benjamin, D., Miloucheva, B., &amp; Vigezzi, N. (2025). The opportunity cost of a PhD: Spending your twenties (Working Paper No. 802). University of Toronto, Department of Economics. https://www.economics.utoronto.ca/public/workingPapers/tecipa-802.pdf</li>
<li>Cooper, P. Is grad school worth it? A comprehensive return on investment analysis. Foundation for Research on Equal Opportunity. https://freopp.org/whitepapers/is-grad-school-worth-it-a-comprehensive-return-on-investment-analysis/</li>
</ul>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/07/24/does-a-phd-pay-off/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>From Institutions to Individuals: the White House Report on Revitalizing U.S. Scientific Leadership</title>
		<link>https://lemire.me/blog/2026/07/22/from-institutions-to-individuals-the-white-house-report-on-revitalizing-u-s-scientific-leadership/</link>
					<comments>https://lemire.me/blog/2026/07/22/from-institutions-to-individuals-the-white-house-report-on-revitalizing-u-s-scientific-leadership/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Wed, 22 Jul 2026 14:32:18 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22727</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/Capture-decran-le-2026-07-22-a-10.24.40-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />In 1945, Vannevar Bush published a report entitled Science: The Endless Frontier. His thesis was that prosperity follows from basic research. The report was highly influential in the United States and elsewhere. It led to the creation of an entirely new government bureaucracy. With this report, Bush popularized the linear model of innovation: innovation (such &#8230; <a href="https://lemire.me/blog/2026/07/22/from-institutions-to-individuals-the-white-house-report-on-revitalizing-u-s-scientific-leadership/" class="more-link">Continue reading <span class="screen-reader-text">From Institutions to Individuals: the White House Report on Revitalizing U.S. Scientific Leadership</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/Capture-decran-le-2026-07-22-a-10.24.40-150x150.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>In 1945, Vannevar Bush published a report entitled <a href="https://www.govinfo.gov/content/pkg/GOVPUB-PR32_400-e7966ee70a4f7b47f862431c9776f727/pdf/GOVPUB-PR32_400-e7966ee70a4f7b47f862431c9776f727.pdf">Science: The Endless Frontier</a>. His thesis was that prosperity follows from basic research. The report was highly influential in the United States and elsewhere. It led to the creation of an entirely new government bureaucracy.</p>
<p>With this report, Bush popularized the linear model of innovation: innovation (such as medical cures) flows sequentially from basic research to applied research to development to production and diffusion. Grow basic research, and the rest will follow.</p>
<p>When Bush wrote his report, basic research was not usually supported directly by the state. We did not have a large basic research infrastructure. And yet, the West had just lived through an unprecedented period of rapid scientific progress: the theory of evolution, electromagnetism, radio communication, special and general relativity, quantum mechanics, nuclear technology, rockets, the combustion engine, and more. We would get the invention of the transistor only two years after Bush’s report. We also did not have today’s peer-review mechanism.</p>
<p>Even though Bush’s report has been viewed as a piece of genius that unlocked a golden era of scientific prosperity, I believe that the linear model of innovation is hopelessly naïve. I believe the thesis that a large bureaucracy delivering funding to other bureaucracies (such as universities) is how we get innovation is absurd. Except perhaps in the domain of computing (“bits”), we have been largely stagnant technologically since about the 1970s. So Bush’s model failed over time. To be clear, it might have worked for a while by encouraging more young people to study engineering and science. It may also have shone a favorable light on a few enterprising professors who got to promote useful ideas.</p>
<p>If you visit a research lab today in a leading university, what you are most likely to see is a boring bureaucracy that caters to whatever is politically favorable at the moment—a bureaucracy that plays it safe and avoids controversy. You see young people seeking well-paid jobs, going through the motions with often little genuine interest in, say, curing cancer. We have never published so many research papers—the volume has been growing exponentially ever since Bush wrote his report—but it is doubtful that this is how technological breakthroughs are achieved.</p>
<p>The evidence is overwhelming that shoddy science is widespread. We have a severe reproducibility crisis: if you redo an experiment (even a highly cited one), you are likely to fail to reproduce the results. This affects psychology, medicine, and many other fields. The system does not particularly care because the incentives to get things right are not there. As long as the work is politically aligned, solidity of the results seems secondary.</p>
<p>There was a TV show (<em>The Big Bang Theory</em>) where the main character, Sheldon Cooper—an awkward genius—gets to work on crazy ideas. That is how Bush imagined it: fund young people like Sheldon Cooper, and you will get extraordinary breakthroughs. In the real world, Sheldon would not get very far on campus. I have met misfits like him. When they are incapable of playing the political game, the system crushes them. But even if that were not the case, extraordinary intelligence needs to be applied to the right problems to be of value. You could have a ChatGPT that is brighter than any of us in every possible way, and it could still be deployed simply to fill out forms faster and better than we do—it may not cure cancer.</p>
<p>The American government has just released what might be considered an update to Bush’s report. Michael Kratsios wrote a report entitled <a href="https://www.whitehouse.gov/wp-content/uploads/2026/07/Science-A-New-Golden-Age.pdf">Science, A New Golden Age</a>. The report states outright that the linear model no longer holds. It states what I have argued for vehemently: innovation is not a linear process. Take large-language models, for example, which can be used by engineers and scientists to further their research. I have also argued that the success of large-language models today has as much to do with the users as with the researchers.</p>
<p>At this point, some people engage in the following type of rhetoric: if we had not invented calculus, we would not have AI today; therefore, calculus caused AI. But you could also say that the subsidized nail factory in the Soviet Union, which made overpriced and bad nails, was necessary to hold Landau’s house together, and that without those nails we would not have the theory of Landau levels. The causality argument goes in all directions.</p>
<p>Innovation is the result of a complex system. We see that the United States and, more recently, China are innovative countries. In 2026, you do not go to France for the latest advances. The evidence is overwhelming that scientific and technological progress depends as much on culture as on anything else. It is not something to be managed by bureaucrats.</p>
<p>One of the cultural ingredients that seems essential is meritocracy. You must put the people who are good at building on top of your hierarchy. This does not happen magically. You need a set of incentives in which rewarding the wrong people is costly.</p>
<p>What does Kratsios propose? Many interesting ideas that, I expect, could renew our culture. He proposes to break out of the Cold War–era funding model. Today, the research funding mechanism is centered around the government giving money to the university bureaucracy. The grant might be in the name of one professor, but the recipient is still the university. In the new model, instead of funding universities, the government would assign money directly to individuals in various ways (short grants, prizes, and so forth). This would shift power away from administrators toward individuals who know how to get things done. It would also neutralize some of the political power of the current mandarin class of scientists who control access to the top positions.</p>
<p>The report recommends restoring permissionless innovation. It is sometimes poorly understood how limited the system has become. I once had a graduate student undertake interviews with practitioners. This required an ethics approval which, in her case, took a few months to obtain. Again, the system has built up political structures that seek to block innovation it does not like. They need to be torn down, the sooner the better.</p>
<p>The report has many other interesting recommendations. One that I particularly like is an AI-guided agenda. We need to hook up our brand-new AIs to experimental devices. We are not going to cure aging with chatbots. We need experiments on a massive scale.</p>
<p>Will Kratsios’s vision move from report to reality? History shows that cultural and institutional change is never easy. Yet the stakes could not be higher. By embracing meritocracy, permissionless innovation, and ambitious AI-augmented experimentation, we have a genuine chance to escape decades of stagnation and rekindle the spirit of discovery that once defined the West. The opportunity is before us. It must not be squandered.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/07/22/from-institutions-to-individuals-the-white-house-report-on-revitalizing-u-s-scientific-leadership/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Using AI to build your own software</title>
		<link>https://lemire.me/blog/2026/07/16/using-ai-to-build-your-own-software/</link>
					<comments>https://lemire.me/blog/2026/07/16/using-ai-to-build-your-own-software/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Thu, 16 Jul 2026 20:00:21 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22724</guid>

					<description><![CDATA[A few years ago, a friend of mine was stuck. He needed to quickly process over a hundred high-quality images according to a complicated sequence. He was using Photoshop, but it was going to take him days. Initially, he asked for my help, could I do the manual labor? I spent 15 minutes writing a &#8230; <a href="https://lemire.me/blog/2026/07/16/using-ai-to-build-your-own-software/" class="more-link">Continue reading <span class="screen-reader-text">Using AI to build your own software</span></a>]]></description>
										<content:encoded><![CDATA[<p>A few years ago, a friend of mine was stuck. He needed to quickly process over a hundred high-quality images according to a complicated sequence. He was using Photoshop, but it was going to take him days. Initially, he asked for my help, could I do the manual labor? I spent 15 minutes writing a script with ImageMagick that processed all the images in seconds, but in a completely automated way.</p>
<p>When my kids were young, instead of helping them study algebra and grammar, I wrote small JavaScript apps for them to use. I built a small collection of educational tools.</p>
<p>The great success story of AI for me is exactly this: AI helps you write your own tools, faster and better.</p>
<p>Last night, I was struggling with videos I had to process. I wanted to add nice subtitles to them. There are software applications for that, but they require manual labor and don’t always work the way I want them to. After a long night, I had an insight: why don’t I ask my AI to help build the automated tool I need? So I did—and it worked really well, very quickly. So what’s the lesson here? Maybe that we should spend more time building our own software for our own personal use than we used to.</p>
<p><iframe loading="lazy" width="560" height="315" src="https://www.youtube.com/embed/Falc7QnHy7k?si=REbbbCN_EYdOgdaL" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"></iframe></p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/07/16/using-ai-to-build-your-own-software/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>X just gave us an interface that AI agents can use. I pointed it at my own posts.</title>
		<link>https://lemire.me/blog/2026/07/11/x-just-gave-us-an-interface-that-ai-agents-can-use-i-pointed-it-at-my-own-posts/</link>
					<comments>https://lemire.me/blog/2026/07/11/x-just-gave-us-an-interface-that-ai-agents-can-use-i-pointed-it-at-my-own-posts/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sat, 11 Jul 2026 19:53:17 +0000</pubDate>
				<category><![CDATA[]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22714</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/image-5-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />I have been on X for a long time. Like most people who post regularly, I have a gut feeling for what might interest people. I post in the morning. Longer posts seem to do better. But gut feelings are not measurements. And until recently, digging into your own posting data meant either clicking around &#8230; <a href="https://lemire.me/blog/2026/07/11/x-just-gave-us-an-interface-that-ai-agents-can-use-i-pointed-it-at-my-own-posts/" class="more-link">Continue reading <span class="screen-reader-text">X just gave us an interface that AI agents can use. I pointed it at my own posts.</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/07/image-5-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p>I have been on X for a long time. Like most people who post regularly, I have a gut feeling for what might interest people. I post in the morning. Longer posts seem to do better.</p>
<p>But gut feelings are not measurements. And until recently, digging into your own posting data meant either clicking around the web UI or writing custom scripts. Neither is particularly friendly when you want to ask <em>ad hoc</em> questions with an AI assistant.</p>
<p>X recently launched hosted <a href="https://modelcontextprotocol.io/">MCP</a> servers: official endpoints that AI tools can connect to. MCP is a protocol for plugging tools into language models: the model can search posts, manage bookmarks, fetch trends, and so on. In practice, I connected an AI coding agent to the X MCP server and simply started asking questions about my account.</p>
<p>I spent a session exploring about two months of my own activity. Here is what I found interesting.</p>
<p>Over roughly sixty days (mid-May through mid-July 2026), I published on the order of 435 posts that were not pure retweets of other people—mostly a mix of original posts, replies, and a few X Articles. The agent pulled them through the MCP tools, kept the public metrics (likes, views, reposts), and ran simple analyses.</p>
<p>I asked for every post to be binned by local hour of day (America/Toronto, Eastern time), and for each hour: how many posts, and the min / median / max view count.</p>
<p>My posting is heavily skewed toward the morning:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: left; font-weight: 600; border-bottom: 2px solid #d0d7de; padding: 0.65rem 0.8rem; vertical-align: top;">Local hour</th>
<th style="text-align: right;">Posts</th>
<th style="text-align: right;">Median views</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">08:00–08:59</td>
<td style="text-align: right;">45</td>
<td style="text-align: right;">454</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">09:00–09:59</td>
<td style="text-align: right;">58</td>
<td style="text-align: right;">1,067</td>
</tr>
<tr style="background: #ffffff;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">10:00–10:59</td>
<td style="text-align: right;">30</td>
<td style="text-align: right;">194</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="padding: 0.65rem 0.8rem; border-bottom: 1px solid #eaeef2; vertical-align: top;">11:00–11:59</td>
<td style="text-align: right;">42</td>
<td style="text-align: right;">284</td>
</tr>
</tbody>
</table>
<p>The 9 a.m. hour is both my busiest and, among busy hours, my strongest by median views. The overall median across all hours was only about 188 views, so most of what I write is quiet. The distribution is heavy-tailed: a few posts get tens or hundreds of thousands of impressions; the rest are background noise.</p>
<p>I then binned posts by character length in steps of 25 characters (using the text as returned by the API, including short <code>t.co</code> URLs).</p>
<p>The bulk of my writing is short, often a reply of a few dozen characters:</p>
<table style="width: 100%; border-collapse: collapse; margin: 1.2rem 0; border: 1px solid #d0d7de; border-radius: 8px; overflow: hidden; box-shadow: 0 1px 2px rgba(0,0,0,0.04);">
<thead style="background: #f6f8fa; color: #24292f;">
<tr>
<th style="text-align: right;">Characters</th>
<th style="text-align: right;">Posts</th>
<th style="text-align: right;">Median likes</th>
<th style="text-align: right;">Max likes</th>
</tr>
</thead>
<tbody>
<tr style="background: #ffffff;">
<td style="text-align: right;">0–25</td>
<td style="text-align: right;">44</td>
<td style="text-align: right;">1</td>
<td style="text-align: right;">195</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">25–50</td>
<td style="text-align: right;">87</td>
<td style="text-align: right;">1</td>
<td style="text-align: right;">60</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">50–75</td>
<td style="text-align: right;">69</td>
<td style="text-align: right;">0</td>
<td style="text-align: right;">98</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">75–100</td>
<td style="text-align: right;">51</td>
<td style="text-align: right;">1</td>
<td style="text-align: right;">61</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">100–125</td>
<td style="text-align: right;">33</td>
<td style="text-align: right;">1</td>
<td style="text-align: right;">32</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">…</td>
<td style="text-align: right;"></td>
<td style="text-align: right;"></td>
<td style="text-align: right;"></td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">175–200</td>
<td style="text-align: right;">12</td>
<td style="text-align: right;">5</td>
<td style="text-align: right;">58</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">200–225</td>
<td style="text-align: right;">17</td>
<td style="text-align: right;">4</td>
<td style="text-align: right;">456</td>
</tr>
<tr style="background: #ffffff;">
<td style="text-align: right;">275–300</td>
<td style="text-align: right;">19</td>
<td style="text-align: right;">4</td>
<td style="text-align: right;">385</td>
</tr>
<tr style="background: #fbfcfd;">
<td style="text-align: right;">300–325</td>
<td style="text-align: right;">48</td>
<td style="text-align: right;">46.5</td>
<td style="text-align: right;">470</td>
</tr>
</tbody>
</table>
<p>Under about 175 characters, the median stays at zero or one like. Around full-length posts (roughly the old 280-character regime and a bit beyond), engagement jumps. The 300–325 character band is where a large fraction of my “serious” posts live, and the median likes there are an order of magnitude higher than for short replies.</p>
<p>I also asked the AI to identify the posts that had the most likes, the following types of posts were liked:</p>
<ul>
<li>AI vs. “experts” claiming models are nowhere near human intelligence</li>
<li>Go adding SIMD-style data-parallelism to the standard library</li>
<li>SIMD-accelerated data processing talks and library notes (JSON, string→integer maps, vulnerability-report fatigue)</li>
<li>Nvidia hardware, university AI-cheating, C++ contracts</li>
</ul>
<p>The interesting part is the workflow. I did not export a CSV by hand and open Excel. I asked an agent, connected to X’s MCP server. If AI agents can do this for one account’s metrics, they can do it for bug trackers, logs, paper drafts, and codebases.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/07/11/x-just-gave-us-an-interface-that-ai-agents-can-use-i-pointed-it-at-my-own-posts/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Chatting with an AI Won’t Make You a Top Programmer</title>
		<link>https://lemire.me/blog/2026/06/21/chatting-with-ai-wont-make-you-a-top-programmer/</link>
					<comments>https://lemire.me/blog/2026/06/21/chatting-with-ai-wont-make-you-a-top-programmer/#comments</comments>
		
		<dc:creator><![CDATA[Daniel Lemire]]></dc:creator>
		<pubDate>Sun, 21 Jun 2026 17:51:16 +0000</pubDate>
				<category><![CDATA[]]></category>
		<category><![CDATA[essay]]></category>
		<guid isPermaLink="false">https://lemire.me/blog/?p=22706</guid>

					<description><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/06/image-4-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" />When I was a kid, most people did not know how to type. We took typing class. The final exam was a speed test: words per minute. Today, you will not impress anyone by saying you can type. In fact, cursive writing is fading. Kids increasingly cannot read or write it. We type constantly. We &#8230; <a href="https://lemire.me/blog/2026/06/21/chatting-with-ai-wont-make-you-a-top-programmer/" class="more-link">Continue reading <span class="screen-reader-text">Chatting with an AI Won’t Make You a Top Programmer</span></a>]]></description>
										<content:encoded><![CDATA[<img width="150" height="150" src="https://lemire.me/blog/wp-content/uploads/2026/06/image-4-150x150.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" /><p dir="auto">When I was a kid, most people did not know how to type. We took typing class. The final exam was a speed test: words per minute. Today, you will not impress anyone by saying you can type. In fact, cursive writing is fading. Kids increasingly cannot read or write it. We type constantly. We forget how many skills are learned, and how often some of these skills have faded.</p>
<p dir="auto">But not everything fades. Socrates would be immensely popular today as a teacher. I still buy and recommend paper books.</p>
<p dir="auto">Is reading and writing code more like Socrates, or more like cursive writing? There are clear signs that code could become like cursive writing. This year, I have met more than one student who could use AI to build an application but could not read or write code. It is not new. Software has long had non-technical people who describe what they built or designed. In fact, in much of the industry, the standard view was that once you had a university degree, you no longer coded. Coding was for monkeys or low-status employees. Top engineers paid a million dollars a year at Google or Meta know how to write code. They often read and write assembly and TypeScript. They know it all.</p>
<p dir="auto">Why the discrepancy?</p>
<p dir="auto">We pay an engineer a million dollars because he understands concepts few others grasp. He outruns others because he sees the problems more deeply. Reading and writing large amounts of code is part of how you gain those insights. Chatting with an AI will not make you a top 1% programmer. In the future, top engineers might read more code than anyone could in the past. These engineers will not be everywhere, but they will pack a punch. “But Daniel, people say programming is solved. Why read or write code?” Be careful with your models. When television arrived, some predicted it would replace the university lecturer. In some respects the model was correct, yet it did not happen. The lecturer’s job was never to deliver a TV show. The Google engineer paid a million dollars was never a machine that produces code. Nobody actually wants code, any more than they want raw text.</p>
<p dir="auto">In fact, I predict a bifurcation in the tooling. The best engineers will work with tools that maximize their understanding of the code. I believe that reading and writing code, at a high level, is more like studying Socrates than like cursive writing. It is a necessary mental labor that does not become obsolete just because we have better tools for generating output.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://lemire.me/blog/2026/06/21/chatting-with-ai-wont-make-you-a-top-programmer/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
			</item>
	</channel>
</rss>
