<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>John D. Cook</title>
	<atom:link href="http://www.johndcook.com/blog/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.johndcook.com/blog</link>
	<description>Applied Mathematics Consulting</description>
	<lastBuildDate>Sun, 04 Oct 2026 13:24:42 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://www.johndcook.com/wp-content/uploads/2020/01/cropped-favicon_512-32x32.png</url>
	<title>John D. Cook</title>
	<link>https://www.johndcook.com/blog</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Miquel&#8217;s pentagon theorem</title>
		<link>https://www.johndcook.com/blog/2026/10/04/miquels-pentagon-theorem/</link>
					<comments>https://www.johndcook.com/blog/2026/10/04/miquels-pentagon-theorem/#respond</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 12:34:22 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247975</guid>

					<description><![CDATA[<p>An earlier post presented an elegant plane geometry theorem discovered by the 19th century school teacher Auguste Miquel. This post presents his pentagon theorem. Start with a pentagon. It may be irregular, but it needs to be convex. Extend each of the sides of the pentagon to form a star, then draw give circles, one [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/10/04/miquels-pentagon-theorem/">Miquel’s pentagon theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>An <a href="https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/">earlier post</a> presented an elegant plane geometry theorem discovered by the 19th century school teacher Auguste Miquel. This post presents his pentagon theorem.</p>
<p>Start with a pentagon. It may be irregular, but it needs to be convex.</p>
<p>Extend each of the sides of the pentagon to form a star, then draw give circles, one through each of the triangles formed by a side of the pentagon and a vertex of the star.</p>
<p>The five circles intersect in pairs at ten points: the five verices of the pentagon and five new points. The five new points lie on a circle.</p>
<p><img fetchpriority="high" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/miquel_pentagon.png" width="480" height="553" /></p>
<p>The converse of this theorem is known as the five circles theorem.</p>The post <a href="https://www.johndcook.com/blog/2026/10/04/miquels-pentagon-theorem/">Miquel’s pentagon theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/10/04/miquels-pentagon-theorem/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Topological models of modal logic</title>
		<link>https://www.johndcook.com/blog/2026/10/04/topological-models-of-modal-logic/</link>
					<comments>https://www.johndcook.com/blog/2026/10/04/topological-models-of-modal-logic/#respond</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 12:25:19 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247985</guid>

					<description><![CDATA[<p>The previous post discussed a superficial connection between modal logic and topology, that both use the terms regular and normal to indicate added sets of axioms. McKinsey and Tarski developed a deeper connection between modal logic and topology that we&#8217;ll discuss here. Starting with a topological space X and a proposition p, define [[p]] as the set [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/10/04/topological-models-of-modal-logic/">Topological models of modal logic</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The <a href="https://www.johndcook.com/blog/2026/10/04/modal-topology/">previous post</a> discussed a superficial connection between modal logic and topology, that both use the terms <em>regular</em> and <em>normal</em> to indicate added sets of axioms. McKinsey and Tarski developed a deeper connection between modal logic and topology that we&#8217;ll discuss here.</p>
<p>Starting with a topological space <em>X</em> and a proposition <em>p</em>, define [[<em>p</em>]] as the set of points in <em>X</em> at which <em>p</em> is true. Define □<em>p</em> to be true at points in the interior of [[<em>p</em>]] and define ◇<em>p</em> to be true on the closure of [[<em>p</em>]].</p>
<p>You could think of □<em>p</em> as the points where <em>p</em> is robustly true. Not only is <em>p</em> true at <em>x</em>, there&#8217;s some wiggle room around <em>x</em>, i.e. an open set, in which <em>p</em> remains true.</p>
<p>You could think of ◇<em>p</em> as the points where we cannot rule out the possibility of <em>p</em> being true using open sets. If ◇<em>p</em> includes <em>x</em>, any open set containing <em>x</em> also contains part of ◇<em>p</em>, though it may also contain points outside of ◇<em>p.</em></p>
<h2>Regularity</h2>
<p>For any topology on <em>X</em>, the logic constructed above is normal. The axiom</p>
<p style="padding-left: 40px;">◇<em>p</em> ⇔ ¬ (□ ¬ <em>p</em>)</p>
<p>holds because the closure of a set is the complement of the interior of its complement [1].</p>
<p>Note that this is a regularity result for the modal logic, not the topology. The topology could be arbitrary, and not necessarily regular or normal in the topological sense.</p>
<h2>S4</h2>
<p>The logic constructed above also satisfies a couple more axioms. We have</p>
<p style="padding-left: 40px;">□<em>p</em> → <em>p</em></p>
<p>because the interior of a set is a subset of the set, and</p>
<p style="padding-left: 40px;">□<em>p</em> → □□<em>p</em></p>
<p>because the interior of the interior of a set is simply the interior. This means the modal logic corresponding to a topology satisfies the S4 axioms. You could say S4 is the logic that corresponds to the McKinsey and Tarski logic of all topological spaces.</p>
<h2>More logics and more topologies</h2>
<p>So S4 is the logic that corresponds to <em>all</em> topologies. We could look at more restricted topologies and ask what are their corresponding logics. Or we could start with a modal logic and ask whether there&#8217;s a topology that models that logic.</p>
<p>Interesting logics correspond to badly behaved topological spaces. Familiar topological spaces like the real line correspond to S4.</p>
<h3>Trivial modal logic</h3>
<p>The discrete topology corresponds to the trivial modal logic. All sets are open, and closed, so any set is the same as its interior and its closure. So □<em>p</em> and ◇<em>p</em> reduce to just <em>p</em>.</p>
<h3>S5</h3>
<p>For the indiscrete topology, □<em>p</em> corresponds to a proposition holding everywhere and ◇<em>p</em> corresponds to it holding somewhere. If the topological space has infinitely many points, the corresponding modal logic is S5. [2]</p>
<h3>Between S4 and S5</h3>
<p>The cofinite topology on an infinite set <em>X</em> defines a set <em>U</em> to be open if the complement of <em>U</em> is finite. The McKinsey-Tarski logic of the cofinite topology is somewhere between S4 and S5. You can show that the formula</p>
<p style="padding-left: 40px;"><em>p</em> ∧ ◇□<em>p</em> → □<em>p</em></p>
<p>holds, which doesn&#8217;t hold in S4, and the formula</p>
<p style="padding-left: 40px;">◇<em>p</em> → □◇<em>p</em></p>
<p>does not hold, though it must hold in S5.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2022/01/21/modal-logic-and-sf/">Modal logic and science fiction</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2021/12/30/modal-axioms/">Naming and numbering modal logics</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2018/10/30/modal-logic-security/">Modal logic and cybersecurity</a></li>
</ul>
<p>[1] We should also verify that if <em>A</em> ∩ <em>B</em> ⊂ <em>C</em>, then Interior(<em>A</em>) ∩ Interior(<em>B</em>) ⊂ Interior(<em>C</em>).</p>
<p>[2] Propositions can only have a finite number of terms. Having infinite points in the topological space prevents the corresponding logic from proving theorems that don&#8217;t necessarily hold in S5.</p>The post <a href="https://www.johndcook.com/blog/2026/10/04/topological-models-of-modal-logic/">Topological models of modal logic</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/10/04/topological-models-of-modal-logic/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Modal logic and topology</title>
		<link>https://www.johndcook.com/blog/2026/10/04/modal-topology/</link>
					<comments>https://www.johndcook.com/blog/2026/10/04/modal-topology/#respond</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 12:24:33 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Logic]]></category>
		<category><![CDATA[Topology]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247977</guid>

					<description><![CDATA[<p>You can&#8217;t say much about modal logic in general. You have to be more specific to get anywhere. You have to choose some axioms. Ideally the axioms you need for your application correspond to a named set of axioms that has been studied before. The situation is similar in point-set topology. You can&#8217;t say very [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/10/04/modal-topology/">Modal logic and topology</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>You can&#8217;t say much about modal logic in general. You have to be more specific to get anywhere. You have to choose some axioms. Ideally the axioms you need for your application correspond to a named set of axioms that has been studied before.</p>
<p>The situation is similar in point-set topology. You can&#8217;t say very much about a general topological space. You have to specify some separation axioms to get going.</p>
<h2>Bare bones</h2>
<h3>Modal logic</h3>
<p>A modal logic is any set of formulas in the modal language that:</p>
<ol>
<li>contains all propositional tautologies,</li>
<li>is closed under modus ponens, and</li>
<li>is closed under uniform substitution.</li>
</ol>
<p>In particular, this definition requires <em>nothing</em> of the modal operator □ (“box”). You just have propositional logic with a funny symbol added that could mean anything.</p>
<h3>Topology</h3>
<p>A topological space is a set <em>X</em> along with a set of subsets of <em>X</em> called open sets. The empty set and the full space <em>X</em> are open sets. Furthermore, the set of open sets is closed under finite intersections and arbitrary unions.</p>
<p>There&#8217;s not much you can say about topological spaces in general because, for example, the definition includes extreme cases such as the discrete topology (every subset of <em>X</em> is open) and the indiscrete topology (only the empty set and <em>X</em> are open).</p>
<h2>Regular and normal</h2>
<p>Like many areas of mathematics, logic and topology use the terms &#8220;regular&#8221; and &#8220;normal&#8221; to refer to systems with common choices of extra structure.</p>
<h3>Modal logic</h3>
<p>A regular modal logic is a normal modal logic with a second modal operator ◇ (&#8220;diamond&#8221;) that satisfies</p>
<p style="padding-left: 40px;">◇ <em>p</em> ⇔ ¬ (□ ¬ <em>p</em>)</p>
<p>and has the inference rule (<em>p</em> ∧ <em>q</em>) → <em>r</em> implies (□<em>p</em> ∧ □<em>q</em>) → □<em>r.</em></p>
<p>A modal logic is normal if it satisfies the axiom</p>
<p style="padding-left: 40px;">□ (<em>p</em> → <em>q</em>) → (□ <em>p</em> → □ <em>q</em>)</p>
<p>and the inference rule that if <em>p</em> is a theorem, □<em>p</em> is also a theorem.</p>
<h3>Topology</h3>
<p>Topology also uses <em>regular</em> and <em>normal</em> to refer to adding a few axioms.</p>
<p>A regular topological space is one in which you can separate points from closed sets. Given a point <em>x</em> and a closed set <em>F</em> not containing <em>x</em>, there exist disjoint open sets <em>U</em> and <em>V</em> such that <em>x</em> is contained in <em>U</em> and <em>F</em> is contained in <em>V</em>. [1]</p>
<p>A normal topological space is one in which you can separate disjoint closed sets.</p>
<p>For many mathematicians, a metric space is the weakest topology they&#8217;re interested in, and metric spaces are normal. But weaker topologies come up. The Zariski topology in algebraic geometry is not regular, and the weak topology on an infinite dimensional Banach space is regular but not normal.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2022/01/21/modal-logic-and-sf/">Modal logic and science fiction</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2021/12/30/modal-axioms/">Naming and numbering modal logics</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2018/10/30/modal-logic-security/">Modal logic and cybersecurity</a></li>
</ul>
<p>[1] Why do we use <em>F</em> to denote a closed set? It&#8217;s a convention that goes back to the French word <em lang="fr">fermé</em> for &#8220;closed.&#8221;</p>
<p>&nbsp;</p>The post <a href="https://www.johndcook.com/blog/2026/10/04/modal-topology/">Modal logic and topology</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/10/04/modal-topology/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Miquel&#8217;s pivot theorem</title>
		<link>https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/</link>
					<comments>https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/#respond</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 22:53:41 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247973</guid>

					<description><![CDATA[<p>Euclidean geometry dates back at least to Euclid (circa 300 BC), and so you might think it&#8217;s been pretty well picked over by now. And yet people still occasionally discover new plane geometry theorems. Some of these new theorems are complicated, asking question that the ancients would not have asked. But once in a while [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/">Miquel’s pivot theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Euclidean geometry dates back at least to Euclid (circa 300 BC), and so you might think it&#8217;s been pretty well picked over by now. And yet people still occasionally discover new plane geometry theorems.</p>
<p>Some of these new theorems are complicated, asking question that the ancients would not have asked. But once in a while someone discovers a gem that the ancients could have appreciated but didn&#8217;t find.</p>
<p>One example is Miquel’s pivot theorem [1]. The theorem was discovered in 1838, which relative to the timeline of Euclidean geometry makes it a recent discovery.</p>
<p>Choose a point on each side of a triangle. Then for each vertex draw a circle through it and the chosen points on the adjacent sides. Miquel&#8217;s theorem says the three circles meet in one point.</p>
<p>Here&#8217;s an example. For a trangle <em>ABC</em>, choose points <em>D</em>, <em>E</em>, and <em>F</em> on each side. The three circles described in the theorem intersect at <em>M</em>.</p>
<p><img decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/miquel1.png" width="480" height="440" /></p>
<p>Now the three points <em>D</em>, <em>E</em>, and <em>F</em> don&#8217;t have to be limited to the sides of the triangle; they can be on the line segment containing the side. Here&#8217;s an example where <em>D</em> is outside the triangle.</p>
<p><img decoding="async" class="size-medium aligncenter" src="https://www.johndcook.com/miquel2.png" width="480" height="438" /></p>
<p>And here&#8217;s an example where two of the chosen points, <em>D</em> and <em>F</em>, are outside the triangle. The three circles still intersect at one point <em>M</em>.</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/miquel3.png" width="480" height="527" /></p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2023/06/18/circle-through-three-points/">Equation of a circle through three points</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2023/01/23/van-aubels-theorem/">Van Aubel&#8217;s theorem</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2026/08/18/big-little-hexagon/">Big little hexagon</a></li>
</ul>
<p>[1] Miquel, Auguste (1838), <span lang="fr">&#8220;Mémoire de Géométrie&#8221;</span>, <span lang="fr">Journal de Mathématiques Pures et Appliquées</span>, 1: 485–487</p>The post <a href="https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/">Miquel’s pivot theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/10/03/miquels-pivot-theorem/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Servers in dawn-dusk orbit</title>
		<link>https://www.johndcook.com/blog/2026/09/25/dawn-dusk-orbit/</link>
					<comments>https://www.johndcook.com/blog/2026/09/25/dawn-dusk-orbit/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 11:45:31 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<category><![CDATA[Orbital mechanics]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247956</guid>

					<description><![CDATA[<p>Despite the predictions that no one would ever put build data centers in space, Google is starting on Thursday. Google&#8217;s prototype satellite will be one of 130 payloads on SpaceX&#8217;s Transporter 18 mission on October 1. The server will follow a dawn-dusk orbit, a special case of a sun-synchronous orbit (SSO), following the terminator line [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/25/dawn-dusk-orbit/">Servers in dawn-dusk orbit</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Despite the predictions that no one would ever put build data centers in space, Google is starting on Thursday. Google&#8217;s prototype satellite will be one of 130 payloads on SpaceX&#8217;s Transporter 18 mission on October 1.</p>
<p>The server will follow a dawn-dusk orbit, a special case of a sun-synchronous orbit (SSO), following the terminator line between daylight on dark on the earth below. A dawn-dusk orbit allows the satellite&#8217;s solar panels to stay in nearly continuous daylight, while also being in a relatively inexpensive low earth orbit (LEO). Geostationary orbit (GEO) would allow solar panels to always receive sunlight, but launching a satellite into GEO requires more fuel and so is more expensive.</p>
<p>Another advantage of LEO is that radiation levels are a couple orders of magnitude less than at GEO. Lower radiation means electronics do not need to be as hardened against radiation.</p>
<p>A dawn-dusk orbit would not be possible if the earth were perfectly spherical. The earth&#8217;s equatorial bulge makes it possible to design an orbit that precesses once per year. David Hammen explains this in an answer to a <a href="https://space.stackexchange.com/questions/60763/how-to-better-understand-how-dawn-dusk-orbits-work">question</a> on the Space Exploration Stack Exchange site.</p>
<blockquote><p>If the Earth had a spherically distributed gravitational field, a satellite&#8217;s right ascension of ascending node would be constant. … Fortunately, the Earth&#8217;s gravitational field is not spherical. The Earth&#8217;s rotation results in an equatorial bulge. This equatorial bulge causes RAAN to precess (or recess). …</p>
<p>Sun synchronous orbits are chosen so that RAAN precesses by 360 degrees per year, or a bit less than one degree per day. …</p>
<p>A dawn-dusk satellite is a special case of a sun synchronous orbit. … A dawn-dusk orbit typically does not quite follow the terminator. Following the terminator would require a rather high orbit.</p></blockquote>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2026/02/02/satellites-have-a-lot-of-room/">Satellites have a lot of room</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/02/28/max-min-orbital-speed/">Max and min orbital speed</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2024/12/23/starlink-configurations/">Starlink configurations</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/25/dawn-dusk-orbit/">Servers in dawn-dusk orbit</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/25/dawn-dusk-orbit/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Navigation with only addition, subtraction, and tables</title>
		<link>https://www.johndcook.com/blog/2026/09/23/navigation-minimum/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 14:11:52 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247950</guid>

					<description><![CDATA[<p>In the novel Carry On, Mr. Bowditch, a sailor asked Nathaniel Bowditch to teach him how to do navigational calculations, but the man only knows how to add and subtract by counting on his fingers. He had not heard of multiplication. Bowditch is surprised, but realizes if he made tables of logs of trig functions, [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/23/navigation-minimum/">Navigation with only addition, subtraction, and tables</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>In the novel <em>Carry On, Mr. Bowditch</em>, a sailor asked Nathaniel Bowditch to teach him how to do navigational calculations, but the man only knows how to add and subtract by counting on his fingers. He had not heard of multiplication. Bowditch is surprised, but realizes if he made tables of logs of trig functions, then it would be possible for barely numerate sailors to calculate their position.</p>
<p><em>Carry On, Mr. Bowditch</em> is fiction, but it&#8217;s essentially factual, and so I imagine something like the conversation above did happen. I thought about how this might work, and it would be difficult.</p>
<p>It is true that if someone can add, subtract, and look-up numbers in a table of logarithms, they can effectively multiply. To find the product <em>xy</em>, they would look up the logarithms of <em>x</em> and <em>y</em>, add the results, then use the table in reverse to find what number has a logarithm equal to the sum.</p>
<h2>Difficulties</h2>
<p>But there are a couple difficulties in this imagined scheme. First, every calculation would require a lot of steps, including looking up numbers in multiple tables. It&#8217;s likely someone who cannot multiply also cannot read, so writing down instructions might not be viable. Second, carrying out calculations using tables is usually not simply a matter of looking up numbers; there are other things someone would need to know, such as <a href="https://www.johndcook.com/blog/2024/06/25/trig-tables/">interpolation and range reduction</a>. </p>
<p>Bowditch was trying to train innumerate sailors to do specific calculations, not general mathematics, and so there would be ways to mitigate the problems above. Maybe he could create diagrams that would allow a semi-literate person to carry out an algorithm. The sailor wouldn&#8217;t need to be able to read <em>per se</em>. The instructions could be aids to help him recall memorized steps. The specialized nature of the calculations might also eliminate the need for range reduction and interpolation.</p>
<h2>Tables</h2>
<p>How many tables would be necessary? Someone who understands trigonometry doesn&#8217;t need separate tables for sine and cosines. And they wouldn&#8217;t need a table with entries for all angles. A table of sines for angles between 0 and 45° would be enough. But someone who doesn&#8217;t know multiplication would need more tables and bigger tables. Or they would need instruction in how to get by with less. It would be an interesting trade-off.</p>
<p>The novel mentioned tabulating logs of trig functions. For example, if you need to calculate</p>
<p style="padding-left: 40px;">cos(<em>a</em>) cos(<em>b</em>)</p>
<p>it would be convenient to be able to look up log(cos(<em>a</em>)) and log(cos(<em>b</em>)) rather than look up the cosines and then look up their logs. But you&#8217;d still need to be able to convert</p>
<p style="padding-left: 40px;">log( cos(<em>a</em>) cos(<em>b</em>) )</p>
<p>into</p>
<p style="padding-left: 40px;">cos(<em>a</em>) cos(<em>b</em>).</p>
<p>This would require a table of logarithms, if the user is able to infer exponentials by reading a table in reverse. Otherwise you&#8217;d need a table of exponentials.</p>
<p>In general, you can assume less sophistication from a user by increasing the number of tables. But this also complicates the instructions the user must follow.</p>
<p>To give a specific example, suppose a sailor wanted to calculate his position using the <a href="https://www.johndcook.com/blog/2026/09/21/haversine-law/">law of haversines</a>:</p>
<p style="padding-left: 40px;">hav(<em>c</em>) = hav(<em>a</em> − <em>b</em>) + sin(<em>a</em>) sin(<em>b</em>) hav(<em>C</em>).</p>
<p>A mathematically sophisticated sailor would only need a table of sines to infer <em>c</em> from <em>a</em>, <em>b</em>, and <em>C</em>. He could calculate haversine via</p>
<p style="padding-left: 40px;">hav(θ) = sin²(θ/2),</p>
<p>though inverting hav(<em>c</em>) to solve for <em>c</em> would require calculating a square root, either directly or via a table.</p>
<p>If one were to minimize the amount of sophistication needed by maximizing the use of tables, the algorithm for finding <em>c</em> would be</p>
<ol>
<li>Subtract <em>b</em> from <em>a</em> and look up the haversine of the difference.</li>
<li>Look up log(sin(<em>a</em>)) and log(sin(<em>b</em>)) from one table and log(hav(<em>C</em>)) from another and add the results.</li>
<li>Take the exponential of the result in the previous step using a table of exponentials.</li>
<li>Add the results of steps 1 and 3, and look up the result in a table of inverse haversine values.</li>
</ol>
<p>This would require five tables: sine, log sine, log haversine, exponential, and inverse haversine.</p>
<h2>Condescension</h2>
<p>Bowditch&#8217;s effort to make navigation accessible to the uneducated is an example of condescension in its literal and positive sense. If we say a person is condescending, we imagine an arrogant person who belittles those around him. But condescension literally means coming down to be with someone. Theologians use the word to describe the incarnation of Christ.</p>
<p>Like all scholars, Bowditch wrote for his peers, notably in his English edition of Laplace&#8217;s magnum opus on celestial mechanics. But unlike most scholars, he also devoted years of his life to a making knowledge accessible to uneducated men, culminating in his book <em>The New American Practical Navigator</em>, still in print <a href="https://msi.nga.mil/Publications/APN">here</a> [1].</p>
<h2>Related posts</h2>
<ul>
<li class='link'><a href='https://www.johndcook.com/blog/2024/06/03/using-a-table-of-logarithms/'>Using a table of logarithms</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2024/06/25/trig-tables/'>Using a table of trig functions</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2022/04/10/why-a-slide-rule-works/'>Why a slide rule works</a></li>
</ul>
<p>[1] The book has been updated over the last couple centuries. Obviously the section on GPS, for example, does not date back to Bowditch.</p>The post <a href="https://www.johndcook.com/blog/2026/09/23/navigation-minimum/">Navigation with only addition, subtraction, and tables</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Nathaniel Bowditch</title>
		<link>https://www.johndcook.com/blog/2026/09/22/nathaniel-bowditch/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:01:30 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<category><![CDATA[History]]></category>
		<category><![CDATA[Navigation]]></category>
		<category><![CDATA[Orbital mechanics]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247946</guid>

					<description><![CDATA[<p>A couple days ago a friend told me about the book Carry On, Mr. Bowditch, a fictional account of the life of Nathaniel Bowditch (1773–1838). I&#8217;ve been listening to the book on Audible, and apparently it&#8217;s only lightly fictionalized. Bowditch was a self-educated mathematician and astronomer, best known for his book The American Practical Navigator, [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/22/nathaniel-bowditch/">Nathaniel Bowditch</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>A couple days ago a friend told me about the book <em>Carry On, Mr. Bowditch</em>, a fictional account of the life of Nathaniel Bowditch (1773–1838). I&#8217;ve been listening to the book on Audible, and apparently it&#8217;s only lightly fictionalized.</p>
<p>Bowditch was a self-educated mathematician and astronomer, best known for his book <em>The American Practical Navigator</em>, first published in 1801. The book has been continually revised over the last two centuries and is still in print, available for <a href="https://msi.nga.mil/Publications/APN">download</a> from the National Geospatial-Intelligence Agency. The latest edition begins with a brief account of Bowditch&#8217;s life, confirming the essential details of the fictional biography.</p>
<p>Two things stand out about Bowditch: his attention to detail and his desire to make ideas accessible. He taught himself Latin in order to read Newton&#8217;s <em>Principia</em> and followed the text so closely that he found a number of errors.</p>
<p>Bowditch&#8217;s navigation book grew out of the numerous corrections he made to error he found in John Hamilton Moore’s <em>The Practical Navigator</em>, the leading navigation text of the time.</p>
<p>At the beginning of the 19th century it was theoretically possible to determine time, and hence longitude, from lunar observation. However, the method required ideal observation conditions and laborious calculation. Bowditch developed a way to make the necessary measurements under more general conditions, and simplified the necessary calculations. According to the biographical preface mentioned above,</p>
<blockquote><p>Bowditch vowed while writing this edition [of his navigation text] to “put down in the book nothing I can’t teach the crew,” and it is said that every member of his crew including the cook could take a lunar observation and plot the ship’s position.</p></blockquote>
<p>After completing <em>The American Practical Navigator</em>, Bowditch began an English translation of Pierre Laplace&#8217;s encyclopedic <em lang="fr">Mecanique Celeste</em>, filling in details to make the work accessible to a wider audience. He was able to translate four out of the five volumes by the end of his life.</p>
<h2>Related posts</h2>
<ul>
<li class='link'><a href='https://www.johndcook.com/blog/2025/12/04/the-navigational-triangle/'>The navigational triangle</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2025/12/04/line-of-position/'>Line of position</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2026/04/14/intersecting-spheres-and-gps/'>Intersecting spheres and GPS</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/22/nathaniel-bowditch/">Nathaniel Bowditch</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Haversine law</title>
		<link>https://www.johndcook.com/blog/2026/09/21/haversine-law/</link>
					<comments>https://www.johndcook.com/blog/2026/09/21/haversine-law/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 23:13:22 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<category><![CDATA[Navigation]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247940</guid>

					<description><![CDATA[<p>Suppose you want to solve a triangle. You know two sides and the angle between them. Then you can solve for the third side using the law of cosines. Now suppose you want to solve a big triangle, a triangle on the surface of the earth so large that the curvature of the earth matters. You [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/21/haversine-law/">Haversine law</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Suppose you want to solve a triangle. You know two sides and the angle between them. Then you can solve for the third side using the law of cosines.</p>
<p>Now suppose you want to solve a <strong>big</strong> triangle, a triangle on the surface of the earth so large that the curvature of the earth matters. You can still use the law of cosines, but you&#8217;ll need <a href="https://www.johndcook.com/blog/2022/08/24/law-of-cosines-on-a-sphere/">the spherical law of cosines</a>:</p>
<p style="padding-left: 40px;">cos(<em>c</em>) = cos(<em>a</em>) cos(<em>b</em>) + sin(<em>a</em>) sin(<em>b</em>) cos(<em>C</em>).</p>
<p>If you know the (angular) lengths of sides <em>a</em> and <em>b</em>, and (tangential) angle <em>C</em> between the two sides, you can solve for <em>c</em> by taking the inverse cosine of the right hand side above.</p>
<p>Now suppose you want to solve this big triangle because you&#8217;re a <strong>navigator</strong> on a ship a couple centuries ago, doing calculations by looking up trig functions and inverse trig functions in a table. You&#8217;re interested in triangles that are so big that you have to account for the fact that you&#8217;re living on a sphere. But at the same time, your triangles are still fairly small relative to the size of the globe.</p>
<h2>The problem with the law of cosines</h2>
<p>The numbers <em>a</em> and <em>b</em> will often be fairly small, and so their cosines will be near 1 and their sines are near zero. So the calculation</p>
<p style="padding-left: 40px;">cos(<em>a</em>) cos(<em>b</em>) + sin(<em>a</em>) sin(<em>b</em>) cos(<em>C</em>)</p>
<p>will add a number near 1 and a number near zero. That&#8217;s a problem.</p>
<p>Say you&#8217;re working with five decimal place arithmetic. Then if the second term above is less than 10<sup>−5</sup>, its contribution to the sum gets completely lost in the addition to the first term. If the second term is larger than 10<sup>−5</sup> but still small, its contribution to the sum will be partially lost.</p>
<h2>Law of haversines</h2>
<p>Enter the <a href="https://www.johndcook.com/blog/2009/09/25/how-many-trig-functions/">haversine</a>, defined by</p>
<p style="padding-left: 40px;">hav(θ) = (1 − cos(θ))/2.</p>
<p>The expression 1 − cos θ was called the versine, and so half of the versine is the haversine.</p>
<p>In terms of the haversine, the law of cosines above becomes the law of haversines:</p>
<p style="padding-left: 40px;">hav(<em>c</em>) = hav(<em>a</em> − <em>b</em>) + sin(<em>a</em>) sin(<em>b</em>) hav(<em>C</em>).</p>
<p>Now suppose you have a table of haversines and inverse haversines. The law of haversines requires a little less work: you have one less table lookup, and you trade a product for a subtraction.</p>
<p>But the primary advantage is numerical accuracy: the terms on the right side have roughly the same size.</p>
<h2>Tables</h2>
<p>Note that we&#8217;re assuming the values in your table of haversines have been calculated correctly to the given precision. If you calculated your own values of haversines from the definition above, you&#8217;d lose precision in the subtraction 1 − cos θ, defeating the advantage of the law of haversines [1].</p>
<h2>History</h2>
<p>According to <a href="https://en.wikipedia.org/wiki/Haversine_formula">Wikipedia</a>.</p>
<blockquote><p>The first table of haversines in English was published by James Andrew in 1805, but Florian Cajori credits an earlier use by José de Mendoza y Ríos in 1801. The term <em>haversine</em> was coined in 1835 by James Inman.</p></blockquote>
<h2>Experiments</h2>
<p>I ran some experiments that carried out arithmetic in float16 (11 bits of precision) to approximate what someone might have done by hand. When the difference between <em>a</em> and <em>b</em> was on the order of 1° or 0.1°, the law of cosine method often overflowed: the right-hand side evaluated to something larger than 1 even though theoretically it should be less than 1. The haversine method never overflowed.</p>
<p>The median error for the haversine method was a couple orders of magnitude less than that of the cosine method.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2025/12/01/lewis-clark-geolocation/">Lews &amp; Clark navigation</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2009/09/25/how-many-trig-functions/">How many trig functions are there?</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/11/08/heron-on-a-sphere/">Analog of Heron&#8217;s formula for a sphere</a></li>
</ul>
<p>[1] hav(θ) = (1 − cos(θ))/2 = sin²(θ/2). If you calculated hav θ by looking up sin(θ/2) and squaring it, you&#8217;d be doing extra work, but you wouldn&#8217;t have numerical problems.</p>The post <a href="https://www.johndcook.com/blog/2026/09/21/haversine-law/">Haversine law</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/21/haversine-law/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Why fitting a logistic is nearly impossible from early data</title>
		<link>https://www.johndcook.com/blog/2026/09/18/logistic-fit-sensitivity/</link>
					<comments>https://www.johndcook.com/blog/2026/09/18/logistic-fit-sensitivity/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 01:17:36 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Probability and Statistics]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247934</guid>

					<description><![CDATA[<p>Nothing grows exponentially forever. What appears to be an exponential curve often turns out to be some sort of S curve, such as a logistic curve. Suppose you&#8217;re collecting data on the left side of the curve. If there&#8217;s even a small amount of error in your data, you won&#8217;t be able to predict the [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/18/logistic-fit-sensitivity/">Why fitting a logistic is nearly impossible from early data</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Nothing grows exponentially forever. What appears to be an exponential curve often turns out to be some sort of S curve, such as a logistic curve.<img loading="lazy" decoding="async" class="aligncenter" src="https://www.johndcook.com/logistic_tangents.png" alt="logistic curve with extrapolations" width="480" height="360" /></p>
<p>Suppose you&#8217;re collecting data on the left side of the curve. If there&#8217;s even a small amount of error in your data, you won&#8217;t be able to predict the asymptotic value with any accuracy. But if you have data on both sides of the inflection point, you can make a good prediction of the limiting value.</p>
<p>I&#8217;ve written about this <a href="https://www.johndcook.com/blog/2025/12/20/fit-logistic-curve/">before</a>, explaining that the problem is hard, but I didn&#8217;t say <em>why</em> it&#8217;s hard. Here I&#8217;d like to give an idea why it&#8217;s hard.</p>
<p>Suppose you want to fit a logistic equation</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/logistic_fit1.svg" alt="y(t) = \frac{L}{1 + \exp(-k(t - t_0))}" width="197" height="43" /></p>
<p>to three distinct values of <em>t</em> and the corresponding values of <em>y</em>. There is a unique solution, but in general you cannot find a solution in closed form. However, if the values of <em>t</em> are evenly spaced</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/logistic_fit2.svg" alt="y_2 - y_1 = y_1 - y_0 = h" width="154" height="20" /></p>
<p>there is a method [1] to solve for the parameters <em>L</em>, <em>k</em>, and <em>t</em><sub>0</sub>. For this post we&#8217;re only interested in the limiting value <em>L</em>, and it can be found by</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/logistic_fit3.svg" alt="L = \frac{y_1^2(y_0 + y_2) - 2y_0 y_1 y_2}{y_1^2 - y_0 y_2}" width="193" height="54" /></p>
<p>independent of <em>h</em>.</p>
<p>To find out how small changes in the <em>y</em>&#8216;s change the estimate of <em>L</em>, we take the partial derivatives of <em>L</em> with respect to the <em>y</em>&#8216;s and find</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/logistic_fit4.svg" alt="\frac{\partial L}{\partial y_0} = \frac{\partial L}{\partial y_2} = \frac{y_1^2\, h^2}{\left(y_1^2 - y_0 y_2\right)^2} " width="191" height="68" /></p>
<p>and</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/logistic_fit5.svg" alt="\frac{\partial L}{\partial y_1} = \frac{-2\, y_0 y_2\, h^2}{\left(y_1^2 - y_0 y_2\right)^2}" width="143" height="66" /><br />
All three derivatives have the same expression in the denominator: <em>y</em><sub>1</sub>² − <em>y</em><sub>0</sub> <em>y</em><sub>2</sub>.</p>
<p>If the function <em>y</em>(<em>t</em>) were an exponential, this expression would be exactly zero [2]. The function <em>y</em>(<em>t</em>) is not exactly exponential, but it is <em>approximately</em> exponential when the <em>t</em>&#8216;s are in the left or right tail of the logistic curve. The further out in either tail the <em>t</em>&#8216;s are, the closer the expression is to zero.</p>
<p>So when all the <em>t</em>&#8216;s come from the same side of the inflection point, <em>y</em>(<em>t</em>) is nearly exponential the partial derivatives are huge and so the fitted value of <em>L</em> is extremely sensitive to changes in the <em>y</em>&#8216;s.</p>
<p>As a concrete example, set <em>L</em> = <em>k</em> = 1 and <em>t</em><sub>0</sub> = 0. Evaluate <em>y</em>(<em>t</em>) at −2, −1.5, and −1. Then the values of <em>y</em> are</p>
<p style="padding-left: 40px;"><em>y</em><sub>0</sub> = 0.11920292<br />
<em>y</em><sub>1</sub> = 0.18242552<br />
<em>y</em><sub>2</sub> = 0.26894142</p>
<p>If you forecast <em>L</em> using exactly these three values you&#8217;ll get <em>L</em> = 1.</p>
<p>But if you change <em>y</em><sub>0</sub> to 0.12374097, the forecasted value of <em>L</em> is infinite. Values of <em>y</em><sub>0</sub> in the interval [0.11920292, 0.12374097] predict values of <em>K</em> in [1, ∞].</p>
<p>[1] Raymond Pearl and Lowell J. Reed. On the Rate of Growth of the Population of the United States Since 1790 and its Mathematical Representation. Proceedings of the National Academy of Sciences of the United States of America, Vol. 6, No. 6 (Jun. 15, 1920), pp. 275-288</p>
<p>[2] exp(<em>x</em> + <em>h</em>)² = exp(<em>x</em>)² exp(<em>h</em>)² = exp(<em>x</em>) exp(<em>x</em> + 2<em>h</em>)</p>
<p>&nbsp;</p>The post <a href="https://www.johndcook.com/blog/2026/09/18/logistic-fit-sensitivity/">Why fitting a logistic is nearly impossible from early data</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/18/logistic-fit-sensitivity/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Empirical fractal</title>
		<link>https://www.johndcook.com/blog/2026/09/17/empirical-fractal/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 16:48:14 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247924</guid>

					<description><![CDATA[<p>There&#8217;s a common saying in discussion of fractals that the length of a coastline depends on how small a device you use to measure it. I thought this was a hypothetical, say as applied to the steps in the construction of the Koch snowflake. But the saying has its roots in actually surveying. Lewis Fry [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/17/empirical-fractal/">Empirical fractal</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>There&#8217;s a common saying in discussion of fractals that the length of a coastline depends on how small a device you use to measure it. I thought this was a hypothetical, say as applied to the steps in the construction of the Koch snowflake. But the saying has its roots in actually surveying.</p>
<p>Lewis Fry Richardson (1881–1953) noticed that the length of the coast of Scotland depended on the size of segments used to measure it. More specifically, he found that the length followed a power law, i.e. that there&#8217;s a linear relation between the log of the coastline length and the log of the ruler length.</p>
<p>Here&#8217;s a reproduction of Richardson&#8217;s plot, taken from [1].</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/scottish_coast.png" width="500" height="351" /></p>
<p>Mandelbrot built on Richardson&#8217;s observation and defined the idea of fractal dimension.</p>
<p>I was under the impression that fractals were invented as mathematical novelties that researchers later found applications for. But as is often the case, the applications came first. Or at least <em>some</em> applications came first.</p>
<p>Ideally there&#8217;s always a feedback cycle where applications lead to theory and theory leads to applications. As Donald Knuth put it, &#8220;The best theory is inspired by practice. The best practice is inspired by theory.&#8221;</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2025/09/04/minimalist-mandelbrot-set/">Minimalist Mandelbrot set</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2021/07/11/fractal-brownian-motion/">The fractal nature of Brownian motion</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/08/16/randomly-generated-dragon/">Randomly generated dragon</a></li>
</ul>
<p>[1] Eoghan Bradley and Mark McCartney. Four hundred years of the fractal coastline of Scotland. The Mathematical Gazette, November 2019, Vol. 103, No. 558 (November 2019), pp. 518-521</p>The post <a href="https://www.johndcook.com/blog/2026/09/17/empirical-fractal/">Empirical fractal</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Phone words</title>
		<link>https://www.johndcook.com/blog/2026/09/17/phone-words/</link>
					<comments>https://www.johndcook.com/blog/2026/09/17/phone-words/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 14:03:51 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247915</guid>

					<description><![CDATA[<p>I recently bought a copy of Los Alamos Rolodex, a book displaying business cards from Los Alamos Nation Labs from 1967 to 1978. You can find some examples of the cards here. One of the cards in the book is for Eugene Frank, President of B &#38; F Instruments. His card lists his phone number [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/17/phone-words/">Phone words</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>I recently bought a copy of Los Alamos Rolodex, a book displaying business cards from Los Alamos Nation Labs from 1967 to 1978. You can find some examples of the cards <a href="https://clui.org/collections/los-alamos-business-cards/selection-cards">here</a>.</p>
<p>One of the cards in the book is for Eugene Frank, President of B &amp; F Instruments. His card lists his phone number as</p>
<p style="padding-left: 40px;">(215) MErcury 9-7100</p>
<p>At first glance I thought the &#8220;E&#8221; in &#8220;MErcury&#8221; had been accidentally capitalized. Then I realized the intention was that someone would dial ME (i.e. 63) and ignore &#8220;rcury&#8221;. So the phone number would be (215) 639-7100.</p>
<p>This card was from 1968, the height of the space race. Maybe the card was alluding to the Project Mercury or the planet Mercury, or both. [1]</p>
<p>The telephone keypad mapping (ITU E.161 standard) is a poor attempt at making phone numbers more memorable. For starters, there&#8217;s no way to encode 0 or 1 [2]. It&#8217;s unlikely a phone number will correspond to anything memorable unless you come up with the word first and then try to obtain the phone number, such as 800 FLOWERS.</p>
<p>Inserting extra letters, as Mr. Frank did, greatly increases the chances of encoding a phone number as a word. But then you need to denote which letters count and which ones are filler, so there&#8217;s not much advantage. Still, I wanted to play around with it for fun. I found 109 words [3] containing the letters from a telephone encoding of 4228646. (I&#8217;m using the file <code>/usr/share/dict/words</code> on my laptop as my list of words.)</p>
<p>Here are some of the more interesting hits.</p>
<ul>
<li>semicatholicism</li>
<li>heartburning</li>
<li>gladiatorism</li>
<li>diabetogenic</li>
<li>galactogenetic</li>
<li>xanthocreatinine</li>
</ul>
<p>There are over 30,000 words containing an encoding of the area code 832. One of these is <em>traditional</em>, and so I could write my phone number as</p>
<p style="padding-left: 40px;"><code>TraDitionAl semICAThOlIcisM</code>.</p>
<p>Another choice for 832 is <em>intercosmic</em>, so</p>
<p style="padding-left: 40px;"><code>inTErCosmic GAlaCTOGeNetic</code></p>
<p>is another possibility.</p>
<p><em>Galactogentic</em> can refer to the production of milk by the mammary glands or to the formation of galaxies (e.g. the Milky Way). Here <em>intercosmic</em> fits with the later sense.</p>
<p>I got greedy and tried to find a word containing the full phone number, 8324228646, but didn&#8217;t find anything.</p>
<p>Here&#8217;s my business card in the style of the Los Alamos Rolodex cards, created by Grok, using (832) GlAdiATOrIsM as the phone number.</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/vintage_card.png" width="600" height="366" /></p>
<p>Now suppose you remembered &#8220;gladiatorism&#8221; but not which letters were capitalized. Then you&#8217;d have to try up to 792, i.e. 12 choose 7, possible numbers, so this really isn&#8217;t a practical mnemonic. If you remembered &#8220;traditional semicatholicism&#8221; without capitalization it would be worse, with over a million possibilities (11 choose 3 times 15 choose 7). Some possibilities are counted twice, since different ways of selecting letters can lead to the same phone number, but still there are too many possibilities to try.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2021/07/26/major-memory-keypad/">Major memory system telephone keypad</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2023/11/17/phone-number-intel/">What can you learn from a phone number?</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2022/03/14/phone-tones-in-musical-notation/">Phone tones inn musical notation</a></li>
</ul>
<p>[1] Thanks to Andrew for pointing out in his comment that it was common at one time to encode the first two numbers of the exchange (the second triplet of numbers in a phone number) as letters, and assign a word to those letters. Sometimes this was standardized, such as Pennsylvania 6 for 736, an example made famous by Glenn Miller. But from what I can tell, not all exchanges had standard names, and proposed standards weren&#8217;t always adopted in practice.</p>
<p>In the example above, I don&#8217;t know whether it was common to encode 639 as Mercury 9, or even ME 9, or whether Mr. Frank chose this. It was common chose <em>some</em> encoding for the first two numbers of the exchange, though that practice was going away by 1968. Perhaps Mr. Frank was an older man who retained a habit he acquired when it was more common. None of the other cards in the book spelled out the exchange.</p>
<p><strong>Update</strong>: Thanks to Chuck for pointing out this <a href="https://en.wikipedia.org/wiki/Telephone_exchange_names#Standardization">list</a> of recommended words for exchanges. Note that there are multiple suggestions for most exchanges, including six for 63X.</p>
<p>[2] Not only are there no letters for 0 and 1, the letters O and I represent digits. At one point in time the first digit of an exchange (the middle three digits) could not be a 0 or 1, but these digits could appear anywhere else.</p>
<p>[3] I initially found a list of 185 words, but some of these were duplicates: a word can represent a phone number in more than one way.</p>The post <a href="https://www.johndcook.com/blog/2026/09/17/phone-words/">Phone words</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/17/phone-words/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
			</item>
		<item>
		<title>Converting between cosine similarity and concentration ratio</title>
		<link>https://www.johndcook.com/blog/2026/09/16/concentration-ratio/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 16:05:11 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<category><![CDATA[Machine learning]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247909</guid>

					<description><![CDATA[<p>I&#8217;ve written three posts on cosine similarity lately. The first looked at interpreting cosine similarity. The second looked at an approximation related to the first. The third looked at how ranking according to cosine similarity works better than cosine similarity itself. Normalized word vectors are points on a high dimensional sphere, and geometry in high [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/16/concentration-ratio/">Converting between cosine similarity and concentration ratio</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>I&#8217;ve written three posts on cosine similarity lately. The <a href="https://www.johndcook.com/blog/2026/09/15/cosine-similarity/">first</a> looked at interpreting cosine similarity. The <a href="https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/">second</a> looked at an approximation related to the first. The <a href="https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/">third</a> looked at how ranking according to cosine similarity works better than cosine similarity itself.</p>
<p>Normalized word vectors are points on a high dimensional sphere, and geometry in high dimensions is counterintuitive. See the first post in this series for an explanation.</p>
<p>The set of points within a given angular distance of a point on a hypersphere is called a <a href="https://www.johndcook.com/blog/2023/08/09/hypersphere-cap/">spherical cap</a>. The ratio of the area of this spherical cap to that of the whole sphere is called <strong>cap fraction</strong> or <strong>concentration ratio</strong>. Concentration ratio explains why a modest cosine similarity value corresponds to a tiny portion of the area of the sphere and should be interpreted as a close match.</p>
<p>For this post, I wanted to share a plot of concentration ratio as a function of cosine similarity.</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/concentration_ratio.png" width="480" height="360" /></p>
<p>This shows that moderate values of cosine similarity correspond to infinitesimal concentration ratios. And yet, as the third post linked at the top showed, word vectors are very unevenly distributed, and even extremely small regions of the sphere can contain multiple word vectors.</p>
<p>I only included cosine similarity values up to 0.8 because the function plotted above takes a nosedive for larger values, even on a logarithmic scale.</p>
<p>Here&#8217;s the Python code to make the plot, using the function <code>cap_fraction</code> from <a href="https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/">here</a>.</p>
<pre>
s = np.linspace(0, 0.8, 500)
plt.plot(s, cap_fraction(np.acos(s), 200))
plt.yscale("log")
plt.xlabel("cosine similarity")
plt.ylabel("concentration ratio")
plt.show()
</pre>The post <a href="https://www.johndcook.com/blog/2026/09/16/concentration-ratio/">Converting between cosine similarity and concentration ratio</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Coffee + milk ≠ latte</title>
		<link>https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/</link>
					<comments>https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 15:06:58 +0000</pubDate>
				<category><![CDATA[Statistics]]></category>
		<category><![CDATA[Machine learning]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247906</guid>

					<description><![CDATA[<p>Yesterday I wrote about the canonical example of how vector embeddings of words add: “king” − “man” + “woman” ≈ “queen” This should be interpreted as saying that the word vector for king, minus the word vector for man, plus the word vector for woman, is in some sense close to the word vector for queen. This post will [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/">Coffee + milk ≠ latte</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p><a href="https://www.johndcook.com/blog/2026/09/15/cosine-similarity/">Yesterday</a> I wrote about the canonical example of how vector embeddings of words add:</p>
<p style="padding-left: 40px;">“king” − “man” + “woman” ≈ “queen”</p>
<p>This should be interpreted as saying that the word vector for <em>king</em>, minus the word vector for <em>man</em>, plus the word vector for <em>woman</em>, is in some sense close to the word vector for <em>queen</em>.</p>
<p>This post will look at another example. Is the expression</p>
<p style="padding-left: 40px;">&#8220;coffee&#8221; + &#8220;milk&#8221; ≈ &#8220;latte&#8221;</p>
<p>true in some sense?</p>
<h2>Notation</h2>
<p>In this post I will use &#8220;foo&#8221; to mean the vector embedding of the word <em>foo</em>.</p>
<h2>Coffee + milk</h2>
<p>The cosine similarity between &#8220;coffee&#8221; + &#8220;milk&#8221; and &#8220;latte&#8221; is about 0.63. And for reasons given in the previous post, this is a large value of cosine similarity. But there are 11 words that are more similar to &#8220;milk&#8221; + &#8220;coffee&#8221; than &#8220;latte&#8221;. Here are the top 12 matches in order.</p>
<ol>
<li>coffee</li>
<li>milk</li>
<li>tea</li>
<li>drink</li>
<li>chocolate</li>
<li>cream</li>
<li>breakfast</li>
<li>ice</li>
<li>beer</li>
<li>vanilla</li>
<li>starbucks</li>
<li>latte</li>
</ol>
<p>There are two questions to resolve. First, why isn&#8217;t <em>latte</em> one of the closest words? Second, why is the cosine similarity large even though <em>latte</em> is not one of the best matches?</p>
<h2>Concept arithmetic</h2>
<p>When word vector arithmetic works, as in the king and queen example, the vectors combine <em>concepts</em>. If you replace the male gender component of <em>king</em> with a female component, you get a vector close to the vector for <em>queen</em>.</p>
<p>But when you add the vectors for <em>milk</em> and <em>coffee</em>, you&#8217;re not adding concepts, you&#8217;re adding ingredients.</p>
<p>The concepts of <em>milk</em> and <em>coffee</em> are similar in that they&#8217;re both common beverages, as are tea and even beer. A latte is a beverage, but it&#8217;s not as common as milk, coffee, tea, or beer.</p>
<h2>Extremely uneven distribution</h2>
<p>If you divide word vectors by their norm, you get a point on a high-dimensional sphere. In the case of the glove-twitter-200 vector embedding, you get a point on a sphere in 200 dimensions. As explained in the earlier post, a fairly large cosine similarity corresponds to a tiny portion of the sphere&#8217;s surface area.</p>
<p>In the example of “king” − “man” + “woman”, the vector &#8220;queen&#8221; is the closest match (except for &#8220;king&#8221; itself).</p>
<p>But there are a lot of words whose vectors are within a tiny region around &#8220;coffee&#8221; + &#8220;milk&#8221;. And by tiny, I mean a region that accounts for a proportion of the sphere on the order of 10<sup>−23</sup>.</p>
<p>The glove-twitter-200 vector list contains vectors for 1.2 million words. If these vectors were roughly evenly distributed on the sphere when normalized, you&#8217;d expect each patch representing 10<sup>−6</sup> of the sphere to contain about a word or two. You wouldn&#8217;t expect a patch taking up 10<sup>−12 </sup>of the sphere to contain more than one word, and you certainly wouldn&#8217;t expect a patch taking up 10<sup>−23 </sup>of the sphere to contain 12 words [1].</p>
<h2>Rank order</h2>
<p>Rank order based on cosine similarity is more robust than cosine similarity itself. This is an example of a phenomenon that occurs regularly: a metric whose values are dubious might still rank things well. Naive Bayes is another example. It naively computes probabilities in a way that is blatantly wrong, and yet ranking things by these spurious probabilities works well in some cases.</p>
<p>The cosine similarity between “king” − “man” + “woman” and &#8220;queen&#8221; is roughly the same as the cosine similarity between &#8220;coffee&#8221; + &#8220;milk&#8221; and &#8220;latte.&#8221; But in the former example, rank order picks out <em>queen</em> as the best match; rank order works like you&#8217;d expect, because you&#8217;re working with attributes that can be decomposed.</p>
<h2>Dog + infant = puppy?</h2>
<p>I wouldn&#8217;t be surprised if the Anglo-Saxon word for <em>puppy</em> was something like <em>dogchild</em>. The language was full of colorful compound words, such as <em>hronrad</em> (&#8220;whale-road&#8221;) for the sea and <em>nosethyrl</em> (&#8220;nose-hole&#8221;) for nostril.</p>
<p>Here are the top ten matches for &#8220;dog&#8221; + &#8220;infant&#8221; along with their cosine similarities.</p>
<ol>
<li>dog, 0.819</li>
<li>infant, 0.809</li>
<li>toddler, 0.734</li>
<li>dogs, 0.697</li>
<li>puppy, 0.688</li>
<li>cat, 0.682</li>
<li>pet, 0.676</li>
<li>child, 0.671</li>
<li>newborn, 0.670</li>
<li>baby, 0.650</li>
</ol>
<p>This shows that &#8220;puppy&#8221; is close to &#8220;dog&#8221; + &#8220;infant&#8221;, both in terms of cosine similarity and rank order, though it&#8217;s not the closet.</p>
<p>This also shows that you have to take the addition of word vectors with a grain of salt. It&#8217;s no surprise that <em>puppy</em> was a good match, but it&#8217;s surprising that <em>cat</em> is nearly as good.</p>
<p>[1] I poked around a little to get an idea just how unevenly words are distributed. The closest pair of words is <em>jajaja</em> and <em>jajajaja</em> with a cosine similarity of 0.993. The most isolated word, meaning the word whose nearest neighbor is furthest away, the the Thai word <span lang="th">เคยไหม</span>. It&#8217;s nearest neighbor is the Russian word <span lang="ru">боль</span> with a cosine similarity of 0.283.</p>
<p>The glove-twitter-200 vectors were created from a corpus that is about half English and about other languages and strings of symbols that are not words in any language. Presumably <span lang="th">เคยไหม</span> would have a much closer neighbor in a corpus containing more Thai words.</p>
<p>I didn&#8217;t search the entire corpus, only the 50,000 most frequently occurring vectors, because a full search would require running an <em>O</em>(<em>N</em>²) search with <em>N</em> = 1,200,000.</p>The post <a href="https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/">Coffee + milk ≠ latte</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>Fibonacci product</title>
		<link>https://www.johndcook.com/blog/2026/09/16/fibonacci-product/</link>
					<comments>https://www.johndcook.com/blog/2026/09/16/fibonacci-product/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 12:04:09 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Number theory]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247901</guid>

					<description><![CDATA[<p>The product of four consecutive Fibonacci numbers equals the product of two consecutive integers. For example, 3 × 5 × 8 × 13 = 39 × 40. I ran across this theorem in a note [1] that says &#8220;The product of any four consecutive Fibonacci numbers is twice a triangular number.&#8221; Since triangular numbers have [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/16/fibonacci-product/">Fibonacci product</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The product of four consecutive Fibonacci numbers equals the product of two consecutive integers.</p>
<p>For example,</p>
<p style="padding-left: 40px;">3 × 5 × 8 × 13 = 39 × 40.</p>
<p>I ran across this theorem in a note [1] that says &#8220;The product of any four consecutive Fibonacci numbers is twice a triangular number.&#8221; Since triangular numbers have the form <em>n</em>(<em>n</em> + 1)/2, twice a triangular number is the product of two consecutive integers.</p>
<p>The note also gives a way to find the numbers on the right hand side. We have</p>
<p style="padding-left: 40px;"><em>F</em><sub><em>n</em></sub> <em>F</em><sub><em>n</em>+1</sub> <em>F</em><sub><em>n</em>+2</sub> <em>F</em><sub><em>n</em>+3</sub> = <em>m</em>(<em>m</em> + 1)</p>
<p>where <em>m</em> equals</p>
<p style="padding-left: 40px;"><em>F</em><sub><em>n</em>+1</sub> <em>F</em><sub><em>n</em>+2</sub></p>
<p>if <em>n</em> is odd and</p>
<p style="padding-left: 40px;"><em>F</em><sub><em>n</em></sub> <em>F</em><sub><em>n</em>+3</sub></p>
<p>if <em>n</em> is even.</p>
<p>In the example at the top, 3 is the 4th Fibonacci number, so <em>n</em> = 4. Since 4 is even, <em>m</em> is the product of the 4th and 7th Fibonacci numbers, i.e. <em>m</em> = 3 × 13 = 39.</p>
<h2>More Fibonacci posts</h2>
<ul>
<li class='link'><a href='https://www.johndcook.com/blog/2018/07/13/fibonacci-meets-pythagoras/'>Fibonacci meets Pythagoras</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2026/02/05/fibonacci-certificate/'>Certified Fibonacci numbers</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2025/10/17/trig-fibonacci/'>Turning trig identities into Fibonacci identities</a></li>
</ul>
<p>[1] K. B. Subramaniam. On a link between Triangular and Fibonacci numbers. The Mathematical Gazette, Vol. 103, No. 558 (November 2019), p. 489.</p>The post <a href="https://www.johndcook.com/blog/2026/09/16/fibonacci-product/">Fibonacci product</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/16/fibonacci-product/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Simple approximation for spherical cap area</title>
		<link>https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 22:00:44 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Geometry]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247896</guid>

					<description><![CDATA[<p>The previous post looked at how to interpret cosine similarity, or equivalently angles between word vectors. In a high-dimensional space, randomly chosen vectors are likely nearly perpendicular, and so relatively large angles, such as 50°, indicate very closely related words. Another way to look at this, as explained in the previous post, is that in [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/">Simple approximation for spherical cap area</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The <a href="https://www.johndcook.com/blog/2026/09/15/cosine-similarity/">previous post</a> looked at how to interpret cosine similarity, or equivalently angles between word vectors. In a high-dimensional space, randomly chosen vectors are likely nearly perpendicular, and so relatively large angles, such as 50°, indicate very closely related words.</p>
<p>Another way to look at this, as explained in the previous post, is that in high dimensions, a spherical cap of angular radius θ represents a small portion of a sphere, even for moderately large θ.</p>
<p>The proportion of the area inside the spherical cap, given <a href="https://www.johndcook.com/blog/2023/08/09/hypersphere-cap/">here</a>, involves the &#8220;regularized incomplete beta function&#8221; and so it&#8217;s hard to have an intuition for the value.</p>
<p>For large dimension <em>n</em>, the approximation</p>
<p style="padding-left: 40px;"><em>n</em><sup>−1/2</sup> sin<sup><em>n</em> − 1</sup>(θ)</p>
<p>gives the proportion of the area inside the cap to within an order of magnitude. It&#8217;s easy to see that this function goes to zero quickly as <em>n</em> increases, provided |θ| &lt; π/2.</p>
<p>If you have the cosine similarity <em>c</em> = cos θ rather than θ itself, the approximation becomes</p>
<p style="padding-left: 40px;"><em>n</em><sup>−1/2</sup> (1 − <em>c</em>²)<sup>(<em>n</em> − 1)/2</sup>.</p>
<h2>Python script</h2>
<p>Let&#8217;s try it on the example from the previous post, in which <em>n</em> = 200 and θ = 49°.</p>
<pre>import numpy as np
from scipy.special import betainc

# Fraction of S^{n-1} inside a spherical cap of angular radius theta
# theta is measured from the pole
# Assume 0 &lt; theta &lt; pi/2

def cap_fraction(theta, n):
    x = np.sin(theta) ** 2
    return 0.5 * betainc(0.5 * (n - 1), 0.5, x)

def cap_fraction_approx(theta, n):
    return n**(-0.5) * np.sin(theta)**(n-1)

theta = np.deg2rad(49)
print(cap_fraction(theta, 200)) 
print(cap_fraction_approx(theta, 200)) 
</pre>
<p>This prints 2.03e-26 and 3.37e-26. The order of magnitude is correct as advertised.</p>The post <a href="https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/">Simple approximation for spherical cap area</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>What counts as a large cosine similarity?</title>
		<link>https://www.johndcook.com/blog/2026/09/15/cosine-similarity/</link>
					<comments>https://www.johndcook.com/blog/2026/09/15/cosine-similarity/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 16:06:02 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[Differential geometry]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247894</guid>

					<description><![CDATA[<p>Machine learning represents words as vectors and measures the similarity of words by the angles between the vectors. For vectors x and y, where θ is the angle between the vectors, and so This is the cosine similarity between the words represented by x and y. Small angles have large cosines, and so words with larger cosine similarities [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/15/cosine-similarity/">What counts as a large cosine similarity?</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Machine learning represents words as vectors and measures the similarity of words by the angles between the vectors.</p>
<p>For vectors <strong>x</strong> and <strong>y</strong>,</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" style="background-color: white;" src="https://www.johndcook.com/dotproduct3.svg" alt="\mathbf{x} \cdot \mathbf{y} = ||\mathbf{x} || \,||\mathbf{y} || \, \cos(\theta)" width="173" height="18" /></p>
<p>where θ is the angle between the vectors, and so</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" style="background-color: white;" src="https://www.johndcook.com/dotproduct4.svg" alt="\cos(\theta) = \frac{\mathbf{x} \cdot \mathbf{y}}{ ||\mathbf{x} || \,||\mathbf{y} || }" width="136" height="37" /></p>
<p>This is the cosine similarity between the words represented by <strong>x</strong> and <strong>y</strong>.</p>
<p>Small angles have large cosines, and so words with larger cosine similarities are closer together than words with smaller cosine similarities. The cosine similarity between a word and itself equals 1, and we&#8217;d expect unrelated words to have a cosine similarity near 0.</p>
<p>You can do a sort of arithmetic with vector embeddings of words. The canonical example is that</p>
<p style="padding-left: 40px;">&#8220;king&#8221; − &#8220;man&#8221; + &#8220;woman&#8221; ≈ &#8220;queen&#8221;</p>
<p>Implicit in this equation is that we&#8217;re really adding vector representations of the words. Let <strong>a</strong>, <strong>b</strong>, <strong>c</strong>, and <strong>d</strong> be the vector embeddings of the words <em>king</em>, <em>man</em>, <em>woman</em>, and <em>queen</em>. What we&#8217;re really asserting is that</p>
<p style="padding-left: 40px;"><strong>a</strong> − <strong>b</strong> + <strong>c</strong> ≈ <strong>d</strong>,</p>
<p>except that&#8217;s not true! Or at least it&#8217;s not true unless you view it in the right context.</p>
<p>The angle between <strong>a</strong> − <strong>b</strong> + <strong>c</strong> and <strong>d</strong> is about 49°, which corresponds to a cosine similarity of 0.656. Here I&#8217;m using the gensim glove-twitter-200 embedding that represents words as 200-dimensional vectors.</p>
<p>The way to interpret the equation above is not that a 49° degree angle is approximately 0, or that a similarity of 0.656 is approximately 1.</p>
<p>In high dimensions, such as 200-dimensional word embeddings, nearly all vectors are nearly perpendicular. I wrote a post about this <a href="https://www.johndcook.com/blog/2023/08/09/random-points-hypersphere-orthant/">here</a>. So the angle between randomly selected words will usually be close to 90°, and so in that context an angle of 49° is relatively small. For example, the angle between the vector representations of <em>king</em> and <em>fireplace</em> is 89.25°.</p>
<p>If you divide word vectors by their norm, you can think of each vector as a point on a high-dimensional sphere, in our case a sphere in 200 dimensions. The proportion of vectors within 49° of a given point is surprisingly small in high dimensions.</p>
<p>Let&#8217;s say our point of interest is the north pole of an <em>n</em>-dimensional sphere. We&#8217;d like to calculate the proportion of the area of the sphere that is within an angle θ of the pole. I go through the calculations <a href="https://www.johndcook.com/blog/2023/08/09/hypersphere-cap/">here</a>. (Update: I give an approximation <a href="https://www.johndcook.com/blog/2026/09/15/simple-approximation-for-spherical-cap-area/">here</a> that&#8217;s easier to work with than the exact formula.)</p>
<p>When <em>n</em> = 3, 17% of the area is with 49 degrees of the pole. But when <em>n</em> = 200, the proportion is on the order of 10<sup>−26</sup>, essentially zero.</p>
<p>The vector <strong>d</strong> above representing <em>queen</em> is within a relatively tiny region around the vector <strong>a</strong> − <strong>b</strong> + <strong>c</strong>.</p>
<p>In terms of cosine similarity, 0.656 is a large similarity. Words with a cosine similarity in this range are quite close, even though we wouldn&#8217;t normally think of 0.656 being close to 1. In this context, 0.656 <em>is</em> close to 1.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2023/08/08/angles-between-words/">Angles between words</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2023/08/09/hypersphere-cap/">Area and volume of a hypersphere cap</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2023/08/09/cosine-similarity-not-a-metric/">Cosine similarity does not satisfy the triangle inequality</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/15/cosine-similarity/">What counts as a large cosine similarity?</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/15/cosine-similarity/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Guessing the meaning of a number</title>
		<link>https://www.johndcook.com/blog/2026/09/14/guessing-the-meaning-of-a-number/</link>
					<comments>https://www.johndcook.com/blog/2026/09/14/guessing-the-meaning-of-a-number/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 10:44:45 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247892</guid>

					<description><![CDATA[<p>Suppose I give you an n-digit number and ask you what it represents. This seems impossible, and in theory it is impossible. But in practice it&#8217;s often possible. Apps on a phone may automatically interpret a 10-digit number as a phone number or a 16-digit number as a package tracking number. And very often these interpretations are [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/14/guessing-the-meaning-of-a-number/">Guessing the meaning of a number</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Suppose I give you an <em>n</em>-digit number and ask you what it represents. This seems impossible, and in theory it <em>is</em> impossible. But in practice it&#8217;s often possible.</p>
<p>Apps on a phone may automatically interpret a 10-digit number as a phone number or a 16-digit number as a package tracking number. And very often these interpretations are correct, given the kinds of things most people use their phones for.</p>
<p>It&#8217;s not surprising that a 10-digit number <em>on a phone</em> is a <em>phone number</em>. It&#8217;s more interesting that a 16-digit number is likely a tracking number. It could be other things, such as a credit card number. But people don&#8217;t usually write out credit card numbers in a text note; credit card numbers likely saved in some more opaque way.</p>
<p>I run into a variation of this problem routinely, trying to infer what a number represents inside medical notes.</p>
<p>A five-digit number could be a US postal code, or it could be a <a href="https://www.johndcook.com/blog/2022/09/23/hcpcs-codes/">medical procedure code</a>.</p>
<p>A six-digit number could be a date in MMDDYY format, or it could be a medical record number.</p>
<p>A ten-digit number could be a phone number, or it could be an <a href="https://www.johndcook.com/blog/2024/06/26/npi-number/">NPI</a> (National Provider Identifier) number.</p>
<p>It&#8217;s interesting that it&#8217;s possible make a good guess at what a number means inside unstructured text. Context has been lost, but not all context: you know you&#8217;re looking at medical notes. And that meager bit of context can be surprisingly useful.</p>The post <a href="https://www.johndcook.com/blog/2026/09/14/guessing-the-meaning-of-a-number/">Guessing the meaning of a number</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/14/guessing-the-meaning-of-a-number/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>Bayesian OCR</title>
		<link>https://www.johndcook.com/blog/2026/09/10/bayesian-ocr/</link>
					<comments>https://www.johndcook.com/blog/2026/09/10/bayesian-ocr/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 12:32:31 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<category><![CDATA[Bayesian]]></category>
		<category><![CDATA[Typography]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247884</guid>

					<description><![CDATA[<p>The Greek letter β (beta) and the German letter ß (eszett) look similar, especially in some fonts. Now suppose an OCR program sees some character that could be a beta or could be an eszett. It could calculate some kind of distance between the pixel pattern of the character and the pixel patterns of beta [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/10/bayesian-ocr/">Bayesian OCR</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The Greek letter β (beta) and the German letter ß (eszett) look similar, especially in some fonts.</p>
<p>Now suppose an OCR program sees some character that could be a beta or could be an eszett. It could calculate some kind of distance between the pixel pattern of the character and the pixel patterns of beta and eszett. But that would be discarding context.</p>
<p>If you&#8217;re scanning a Greek document and run into a beta-like symbol, it&#8217;s very likely a beta. If you&#8217;re scanning a German document and run into a beta-like symbol, it <em>could</em> be a beta. For example, it could be a scientific paper that mentions beta particles or beta carotene. But most likely the symbol is an eszett.</p>
<p>The previous paragraph is saying you should compute the <em>conditional</em> probability of a set of pixels representing a character <em>given</em> the language of the document. You could be more sophisticated and look at the position of the symbol in a word as well. For example, if you see a symbol at the end of a Greek word that could either be ο (omicron) or σ (sigma), it&#8217;s likely an omicron because Greek has a different symbol ς for final sigma.</p>
<p>This post is a follow-on to my <a href="https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/">earlier post</a> on the error rate in Google&#8217;s Ngram database. OCR errors are fairly common in that database, so why don&#8217;t they &#8220;just&#8221; fix the errors by using some sort of Bayesian method? OCR software probably does use some sort of Bayesian method, but it&#8217;s not that simple.</p>
<p>In that post I looked at the use of the word <em>grok</em> in English. The Ngram database shows the word being used before it was coined in 1961 due to OCR errors. Why didn&#8217;t Google compute the probability of a word being &#8220;grok&#8221; conditional on the publication date? That would be circular. We happen to know exactly when <em>grok</em> was coined, but in general we might try to determine when a word was coined by looking at a large set of scanned books, like the Ngram database!</p>
<p>Now we could compute the probable value of an ambiguously scanned word by conditioning on the language of the surrounding text. That would be a reasonable thing to do in general, but it could also lead to exactly the kind of errors we see in the Ngram data for <em>grok</em>.</p>
<p>Suppose you see an ambiguously scanned word in a book written in English. There is a higher prior probability that the word is an English word than a German word. Now suppose you see &#8220;gro?&#8221; where ? could be β, ß, or k. Without any context, perhaps the probability of the symbol being a <em>k</em> is small. But <em>grok</em> is an English word and groß is a German word which may lead you to conclude &#8220;?&#8221; is a <em>k</em> and the ambiguous word is <em>grok</em>.</p>
<p>Assigning higher prior probability to English words in English texts is the best thing to do <em>on average</em>, but in particular instances it will lead to errors. That&#8217;s life.</p>
<p>The Ngram database includes millions of scanned books. Google had to use OCR algorithms that work well on average. A linguist with a special interest in a particular word can be more careful and create a more sophisticated probability model (explicit or implicit) customized for their interests. Google did what they could operating at such a large scale.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2025/08/14/uppercase-eszett/">Uppercase eszett</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2011/09/27/bayesian-amazon/">A Bayesian view of Amazon resellers</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/10/bayesian-ocr/">Bayesian OCR</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/10/bayesian-ocr/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>A 50-year-old computer-assisted proof</title>
		<link>https://www.johndcook.com/blog/2026/09/09/four-colors/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 14:59:48 +0000</pubDate>
				<category><![CDATA[Computing]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247877</guid>

					<description><![CDATA[<p>The idea of using computers to assist with proofs is not new. The first major computer-assisted proof was published in 1976, the proof of the four color theorem by Kenneth Appel and Wolfgang Haken. The authors reduced the proof of the four color theorem to verifying calculations on 1,834 configurations, each checked by a computer [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/09/four-colors/">A 50-year-old computer-assisted proof</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The idea of using computers to assist with proofs is not new. The first major computer-assisted proof was published in 1976, the proof of the four color theorem by Kenneth Appel and Wolfgang Haken. The authors reduced the proof of the four color theorem to verifying calculations on 1,834 configurations, each checked by a computer program.</p>
<p>The proof was simplified over the years, and formalized in Coq in 2005. Everyone is satisfied that the theorem is true, but there has never been a satisfying proof, one that a human could read and say &#8220;I see now why any map can be colored using only four colors.&#8221; And there may never be one, but see <a href="https://www.johndcook.com/blog/2013/09/04/homework-problems-for-2090/">this post</a> for a contrary prediction.</p>
<p>The IBM mainframe that ran the calculations completing the proof of the four color theorem did not generate the proof. It simply executed the FORTRAN program that Haken and Appel (and Koch [1]) gave it.</p>
<p>I don&#8217;t see the recent proof of finite-time blowup for solutions to the Navier-Stokes equations as entirely different. Computers did higher-level tasks for the OpenAI team than the mainframe did for Haken and Appel, and these tasks were not as directly programmed as the tasks that were given to the mainframe, but still machines do what they are told to do.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2013/07/19/the-seven-color-map-theorem/">The seven color map theorem</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2019/09/12/detecting-typos/">Detecting errors with the four color theorem</a></li>
</ul>
<p>[1] John A. Koch was a programmer who worked on the four color proof with Haken and Appel. I don&#8217;t know how much credit he deserves, but I suspect it may be more than he was given.</p>The post <a href="https://www.johndcook.com/blog/2026/09/09/four-colors/">A 50-year-old computer-assisted proof</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI is an intelligence multiplier</title>
		<link>https://www.johndcook.com/blog/2026/09/09/ai-multiplier/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 14:17:20 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247875</guid>

					<description><![CDATA[<p>A rising tide may lift all boats, but the AI tide lifts some boats much more than others. By all accounts, the best programmers have had the biggest productivity boost from AI. And top tier mathematicians are using AI to settle long-standing mathematical conjectures. AI is a powerful tool, but tools don&#8217;t come to life [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/09/ai-multiplier/">AI is an intelligence multiplier</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>A rising tide may lift all boats, but the AI tide lifts some boats much more than others.</p>
<p>By all accounts, the best programmers have had the biggest productivity boost from AI. And top tier mathematicians are using AI to settle long-standing mathematical conjectures. AI is a powerful tool, but tools don&#8217;t come to life and make things on their own.</p>
<p>I routinely have naive amateurs [1] send me proofs of open conjectures, and naturally more recent such proofs involve AI. I&#8217;ll get an email saying something like &#8220;I&#8217;ve solved the Collatz conjecture using ChatGPT, but I&#8217;m not a mathematician so I need some help verifying the proof.&#8221; And of course the supposed proof is rubbish.</p>
<p>The recent Navier-Stokes proof is impressive, but AI didn&#8217;t initiate the proof any more than LaTeX did. Nor did a child steer AI into proving the conjecture. Professional mathematicians were able to use AI to pursue their ideas at superhuman speed. But someone without an understanding of the Navier-Stokes problem, and familiarity with recent ideas for approaching the problem, could not have directed AI to produce a proof.</p>
<p>Computer scientists have been saying &#8220;garbage in, garbage out&#8221; from the beginning. A variation on this aphorism for the age of AI would be &#8220;mediocrity in, mediocrity out.&#8221;</p>
<p style="text-align: center;">***</p>
<p>[1] Amateurs can and do make contributions to mathematics. For example, in 2022 David Smith, a retired print technician, discovered a single shape that can be used to create an aperiodic tiling of the plane. By &#8220;naive amateurs&#8221; I mean people who literally do not know what they are talking about.</p>The post <a href="https://www.johndcook.com/blog/2026/09/09/ai-multiplier/">AI is an intelligence multiplier</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>The part of Navier-Stokes no one is talking about</title>
		<link>https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/</link>
					<comments>https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 12:43:36 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<category><![CDATA[Formal methods]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247870</guid>

					<description><![CDATA[<p>Yesterday OpenAI announced a proof that settled a long-standing question about the Navier-Stokes equations from fluid dynamics. The announcement has created a lot of buzz, as one would expect. But there&#8217;s an aspect of OpenAI&#8217;s work that I haven&#8217;t seen anyone talk about: they posted a Lean 4 formal proof at the same time as [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/">The part of Navier-Stokes no one is talking about</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Yesterday OpenAI announced a proof that settled a long-standing question about the Navier-Stokes equations from fluid dynamics. The announcement has created a lot of buzz, as one would expect. But there&#8217;s an aspect of OpenAI&#8217;s work that I haven&#8217;t seen anyone talk about: they posted a Lean 4 formal proof at the same time as their conventional human-readable proof.</p>
<p>Quite a few other mathematical conjectures have been settled recently using AI, and these have also been accompanied with formal proofs, using Lean 4 in particular.</p>
<p>Until very recently, generating machine-verifiable formal proofs has been <strong>excruciatingly tedious</strong>. In 2005, Henk Barendregt and Freek Wiedijk <a href="https://www.cs.ru.nl/~freek/notes/RSpaper.pdf">wrote</a></p>
<blockquote><p>To give an indication of how much work is needed for formalisation, we estimate that it takes approximately one work-week (five work-days of eight work-hours) to formalise one page from an undergraduate mathematics textbook.</p></blockquote>
<p>That was the rule of thumb: <strong>forty hours per page</strong>. And this in the context of undergraduate textbooks. Research publications are much denser than textbooks. Furthermore, page 100 of a textbook probably depends mostly on material on pages 1 through 99. A sentence in a research article could cite anything that has been published before.</p>
<p>Say a research article takes 20 times more effort to formalize than page in an undergraduate textbook. Then formalizing the 166-page paper from OpenAI would take 132,800 person-hours. It took OpenAI 17 hours to verify their proof in Lean. I hesitate to use the word &#8220;revolutionary,&#8221; but lowering the cost of anything by <strong>four orders of magnitude</strong> is revolutionary.</p>
<p>I&#8217;ve used AI to generate formal proofs to check my work just for a little blog post. I wouldn&#8217;t dream of doing that if I had to pay someone a week&#8217;s salary to check my work.</p>
<p>Formal verification doesn&#8217;t just apply to mathematics. You could, for example, formally verify that a set of security policies are consistent and that, given certain assumptions, they accomplish their purpose. You could formally verify that a smart contract imposes a certain maximum liability. You could verify the correctness of mission-critical algorithms. These problems are easier than formalizing mathematics research, and it is easier to quantify the return on investment.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2025/12/24/automation-and-validation/">Automation and validation</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2016/07/11/formal-methods-let-you-explore-the-corners/">Formal methods let you explore the corners</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2020/12/03/formal-proof-roi/">When are formal methods worth the effort?</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/">The part of Navier-Stokes no one is talking about</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/09/formal-method-revolution/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
			</item>
		<item>
		<title>Navier-Stokes in the news</title>
		<link>https://www.johndcook.com/blog/2026/09/08/navier-stokes-in-the-news/</link>
					<comments>https://www.johndcook.com/blog/2026/09/08/navier-stokes-in-the-news/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Tue, 08 Sep 2026 15:26:08 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[Math]]></category>
		<category><![CDATA[Artificial intelligence]]></category>
		<category><![CDATA[Differential equations]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247864</guid>

					<description><![CDATA[<p>There are rumors that a long-standing math problem, one of the Millennium Prize problems, has been solved. The problem concerns technical properties of solutions to the Navier-Stokes equations [1], a set of equations that describe the dynamics of fluid flow. Popular accounts of the problem are often oversimplified and misleading. Some reports will speak of [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/08/navier-stokes-in-the-news/">Navier-Stokes in the news</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>There are rumors that a long-standing math problem, one of the Millennium Prize problems, has been solved.</p>
<p>The problem concerns technical properties of solutions to the <strong>Navier-Stokes equations</strong> [1], a set of equations that describe the dynamics of fluid flow. Popular accounts of the problem are often oversimplified and misleading.</p>
<p>Some reports will speak of the problem as &#8220;solving the Navier-Stokes equations.&#8221; The task is not to write down a closed-form solution, which can&#8217;t be done, or solve the equations numerically, which has been done for decades. The problem is to prove theoretical properties of solutions which are of little interest in practice.</p>
<p>There has been progress toward settling the Navier-Stokes problem. Terence Tao wrote a post on this <a href="https://terrytao.wordpress.com/2026/09/07/finite-time-blowup-with-smooth-forcing-term-for-the-incompressible-porous-medium-boussinesq-and-incompressible-euler-equations/">yesterday</a>.</p>
<p>What&#8217;s also  interesting is the intrigue around the possible solution. A <a href="https://x.com/etale27/status/2097190864598560892">post</a> this morning says</p>
<blockquote><p>If I am reading this correctly, Tristan Buckmaster is alleging OAI has a resolution of Navier-Stokes … which maybe used info from Buckmaster and Levent Alpöge’s private Codex sessions.</p></blockquote>
<p>Buckmaster asked OpenAI whether they used his private sessions and they have not responded. <strong>Update</strong>: <a href="https://openai.com/index/navier-stokes-solution/">Statement</a> from OpenAI.</p>
<h2>Personal note</h2>
<p>This topic connects parts of my career spanning decades. My graduate work was in PDEs and I had some interest in the Navier-Stokes equations. Here are some <a href="https://www.johndcook.com/NavierStokes.pdf">notes</a> I wrote back in the day.</p>
<p>Now I work more with privacy than with PDEs. The question of whether OpenAI uses private data, contradicting their stated policy, is more relevant to my current work than whether the Navier-Stokes equations have global regular solutions.</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2014/08/04/engineering-a-waterpark/">Engineering a waterpark</a></li>
<li class="link"><a href="https://www.johndcook.com/Euler_Lagrange_Equations.pdf">Euler-Lagrange equations</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/data-privacy/">Data privacy</a></li>
</ul>
<p>[1] I never know whether to say equation or equations. You&#8217;ll hear both. You could think of Navier-Stokes as one vector-valued PDE or three scalar-valued equations.</p>The post <a href="https://www.johndcook.com/blog/2026/09/08/navier-stokes-in-the-news/">Navier-Stokes in the news</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/08/navier-stokes-in-the-news/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Ngram error rate</title>
		<link>https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/</link>
					<comments>https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Mon, 07 Sep 2026 18:35:45 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247861</guid>

					<description><![CDATA[<p>The Online Etymological Dictionary gives the following etymology for grok: grok (v.) &#8220;understand empathically,&#8221; 1961, an arbitrary formation by U.S. science fiction writer Robert A. Heinlein (1907-1988) in his book &#8220;Stranger in a Strange Land.&#8221; In the book it is a transliteration of a Martian word and is said to mean etymologically &#8220;to drink.&#8221; It attained [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/">Ngram error rate</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p><span class="hyphens-auto" lang="en">The Online Etymological Dictionary gives the following etymology for <em>grok</em>:</span></p>
<blockquote><p>grok (v.)</p>
<p>&#8220;understand empathically,&#8221; 1961, an arbitrary formation by U.S. science fiction writer Robert A. Heinlein (1907-1988) in his book &#8220;Stranger in a Strange Land.&#8221; In the book it is a transliteration of a Martian word and is said to mean etymologically &#8220;to drink.&#8221; It attained popular use in 1960s-70s counterculture but is perhaps obsolete now except in internet technology circles.</p></blockquote>
<p>I don&#8217;t believe anything in the statement above is disputed. And yet Google&#8217;s Ngram Viewer tells a very different story.</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/grok_ngram.png" width="600" height="220" /></p>
<p>The plot implies that use of the word <em>grok</em> had been increasing before Heinlein&#8217;s book came out and is now much more common than it was in the 1970s. Note that the plot ends before the Grok AI came out in late 2023.</p>
<p>Apparently the Ngram data is unreliable, mainly for two reasons: OCR errors and inaccurate date attribution. Presumably the blip around 1900 was due to the former, OCR causing words like <em>crok</em> or <em>grog</em> to be cataloged as <em>grok</em>. And presumably the rise in usage before 1961 was due to the latter, misattributing the date of sources published after 1961.</p>
<p>The supposed rise in usage before 1961 is interesting. You&#8217;d expect some lag between the time a word circulates in conversation and when it appears in books, but apparently this lag can be smaller than the effect of date misattribution.</p>
<p>Etymonline speculates that <em>grok</em> is &#8220;perhaps obsolete now except in internet technology circles.&#8221; That matches my experience. Even in technological circles, the word was uncommon before Grok was released. Maybe it was more common in print than in conversation.</p>
<h2>Related posts</h2>
<p>Previous posts with Ngram stats. The effects are so large that they&#8217;re probably directionally correct after adjusting for a substantial error rate.</p>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2021/04/18/duodecimal/">Duodecimal vs. Hexadecimal</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2013/06/07/orwellian-vs-huxleyian/">Orwellian vs. Huxleyian</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/">Ngram error rate</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/07/ngram-error-rate/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>Proof of the rank-trace theorem</title>
		<link>https://www.johndcook.com/blog/2026/09/05/proof-of-the-rank-trace-theorem/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Sat, 05 Sep 2026 17:04:13 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Linear algebra]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247858</guid>

					<description><![CDATA[<p>The previous post discussed the motivation for and application of the rank-trace theorem. This post will give a proof. Suppose A is a real symmetric matrix. The rank-trace inequality says where tr is the trace operator, the sum of the elements along the diagonal of the matrix. Terse proof Here&#8217;s the proof in a nutshell: diagonalize A [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/05/proof-of-the-rank-trace-theorem/">Proof of the rank-trace theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>The <a href="https://www.johndcook.com/blog/2026/09/04/stable-rank/">previous post</a> discussed the motivation for and application of the rank-trace theorem. This post will give a proof.</p>
<p>Suppose <em>A</em> is a real symmetric matrix. The rank-trace inequality says</p>
<p><img loading="lazy" decoding="async" class="aligncenter" src="https://www.johndcook.com/rank_trace1.svg" alt="\operatorname{rank}(A)\ge\frac{(\operatorname{tr} A)^2}{\operatorname{tr}(A^2)}" width="143" height="49" /></p>
<p>where tr is the trace operator, the sum of the elements along the diagonal of the matrix.</p>
<h2>Terse proof</h2>
<p>Here&#8217;s the proof in a nutshell: diagonalize <em>A</em> and use the Cauchy-Schwarz inequality.</p>
<h2>Detailed proof</h2>
<p>Now let&#8217;s unpack that. Any real symmetric matrix <em>A</em> is similar to a matrix <em>D</em> with the eigenvalues of <em>A</em> along the diagonal.</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace2.svg" alt="A = PDP^{-1}" width="91" height="18" /></p>
<p>The trace of a matrix stays the same under a similarity transformation, i.e. multiplying by <em>P</em> on one side and its inverse on the other side. So without loss of generality we may as well assume <em>A</em> is diagonal.</p>
<p>The rank of a matrix equals the number of non-zero eigenvalues, so a vector containing the non-zero eigenvalues of <em>A</em></p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace3.svg" alt="v = [\lambda_1, \lambda_2, \ldots, \lambda_r]" width="147" height="17" /></p>
<p>has length <em>r</em> where <em>r</em> is the rank of <em>A</em>. Define <em>w</em> to be the vector of dimension <em>r</em> consisting of all 1&#8217;s.</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace4.svg" alt="w = [1, 1, \ldots, 1]" width="130" height="17" /></p>
<p>Then by the Cauchy-Schwarz inequality we have</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace6.svg" alt="\operatorname{tr}(A)^2 = \langle v, w \rangle^2 \leq \langle v, v \rangle \, \langle w, w \rangle = r \operatorname{tr}(A^2)" width="332" height="22" /></p>
<h2>Cyclic trace property</h2>
<p>Why should a matrix <em>A</em> and its diagonalization <em>D</em> have the same trace?</p>
<p>The trace of a matrix product <em>AB</em> equals the trace of the product <em>BA</em>. To prove this, write out matrix products and the traces, then note that the two expressions are equal.</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/trace_commute.svg" alt=" \begin{align*} \operatorname{tr}(AB) &amp;= \sum_i(AB)_{ii}=\sum_i\sum_k A_{ik}B_{ki} \\ \operatorname{tr}(BA) &amp;= \sum_j(BA)_{jj}=\sum_j\sum_k B_{jk}A_{kj} \end{align*}" width="291" height="98" /></p>
<p>Therefore</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace7.svg" alt="\operatorname{tr}(A) = \operatorname{tr}((PD)P^{-1}) = \operatorname{tr}(P^{-1}(PD)) = \operatorname{tr}(D)" width="351" height="22" /></p>
<p>More generally, trace has the cyclic property</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/cycle_trace.svg" alt="\operatorname{tr}(ABC) = \operatorname{tr}(CAB) = \operatorname{tr}(BCA)" width="242" height="18" /></p>
<p>However, not all permutations preserve the trace. For example, let</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace8.svg" alt="A=\begin{pmatrix}0&amp;1\\0&amp;0\end{pmatrix},\quad B=\begin{pmatrix}0&amp;0\\1&amp;0\end{pmatrix},\quad C=\begin{pmatrix}1&amp;0\\0&amp;0\end{pmatrix}." width="342" height="48" /></p>
<p>Then</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace11.svg" alt="\operatorname{tr}(ABC) = \operatorname{tr}\begin{pmatrix}1&amp;0\\0&amp;0\end{pmatrix} = 1" width="192" height="48" /></p>
<p>but</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace10.svg" alt="\operatorname{tr}(ACB) = \operatorname{tr}\begin{pmatrix}0&amp;0\\0&amp;0\end{pmatrix} = 0" width="192" height="48" /></p>The post <a href="https://www.johndcook.com/blog/2026/09/05/proof-of-the-rank-trace-theorem/">Proof of the rank-trace theorem</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Computing a lower bound on matrix rank</title>
		<link>https://www.johndcook.com/blog/2026/09/04/stable-rank/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Fri, 04 Sep 2026 14:16:16 +0000</pubDate>
				<category><![CDATA[Computing]]></category>
		<category><![CDATA[Math]]></category>
		<category><![CDATA[Linear algebra]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247851</guid>

					<description><![CDATA[<p>Suppose you want to know the rank of an n × n matrix A, the number of linearly independent rows of A, or equivalently the number of linearly independent columns. There are at least three difficulties. Difficulties in computing rank First of all, rank is not a continuous function of a matrix. Since rank is an [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/04/stable-rank/">Computing a lower bound on matrix rank</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Suppose you want to know the rank of an <em>n</em> × <em>n</em> matrix <em>A</em>, the number of linearly independent rows of <em>A</em>, or equivalently the number of linearly independent columns. There are at least three difficulties.</p>
<h2>Difficulties in computing rank</h2>
<p>First of all, rank is not a continuous function of a matrix. Since rank is an integer, an arbitrarily small change in the matrix could cause a discrete change in the rank [1]. A small error in computing <em>A</em> could produce a matrix with a different rank.</p>
<p>Second, finding the rank takes <em>O</em>(<em>n</em>³) operations, which may or may not be an issue depending on context.</p>
<p>Third, you may not have the matrix <em>A</em> in an explicit form. Maybe you&#8217;re able to compute products <em>Av</em> for vectors <em>v</em> but it&#8217;s not practical to form the entire matrix <em>A</em>.</p>
<h2>Rank-trace inequality</h2>
<p>If you don&#8217;t need to know the rank of <em>A</em> per se, but only need to know whether it is above a certain size, a lower bound on the rank may enough.</p>
<p>Suppose <em>A</em> is a Hermitian matrix. If <em>A</em> is real, this means <em>A</em> is symmetric. If <em>A</em> is complex, this means <em>A</em> equals its conjugate transpose. Then the rank-trace inequality says</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/rank_trace1.svg" alt="\operatorname{rank}(A)\ge\frac{(\operatorname{tr} A)^2}{\operatorname{tr}(A^2)}" width="143" height="49" /><br />
The quantity on the right hand side is known as the <strong>stable rank</strong> of <em>A</em>. It&#8217;s not a rank in any algebraic sense, but it gives a lower bound on rank. And it solves the three problems listed above. See the <a href="https://www.johndcook.com/blog/2026/09/05/proof-of-the-rank-trace-theorem/">next post</a> for a proof of the rank-trace theorem.</p>
<h3>Stability</h3>
<p>First of all, trace <em>is</em> a continuous function of a matrix, and so stable rank is also a continuous function of a matrix, provided the denominator isn&#8217;t zero. A small change to a matrix only makes a small change to its stable rank. That&#8217;s why stable rank is called stable.</p>
<h3>Efficiency</h3>
<p>Second, although computing rank takes <em>O</em>(<em>n</em>³) operations, computing stable rank takes only <em>O</em>(<em>n</em>²) operations, though this isn&#8217;t immediately obvious.</p>
<p>The trace of <em>A</em> takes <em>n</em> operations: simply sum the elements on the diagonal of <em>A</em>. But how do you take the trace of <em>A</em>²? Squaring <em>A</em> takes <em>n</em>³ operations, and so if you had to square <em>A</em> to find the trace of <em>A</em>² the rank-trace inequality would have no efficiency advantage over finding the rank of <em>A</em>. But you can compute the trace of <em>A</em>² via</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/trace_A2.svg" alt="\operatorname{tr}(A^2) = \sum_{i=1}^n \sum_{j=1}^n |a_{ij}|^2" width="171" height="57" /></p>
<h3>Formation</h3>
<p>Now suppose you don&#8217;t have the matrix <em>A</em> per se but you do have a way of probing <em>A</em>, computing the product of vectors with <em>A</em>. Maybe <em>A</em> is too large to fit into memory, or explicitly computing the elements of <em>A</em> would take too long.</p>
<p>There are Monte Carlo algorithms for estimating the traces of <em>A</em> and <em>A</em>² that could be used together to estimate the stable rank of <em>A</em>.</p>
<h2>Demonstration</h2>
<p>The following Python code illustrates the discussion above.</p>
<pre>import numpy as np

np.random.seed(20260904)
n = 5
B = np.random.randn(n, n)
A = B.T @ B + 1e-8 * np.eye(n)  # Gram matrix plus a tiny shift =&gt; SPD

rank_A = np.linalg.matrix_rank(A)
tr_A = np.trace(A)
tr_A2 = np.trace(A @ A) # matrix product 
sum_sq = np.sum(A * A) # element-by-element product
stable_rank = (tr_A ** 2) / tr_A2

print(f"A =\n{A}\n")
print(f"rank(A)              = {rank_A}")
print(f"tr(A)                = {tr_A:.12f}")
print(f"tr(A^2) direct       = {tr_A2:.12f}")
print(f"tr(A^2) indirect     = {sum_sq:.12f}")
print(f"stable rank          = {stable_rank:.12f}")
</pre>
<p>The code above produces the output below.</p>
<pre>A =
[[ 1.09945682  0.4899665   0.98901845  0.66983113 -1.35006341]
 [ 0.4899665   0.98531254  0.35067791  0.89757603 -0.72037507]
 [ 0.98901845  0.35067791  4.31233926  0.94556225 -0.54819048]
 [ 0.66983113  0.89757603  0.94556225  1.3494295  -1.33840786]
 [-1.35006341 -0.72037507 -0.54819048 -1.33840786  3.54858332]]

rank(A)              = 5
tr(A)                = 11.295121449420
tr(A^2) direct       = 51.035447533673
tr(A^2) indirect     = 51.035447533673
stable rank          = 2.499826585688
</pre>
<p>[1] Topological argument: A map from a connected space (such as ℝ<sup><em>n</em>×<em>n</em></sup>) onto a discrete space (such as ℤ) cannot be continuous, otherwise the inverse images of the points in the range would partition the connected space into disjoint open sets, violating the definition of a connected space.</p>The post <a href="https://www.johndcook.com/blog/2026/09/04/stable-rank/">Computing a lower bound on matrix rank</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Hugging Face Easter Egg</title>
		<link>https://www.johndcook.com/blog/2026/09/03/hugging-face-easter-egg/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 23:34:41 +0000</pubDate>
				<category><![CDATA[Uncategorized]]></category>
		<category><![CDATA[Unicode]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247847</guid>

					<description><![CDATA[<p>NVIDIA has offered to buy Hugging Face for $12,930,300,000. 129303 is the Unicode code point for the Hugging Face emoji (U+1F917), which you can verify with the following Python code. &#62;&#62;&#62; import unicodedata &#62;&#62;&#62; 129303 == 0x1F917 True &#62;&#62;&#62; unicodedata.name(chr(0x1F917)) 'HUGGING FACE' Related posts Prevent characters from displaying as emoji Unicode, Tolkien, and Privacy Unicode [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/03/hugging-face-easter-egg/">Hugging Face Easter Egg</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>NVIDIA has offered to buy Hugging Face for $12,930,300,000.</p>
<p>129303 is the Unicode code point for the Hugging Face emoji (U+1F917), which you can verify with the following Python code.</p>
<pre>&gt;&gt;&gt; import unicodedata
&gt;&gt;&gt; 129303 == 0x1F917
True
&gt;&gt;&gt; unicodedata.name(chr(0x1F917))
'HUGGING FACE'
</pre>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/huggingface.png" alt="Hugging Face emoji" width="200" height="200" /></p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2022/09/30/preventing-emoji/">Prevent characters from displaying as emoji</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/03/09/tengwar/">Unicode, Tolkien, and Privacy</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/03/09/unicode-surrogates/">Unicode surrogates</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2022/10/02/flags-unicode/">Making flags in Unicode</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/09/03/hugging-face-easter-egg/">Hugging Face Easter Egg</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>New RSA number factored</title>
		<link>https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/</link>
					<comments>https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 17:21:45 +0000</pubDate>
				<category><![CDATA[Computing]]></category>
		<category><![CDATA[Cryptography]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247845</guid>

					<description><![CDATA[<p>Eric Lu announced on X today that he has factored RSA-260, a number N with 260 digits (862 bits) that is the product of two large primes [1]. RSA numbers are challenge problems posed to gauge the security of RSA encryption, which rests on the difficulty of factoring large numbers [2]. The naming scheme is [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/">New RSA number factored</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>Eric Lu <a href="https://x.com/penlume/status/2095372672356212876?s=20">announced</a> on X today that he has factored RSA-260, a number <em>N</em> with 260 digits (862 bits) that is the product of two large primes [1].</p>
<p>RSA numbers are challenge problems posed to gauge the security of RSA encryption, which rests on the difficulty of factoring large numbers [2]. The naming scheme is confusing because RSA-<em>n</em> might have <em>n</em> digits or <em>n</em> bits. For example, RSA-768 is smaller than RSA-260 because the former has 768 bits and the latter has 260 digits.</p>
<p>RSA-260 is the largest RSA number factored so far. What does the news of its factorization say about the security of RSA?</p>
<p>Based on equations <a href="https://www.johndcook.com/blog/2025/09/30/time-needed-to-factor-large-integers/">here</a>, an RSA key with 862 bits would have a security level of 74 bits, i.e. the same security level as symmetric encryption with a 74-bit key. The minimum recommended RSA key size now is 2048 bits, which has a security level of 107 bits.</p>
<p>Security levels are on a logarithmic scale: each additional bit of security doubles the effort required to break the encryption by brute force. So breaking a 2048-bit RSA key would take 2<sup>34</sup>, roughly 10<sup>10</sup>, times more effort than factoring RSA-260. All this depends on numerous assumptions, such as the state of factorization algorithms and the non-existence of CRQC [3].</p>
<p><strong>Update</strong>: A week after the initial announcement, Eric Lu released an <a href="https://cognition.com/blog/factoring-rsa-260">article</a> explaining how he factored RSA-260. He says &#8220;In total, I estimate that this factorization cost about 4,900 GPU-days, or 13.5 GPU-years, which is about $400k at current market prices.&#8221;</p>
<h2>Related posts</h2>
<ul>
<li class="link"><a href="https://www.johndcook.com/blog/2019/02/11/rsa-duplication-flaws/">RSA implementation flaws</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2023/08/05/rsa-private-key/">Generating and inspecting an RSA key</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2025/08/05/martin-gardners-rsa/">Martin Gardner&#8217;s RSA article</a></li>
<li class="link"><a href="https://www.johndcook.com/blog/2026/06/13/rsa-munitions-t-shirt/">RSA munitions T-shirt</a></li>
</ul>
<p>[1] <em>N</em> = <em>pq</em> = 22112825529529666435281085255026230927612089502470015394413748319128822941402001986512729726569746599085900330031400051170742204560859276357953757185954298838958709229238491006703034124620545784566413664540684214361293017694020846391065875914794251435144458199</p>
<p><em>p</em> = 4397328654844826923795068102505872571721883526553349659561256924505973939597593482272505698004801207988043088656411102133523080581</p>
<p><em>q</em> = 5028695206842569864686141618253083416610081090075366674776775706538324961364412200138116378509733307971876652984898985905923678379</p>
<p>[2] The ability to efficiently factor large primes would break RSA. It&#8217;s possible that there&#8217;s a way to break RSA without being able to factor large numbers. More on that <a href="https://www.johndcook.com/blog/2025/01/06/rsa-factoring/">here</a>.</p>
<p>[3] Cryptographically-relevant quantum computer. Quantum computers exist, but so far they&#8217;re cryptographically irrelevant. So far quantum computers cannot factor 21 without <a href="https://www.johndcook.com/blog/2026/03/31/quantum-y2k/">cheating</a>.</p>The post <a href="https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/">New RSA number factored</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/09/03/new-rsa-number-factored/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Patented application of linear algebra</title>
		<link>https://www.johndcook.com/blog/2026/08/31/patented-application-of-linear-algebra/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Mon, 31 Aug 2026 23:26:12 +0000</pubDate>
				<category><![CDATA[Computing]]></category>
		<category><![CDATA[Linear algebra]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247831</guid>

					<description><![CDATA[<p>I just found out Brian Beckman and I got a patent on work we did for GSI Technology [1]. Nearly all the work I do is under an NDA, so I don&#8217;t often get a chance to talk about my projects. This work is public now that it&#8217;s in a patent; I suppose it has [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/08/31/patented-application-of-linear-algebra/">Patented application of linear algebra</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>I just found out Brian Beckman and I got a patent on work we did for GSI Technology [1]. Nearly all the work I do is under an NDA, so I don&#8217;t often get a chance to talk about my projects. This work is public now that it&#8217;s in a patent; I suppose it has been public since the application was published.</p>
<p>Brian did most of the work on the project. My contribution was to mathematically formalize low-level operations on sheets of bits using linear algebra over a binary field. Lots of Hadamard products and outer products, if I remember correctly. When you can reduce computations to algebra, you can prove that a sequence of operations is correct, and you can find optimizations by simplifying expressions.</p>
<p>The patent mentions a programming language called Tartan. I suggested calling it plaid because it used matrices with mask patterns that reminded me of a plaid pattern, and Brian countered saying we should call it Tartan. I like that name better.</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/tartan2.png" alt="Figure 5B from patent" width="550" height="514" /></p>
<p style="text-align: center;">Figure 5B from the patent.</p>
<p>[1] Brian Beckman and John D. Cook. Compiler for a parallel processor. U.S. Patent 12,717,871 B2. Applicant/Assignee: GSI Technology Inc., Sunnyvale, CA.</p>The post <a href="https://www.johndcook.com/blog/2026/08/31/patented-application-of-linear-algebra/">Patented application of linear algebra</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Making the unnecessary easier</title>
		<link>https://www.johndcook.com/blog/2026/08/28/making-the-unnecessary-easier/</link>
					<comments>https://www.johndcook.com/blog/2026/08/28/making-the-unnecessary-easier/#comments</comments>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Fri, 28 Aug 2026 14:35:36 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247823</guid>

					<description><![CDATA[<p>I watched a few videos this morning, looking for ideas of what I could use AI to do. In one video, someone had an agent monitor tech news sites every 30 minutes to notify him of a variety of developments. No doubt that&#8217;s less effort than visiting a bunch of sites every half hour, but [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/08/28/making-the-unnecessary-easier/">Making the unnecessary easier</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>I watched a few videos this morning, looking for ideas of what I could use AI to do.</p>
<p>In one video, someone had an agent monitor tech news sites every 30 minutes to notify him of a variety of developments. No doubt that&#8217;s less effort than visiting a bunch of sites every half hour, but not monitoring tech news in real time takes even less effort.</p>
<p>Another video mentioned having Grok Bot order a sandwich through DoorDash. If I want a sandwich, I make a sandwich.</p>
<p>One video showed how to manage dozens messaging services. Maybe you could just not use dozens of messaging services.</p>
<p>And of course there are videos on creating agents to monitor the agents that monitor your news, order your sandwiches, and manage your messages.</p>
<p>All the use cases I saw were ways to using technology mitigate problems caused by technology, making it easier to do things that don&#8217;t need to be done, or at least things that I don&#8217;t need to do.</p>
<p>Of course different people have different needs. Some people have a professional need to monitor news in real time, for example, and having an agent help with that could be big win. I suspect, however, that the use cases that you&#8217;ll see most often in YouTube videos have been made up to appeal to a wide audience rather than to scratch the author&#8217;s itch.</p>
<p>Productivity is deeply personal. As I said in an <a href="https://www.johndcook.com/blog/2023/06/03/productive-productivity/">earlier post</a>, the scripts I&#8217;ve found most useful are of <em>zero</em> interest to anyone else because they are so specific to my work. I&#8217;ve mostly automated tasks with Python and bash, not with AI.</p>
<p>I&#8217;m not trying to avoid using AI. As I said at the top of the post I&#8217;m looking for more ways to take advantage of it. But I don&#8217;t want to fall into the trap of doing more easily what doesn&#8217;t need to be done. Or, to put a finer point on it, I don&#8217;t want to find ways for <em>my business</em> to do things that <em>we</em> do not need to do, things that other businesses may need to do.</p>
<h2>Related posts</h2>
<ul>
<li class='link'><a href='https://www.johndcook.com/blog/2023/06/03/productive-productivity/'>Productive productivity</a></li>
<li class='link'><a href='https://www.johndcook.com/blog/2010/12/07/cascading-needs/'>Maybe you only need it because you have it</a></li>
</ul>The post <a href="https://www.johndcook.com/blog/2026/08/28/making-the-unnecessary-easier/">Making the unnecessary easier</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
					<wfw:commentRss>https://www.johndcook.com/blog/2026/08/28/making-the-unnecessary-easier/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Second solutions</title>
		<link>https://www.johndcook.com/blog/2026/08/27/second-solutions/</link>
		
		<dc:creator><![CDATA[John]]></dc:creator>
		<pubDate>Thu, 27 Aug 2026 13:14:35 +0000</pubDate>
				<category><![CDATA[Math]]></category>
		<category><![CDATA[Differential equations]]></category>
		<category><![CDATA[Special functions]]></category>
		<guid isPermaLink="false">https://www.johndcook.com/blog/?p=247814</guid>

					<description><![CDATA[<p>This post provides a couple examples to go along with two earlier posts. The pattern we&#8217;re illustrating is families of polynomials pn(x) that each satisfy a differential equation and a three-term recurrence. The differential equations have a second solution qn(x) that is the larger solution with respect to x but the smaller solution with respect to n. In both [&#8230;]</p>
The post <a href="https://www.johndcook.com/blog/2026/08/27/second-solutions/">Second solutions</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></description>
										<content:encoded><![CDATA[<p>This post provides a couple examples to go along with two <a href="https://www.johndcook.com/blog/2026/08/26/junk-solutions/">earlier</a> <a href="https://www.johndcook.com/blog/2026/08/24/numerical-instability-recurrece/">posts</a>.</p>
<p>The pattern we&#8217;re illustrating is families of polynomials <em>p</em><sub><em>n</em></sub>(<em>x</em>) that each satisfy a differential equation and a three-term recurrence. The differential equations have a second solution <em>q</em><sub><em>n</em></sub>(<em>x</em>) that is the larger solution with respect to <em>x</em> but the smaller solution with respect to <em>n</em>.</p>
<p>In both the examples below <em>p</em><sub><em>n</em></sub>(<em>x</em>) is a polynomial, and so bounded on the interval [−1, 1], and <em>q</em><sub><em>n</em></sub>(<em>x</em>) is not a polynomial, with singularities at ±1. This is analogous to the previous examples with Bessel functions <em>J</em><sub><em>n</em></sub>(<em>x</em>) and <em>Q</em><sub><em>n</em></sub>(<em>x</em>) that satisfy the same differential equation but have contrasting behavior with respect to <em>x</em> versus <em>n</em>.</p>
<h2>Legendre polynomials</h2>
<p>The differential equation</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/legendre_de.svg" alt="(1-x^2)\,y^{\prime\prime} - 2x\,y^\prime + n(n+1)\,y = 0 " width="258" height="23" /></p>
<p>has two solutions for each <em>n</em>, <em>P</em><sub><em>n</em></sub>(<em>x</em>) and <em>Q</em><sub><em>n</em></sub>(<em>x</em>).</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/legendre_P_Q.png" width="600" height="450" /></p>
<p>The solutions <em>P</em><sub><em>n</em></sub>(<em>x</em>) are the Legendre polynomials. The solutions <em>Q</em><sub><em>n</em></sub>(<em>x</em>) are not polynomials but involve a term log((1 + <em>x</em>)/(1 − <em>x</em>)) that blows up at 1 and −1. But for fixed <em>x</em> and increasing <em>n</em>, <em>P</em><sub><em>n</em></sub>(<em>x</em>) grows exponentially and <em>Q</em><sub><em>n</em></sub>(<em>x</em>) decays exponentially, provided |<em>x</em>| &gt; 1.</p>
<h2>Chebyshev polynomials</h2>
<p>The differential equation</p>
<p><img loading="lazy" decoding="async" class="aligncenter" style="background-color: white;" src="https://www.johndcook.com/chebyshev_de.svg" alt="(1-x^2)\,y^{\prime\prime} - x\,y^\prime + n^2\,y = 0" width="204" height="23" /></p>
<p>has two solutions for each <em>n</em>, <em>T</em><sub><em>n</em></sub>(<em>x</em>) and <em>V</em><sub><em>n</em></sub>(<em>x</em>).</p>
<p><img loading="lazy" decoding="async" class="aligncenter size-medium" src="https://www.johndcook.com/chebyshev_T_V.png" width="600" height="438" /></p>
<p>The solutions <em>T</em><sub><em>n</em></sub>(<em>x</em>) are the Chebyshev polynomials. The solutions <em>V</em><sub><em>n</em></sub>(<em>x</em>) are not polynomials but involve a term √(<em>x</em>² — 1) that become vertical up at 1 and −1. But for fixed <em>x</em> with |<em>x</em>| &gt; 1 and increasing <em>n</em>, <em>T</em><sub><em>n</em></sub>(<em>x</em>) grows exponentially and <em>V</em><sub><em>n</em></sub>(<em>x</em>) decays exponentially.</p>The post <a href="https://www.johndcook.com/blog/2026/08/27/second-solutions/">Second solutions</a> first appeared on <a href="https://www.johndcook.com/blog">John D. Cook</a>.]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
