<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Statistical Modeling, Causal Inference, and Social Science</title>
	<atom:link href="https://statmodeling.stat.columbia.edu/feed/" rel="self" type="application/rss+xml" />
	<link>https://statmodeling.stat.columbia.edu</link>
	<description></description>
	<lastBuildDate>Fri, 28 Aug 2026 18:08:49 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.4.10</generator>
	<item>
		<title>What stories should we tell about science now?</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/28/what-stories-should-we-tell-about-science-now/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/28/what-stories-should-we-tell-about-science-now/#comments</comments>
		
		<dc:creator><![CDATA[Jessica Hullman]]></dc:creator>
		<pubDate>Fri, 28 Aug 2026 16:11:15 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Miscellaneous Science]]></category>
		<category><![CDATA[Sociology]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54569</guid>

					<description><![CDATA[This is Jessica. Like many academics, I am concerned about what sort of new steady state U.S. universities will find themselves in after the dust settles on recent transitions. Namely, the last few years have brought funding cuts, targeted visa &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/28/what-stories-should-we-tell-about-science-now/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><span style="font-weight: 400">This is Jessica. Like many academics, I am concerned about what sort of new steady state U.S. universities will find themselves in after the dust settles on recent transitions. Namely, the last few years have brought funding cuts, targeted visa policy, reduced demand for grad degrees, and a general brain drain to industry (particularly noticeable in AI and computer science). It’s disorienting to think that academia has already peaked, and that the prestige ranking of the R1 faculty job over the top industry research positions (at least in computer science) might be inverting. But things feel very different than they did even a year ago. The reality of there being less money available to pay for basic aspects of research really started to hit me in the last six months. Post-covid, working on campus became less lively, but now it also feels like our collective attention is anxiously focused on Silicon Valley or Washington D.C. We hold faculty meetings where we discuss things like, Is there any way we can help local faculty members who were laid off from tenure track jobs? How will we ensure we can fund all of our own PhDs, given that TA quotas stay fixed but faculty are running out of funding runway? </span></p>
<p><span style="font-weight: 400">To some, this is an overdue rebalancing. Nate Silver, for example, calls getting a PhD a “<a href="https://x.com/NateSilver538/status/2089183930385600588">much worse value proposition than 20 years ago</a>”, and predicts that elite higher ed will become “<a href="https://x.com/NateSilver538/status/2090268424488280242">~50% less relevant in the new steady state,</a>” which is in his eyes a good recalibration.</span></p>
<p><span style="font-weight: 400">But it’s worth reflecting on what is lost exactly, if this dwindling of minds and resources continues? How should we think about the value of what universities provide over industry, like intellectual autonomy, or training on how to think scientifically? As a professor, I could make a list of the things that have kept me in academia–being free to work on the problems I find most important, the diversity of topics I can work on at any given time, grad students who care about doing deep work, having time to think about the best solution to a problem. But at an aggregate level, it’s less clear what the equation is.  </span></p>
<p><span style="font-weight: 400">As I was puzzling over all this, I attended a metascience conference, where there was a panel on “the social contract for science.” This is the transactional relationship dating back to at least the 1950s, by which scientists receive public funding and autonomy and society gets the benefits of scientific research. Back in the 19th century, scholars began to make a distinction between “pure” science–research unmotivated by any particular application–and applied science. The social contract takes this distinction and further presupposes a dependence relationship: what is confusingly called the “linear model”, the idea that pure science provides the well from which applied science contributions are drawn. Threaten this foundation, e.g., by letting applied science intercept too much of the resources society puts toward science, and we risk running out of useful innovations. Or so the story goes.</span></p>
<p><span style="font-weight: 400">The social contract for science was an attempt to cement the importance of scientific understanding to society, making it an interesting counterpart to the current moment. If the actions of the current administration to direct funds away from universities, and the possibility of using AI to produce research output without understanding, are threatening our sense of what science should be, the history of science policy provides some perspective on how our expectations got shaped in the first place.</span></p>
<p><b>In search of the mysterious fruits of basic science</b></p>
<p><span style="font-weight: 400">I’ve been reading the work of philosopher Heather Douglas, who has traced and critiqued the basic versus applied science distinction, the linear model as justification, and the idea of scientific freedom as limited social responsibility (see, e.g., </span><a href="https://www.sciencedirect.com/science/article/abs/pii/S0039368114000132"><span style="font-weight: 400">here</span></a><span style="font-weight: 400"> and </span><a href="https://link.springer.com/article/10.1007/s11229-023-04477-9"><span style="font-weight: 400">here</span></a><span style="font-weight: 400">, or </span><a href="https://upittpress.org/books/9780822960263/"><span style="font-weight: 400">her book</span></a><span style="font-weight: 400"> on the value-free ideal). Popularized by Vannevar Bush after WWII, in a report prepared for President Roosevelt, basic science is a reframing of pure science, presented as “scientific capital,” providing the principles and conceptions to power new products and processes years into the future. Bush called for deliberate policy to guard against the otherwise inevitable scenario where applied science drives out the pure. One of the eventual outcomes of his report was the creation of the NSF.</span></p>
<p><span style="font-weight: 400">But despite the pragmatic nature of basic science espoused by Bush, as a derivation of pure science, it is hard to separate from less tangible values. One is that scientific understanding is a good outside of practical application, at both the individual and societal level. The earliest advocates of pure science associated it with being closer to God. Post-Enlightenment, this view gave way to a more secular superiority complex, which implied the strong character of the pure scientist, who chose to eschew wealth. “The highest occupation of mankind”, Henry Rowland called it in his Gilded Age era essay, “A Plea for Pure Science,” which bemoaned the vulgarity of attributing scientific greatness to the applied scientist rather than the pure.</span></p>
<p><span style="font-weight: 400">From a less moralistic point of view, we’ve been encouraged to believe that a society that has rigorous ways of understanding the world is better off over one that doesn’t. Throughout history, understanding the laws of Nature has been portrayed as a good in itself, along with an intellectual life. From this view, by educating people on how to pursue deep understanding of the world, universities provide the general good of scientific thinking to society. If we believe in the intrinsic value of reading, writing, or intellectual discussion, then it would seem we should value the university as a place that provides the kind of timespan and environment needed to develop these skills.</span></p>
<p><a href="https://substack.com/home/post/p-202140738"><span style="font-weight: 400">Some argue</span></a><span style="font-weight: 400"> that the university has come to serve too many conflicting purposes (research engine, job training center, credentialer, incubator of coming-of-age experiences), and should go back to its classical roots: training in oral reasoning and rhetoric, ethics and moral judgment, historical analysis, and the cultivation of taste and discrimination. This may be a useful refocusing, but it offers little consolation for the fact that the elite research infrastructure that helped this country establish and maintain scientific leadership for decades is in the process of being gutted.</span></p>
<p><span style="font-weight: 400">If we take our intuitions from the linear model, we might protest that innovation will suffer if universities’ research purposes are deprioritized. The post WWII science-industrial complex expanded the presence of basic research in industry, but studies suggest that the <a href="https://www.sciencedirect.com/science/article/abs/pii/S0048733304000058">knowledge generating role of</a> <a href="https://sms.onlinelibrary.wiley.com/doi/10.1002/smj.2693">corporate R&amp;D has been on the decline</a> for years. To the extent that basic research is the supplier of downstream applications, it would seem we need universities more than ever.</span></p>
<p><span style="font-weight: 400">But the distinction between basic and applied science that’s become synonymous with how we envision science has never been airtight. Critics questioned how an institution could be built around a distinction that seemed to amount to little more than a difference in intention, since applied research sometimes produced important new general knowledge, and pure science contributions sometimes had direct applicability. </span></p>
<p><span style="font-weight: 400">AI research is a recent example. Not only is serious money being made without necessarily requiring advanced degrees, research positions do not require PhDs. By some accounts, passing 30 years old puts one in the older demographic of researchers at frontier AI companies. Yet much of the visible innovation in frontier model development has been heavily concentrated in industry labs, including transformer models, scaling laws, and AlphaFold.</span></p>
<p><span style="font-weight: 400">Of course, AI owes much to academia. The amazing thing about deep learning and LLMs, to anyone who was paying attention to NLP before these developments, is that after many years of AI research contributing interesting questions but lackluster results, the technology finally seemed to work. Would we have had the foundations for deep learning if perceptrons had not been stubbornly pursued by academics like Frank Rosenblatt at Cornell early on, picked up again in the 1980s by Rumelhart and McClelland’s Parallel Distributed Processing group, despite multiple periods during which consensus said connectionist approaches were unlikely to pay off?</span></p>
<p><span style="font-weight: 400">The challenge is that arguing that “someday the research will pay off,” without being able to point to any hard evidence that basic research is, on average, worth the investment, is not such a convincing argument. According to Douglas, studies have been attempted to show the payoffs of basic research, but without very impressive results. Uncertainty about what time scale we should expect between discovery and application makes this kind of exercise difficult.</span></p>
<p><span style="font-weight: 400">At the same time, it’s hard to dispute that monetary incentives can sometimes discourage exploration that would eventually pay off. In evaluating the role of academic research to AI progress, we should keep in mind the uniquely massive private investments AI companies have received, and be cautious using it as a general example. It would be premature to conclude that because progress (in terms of models’ standalone capabilities as measured by benchmarks) doesn’t seem to depend much on academic research at the moment, cutting off academic research would be immaterial. Particularly unfortunate about the historical contingency of frontier companies defining the direction of the field is that they are focused on a pretty narrow space of methods, evaluations, and design ideas. But the power and resources they hold give newcomers to AI research the impression that ideas outside this narrow space aren’t important.</span></p>
<p><b>Indulgence, autonomy, and social responsibility</b></p>
<p><span style="font-weight: 400">As suggested above, it’s always been tempting to bring moral judgment to bear on the basic versus applied research divide. The latest moment with AI research is no exception. Does the moral high ground belong to those who are staying in academia, underfunded or not, to preserve university culture and their autonomy from corporate interests, or those who are willing to give up a comfortable job to shape the impact of AI as a product in the world?</span></p>
<p><span style="font-weight: 400">From one perspective, the academy, as a haven for basic science, has always been at risk of being seen as indulgent. In practice, building a scientific career is in many ways a process of identity development and fulfillment for the scientist. Historical pure science rhetoric associated the pursuit of scientific truth with self-realization. But talking about personal fulfillment does not go over well when your opponent is promoting the idea that science could do more direct good for the country or humanity. Academic scientists have always been at risk of coming off as being self-indulgent, insular, or dilettante when they defend understanding for understanding’s sake. The current political moment is just rehashing old themes.</span></p>
<p><span style="font-weight: 400">Another unfortunate historical association of basic science is with insularity and shirking responsibility. After WWI era advances in chemical warfare and explosives, the social responsibility of the scientist became a much greater concern. Philosopher John Dewey came down sharply on the idea that an autonomous space for pure science, unhindered by societal concerns, was something to strive for. Instead, he argued that this impetus to protect pure science was partly a convenient abdication of moral responsibility for the downstream outcomes of research, a “shirking of responsibility.’’</span></p>
<p><span style="font-weight: 400">Dewey’s concerns came at a time where philosophy was itself seeking to be more scientific. According to Douglas, Dewey’s views on how philosophers should approach science–through greater integration of societal concerns–lost out to the argument espoused by Bertrand Russell, who instead valorized the “disinterested intellectual curiosity which characterizes the genuine man of science.” The latter view became the more accepted one, and our definition of scientific freedom arose in tandem with expectations of limited social responsibility. It’s not particularly surprising, then, to encounter beliefs that academia is not the place to go if you want to have impact in the world.</span></p>
<p><span style="font-weight: 400">At the same time, it seems hard to deny that at this point of time in AI, where we have a large imbalance of power and resources, there is something to be said for the autonomy afforded by the university or nonprofit. Some beliefs about AGI coming out of Silicon Valley border on religious. My biggest concern if I were to join an AI company at this point in time would be losing my ability to think for myself about what problems deserve priority. Having greater agency and impact are attractive, but not if they come at the expense of one’s internal compass or values. As Brendan McCord </span><a href="https://substack.com/@cosmosinstitute/note/c-316622631"><span style="font-weight: 400">said recently</span></a><span style="font-weight: 400">, “Autonomy is different from agency. Agency is getting things done&#8230;You can be more effective than you’ve ever been, and you can be less the author of your own life than you’ve ever been.” Against the groupthink of Silicon Valley, the value of the intellectual autonomy academia provides does feel real. Though it’s unclear how valuable this autonomy will continue to be if academics and others outside the big labs can’t retain enough funding or visibility into frontier model development to remain relevant.</span></p>
<p><b>The problem with defining progress as prediction and control </b></p>
<p><span style="font-weight: 400">In a 2014 article called </span><a href="https://www.sciencedirect.com/science/article/abs/pii/S0039368114000132"><span style="font-weight: 400">Pure science and the problem of progress</span></a><span style="font-weight: 400">, Douglas suggests that if the pure/applied science distinction doesn’t survive scrutiny (which she argues it does not), we’re left with an account of scientific progress based on our ability to predict, intervene, and control our world. But this is not a definition of progress we should be content with:</span></p>
<blockquote><p><span style="font-weight: 400">&#8220;Any increase in the capacity to predict or control the thoughts and feelings of human beings would count as scientific progress. An increased capacity to destroy human subpopulations (through, say, targeted pathogens) would count as scientific progress. Developing new heinous capacities would count as scientific progress. Unlike Rowland, we should have no illusions that greater causal efficacy, greater power of intervention, will in fact always provide a better society.&#8221; (p. 63)</span></p></blockquote>
<p><span style="font-weight: 400">If misaligned AI, our own creation, changes how we view ourselves and the world, if it convinces some of us it has all of our best interests at heart even as it feeds our insecurities, or pursues its own goals in the background, is that scientific progress?</span></p>
<p><span style="font-weight: 400">Douglas argues that judging real progress requires society to weigh in. When it comes to AI, this is happening through pushback against data centers, and the pace of AI progress, and the culture of Silicon Valley. Adoption matters too, but can’t be a substitute for evaluation. We need institutions independent of the companies to help interpret what’s going on. In the midst of changes to so many of our current institutions, we should expect the story we ultimately tell to take time to sort out.</span></p>
<p><span style="font-weight: 400">In the meantime, defending academia as a category of research, or a moral standard, is a dead end. What seems more reasonable to advocate is a set of conditions — time, autonomy, training in scientific judgment, the evaluation of new approaches independent of their profitability. The value of these ingredients isn’t easily summarizable in some neat story, because what drives scientific progress is not that simple. But institutions that help society judge what’s been achieved seem worth defending.</span></p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/28/what-stories-should-we-tell-about-science-now/feed/</wfw:commentRss>
			<slash:comments>14</slash:comments>
		
		
			</item>
		<item>
		<title>Where&#8217;s the Evel Knievel blockbuster biopic?</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/28/wheres-the-evel-knievel-blockbuster-biopic/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/28/wheres-the-evel-knievel-blockbuster-biopic/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 28 Aug 2026 13:51:13 +0000</pubDate>
				<category><![CDATA[Art]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53515</guid>

					<description><![CDATA[A few years ago we discussed Objects of the class Jacques Cousteau: people who are world famous (or at least world famous in the U.S.) but of a category for which there’s only one famous person. Some examples: • Cousteau &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/28/wheres-the-evel-knievel-blockbuster-biopic/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>A few years ago we discussed <a href="https://statmodeling.stat.columbia.edu/2022/01/09/object-of-the-class-jacques-cousteau/">Objects of the class Jacques Cousteau</a>:  people who are world famous (or at least world famous in the U.S.) but of a category for which there’s only one famous person.  Some examples:</p>
<p>• Cousteau was a world-famous underwater photographer . . . actually the only famous underwater photographer.</p>
<p>• Another example, also from the 70s (sorry): Marcel Marceau. A world-famous mime . . . actually the only famous mime.</p>
<p>• The guy who wrote All Creatures Great and Small was a world-famous veterinarian . . . actually the only world-famous veterinarian.</p>
<p>• Jackson Pollock is a world-famous drip painter, and indeed the only famous drip painter. But it’s not clear to me that “drip painter” should count as a category.</p>
<p>• Stradivari:  same idea, he works if you&#8217;re willing to count &#8220;violin maker&#8221; as a category.</p>
<p>• Tony Hawk:  world-famous skateboarder, the only world-famous skateboarder.</p>
<p>• Luther Burbank:  world-famous agriculturalist, the only world-famous agriculturalist.</p>
<p>• John Philip Sousa:  world-famous composer of marches, the only world-famous composer of marches.</p>
<p>• Margaret Mead:  world-famous anthropologist, the only world-famous anthropologist.</p>
<p>• Alan Turing:  world-famous codebreaker, the only world-famous codebreaker.</p>
<p>• Temple Grandin:  world-famous slaughterhouse designer, the only world-famous slaughterhouse designer.  She&#8217;s kind of like Joseph Joanovici, the only world-famous scrap metal dealer, in that she&#8217;s famous for her story more than for her occupation; nevertheless, as with Joanovici, the occupation is central to the story.</p>
<p>• Noam Chomsky:  world-famous linguist, the only world-famous linguist.</p>
<p>• Jim Henson:  world-famous puppeteer, the only world-famous puppeteer.  OK, maybe we need to count Frank &#8220;Yoda&#8221; Oz too.</p>
<p>• Frederick Law Olmsted:  world-famous park designer, the only world-famous park designer.</p>
<p>• Oscar Pistorius:  world-famous paralympic runner, the only famous paralympic runner.  This works even without adding &#8220;killer&#8221; to his title, but it was the killing that kept him famous.</p>
<p>• Weird Al Yankovic:  world-famous writer of novelty songs, the only . . .</p>
<p>• We could also throw in Dr. Demento:  world-famous DJ of novelty songs, but that seems like too narrow of a category.</p>
<p>• Anna Wintour:  if you&#8217;re willing to count &#8220;fashion magazine editor&#8221; as a category.</p>
<p>• Freddy Mercury is the only world-famous person from Zanzibar, but I don&#8217;t think that really counts.  We&#8217;re talking here about categories defined by what you&#8217;ve done, not where you&#8217;re from.</p>
<p>And, of course:</p>
<p>• Evel Knievel:  world-famous daredevil, the only world-famous daredevil.</p>
<p>There was some discussion of how famous he still is, and commenter Manuel <a href="https://statmodeling.stat.columbia.edu/2022/01/09/object-of-the-class-jacques-cousteau/#comment-2042180">wrote</a>:</p>
<blockquote><p>Evel Knievel is just one blockbuster biopic away from being world famous.</p></blockquote>
<p>This made me wonder:  why hasn&#8217;t there been a blockbuster biopic of Knievel?  Wikipedia informs me that there have been three Knievel biopics already, including a biopic back in 1971 starring George Hamilton! and a 2004 TV movie directed by John Badham, who&#8217;s not a nobody&#8212;he also directed Saturday Night Fever.  Also this:</p>
<blockquote><p>On December 19, 2024, a new biographical film adaptation of Evel&#8217;s life was reported to be in the works with La La Land director Damien Chazelle attached to direct, William Monahan set to pen the script, and negotiations with Leonardo DiCaprio to star as Knievel.</p></blockquote>
<p>That sounds like something.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/28/wheres-the-evel-knievel-blockbuster-biopic/feed/</wfw:commentRss>
			<slash:comments>10</slash:comments>
		
		
			</item>
		<item>
		<title>The incredible shrinking SSRN page</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 27 Aug 2026 19:32:10 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54573</guid>

					<description><![CDATA[I was going through comments on this morning&#8217;s post which involved the amazing productivity evidenced by the list of preprints here: But then when I was checking something, I saw that something happened: The number of papers dropped from 259 &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>I was going through comments on <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/">this morning&#8217;s post</a> which involved the amazing productivity evidenced by the list of preprints <a href="https://papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=1700407">here</a>:<br />
<span id="more-54573"></span><br />
<img fetchpriority="high" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-1024x426.png" alt="" width="584" height="243" class="alignnone size-large wp-image-54571" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-1024x426.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-300x125.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-768x319.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-1536x639.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21-500x208.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-12.11.21.png 2006w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>But then when I was checking something, I saw that something happened:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-1024x417.png" alt="" width="584" height="238" class="alignnone size-large wp-image-54574" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-1024x417.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-300x122.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-768x313.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-1536x625.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47-500x203.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.29.47.png 2010w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>The number of papers dropped from 259 to 99!</p>
<p>Then when preparing this new post, I checked again:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-1024x417.png" alt="" width="584" height="238" class="alignnone size-large wp-image-54575" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-1024x417.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-300x122.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-768x313.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-1536x625.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50-500x204.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.30.50.png 2014w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>And again:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-1024x415.png" alt="" width="584" height="237" class="alignnone size-large wp-image-54576" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-1024x415.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-300x122.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-768x311.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-1536x623.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31-500x203.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-15.31.31.png 2012w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>I guess someone&#8217;s running the bot in reverse.</p>
<p>P.S.  It&#8217;s down to 44.  I wonder what the story is.</p>
<p>P.P.S.  Down to 1:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-1024x416.png" alt="" width="584" height="237" class="alignnone size-large wp-image-54577" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-1024x416.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-300x122.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-768x312.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-1536x623.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01-500x203.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-27-at-16.11.01.png 2006w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>The lone remaining paper is &#8220;Sequential Bayesian Pricing of AI Infrastructure,&#8221; posted 29 May 2026.</p>
<p>P.P.P.S.  And now the page is gone.  I guess the experiment has concluded, or the impersonation has been stopped, or something.</p>
<p>P.P.P.P.S.  According to <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/#comment-2418285">this commenter</a>, it was SSRN that took the papers down.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/feed/</wfw:commentRss>
			<slash:comments>14</slash:comments>
		
		
			</item>
		<item>
		<title>This University of Chicago business school professor has authored 258 academic papers in 2026 (so far).</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/27/258/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/27/258/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 27 Aug 2026 13:05:47 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54558</guid>

					<description><![CDATA[Jeremy Horpedahl tells the story: Nicholas Polson has, by my count using his SSRN page, already written 258 working papers in 2026 alone. He’s already written (or at least published to SSRN), six papers today, August 26, 2026. OK, but &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Jeremy Horpedahl <a href="https://economistwritingeveryday.com/2026/08/26/nicholas-polson-has-written-over-200-academic-papers-in-2026-so-far/">tells the story</a>:</p>
<blockquote><p>Nicholas Polson has, by my count using his SSRN page, already written 258 working papers in 2026 alone. He’s already written (or at least published to SSRN), six papers today, August 26, 2026.</p></blockquote>
<p>OK, but who am I to talk?&#8212;I&#8217;ve written over 200 blog posts this year.  But wait:</p>
<blockquote><p>These aren’t just short notes. Most of the papers are of normal academic length: 32 pages, 27 pages, 58 pages. . . . Obviously the research productivity of Polson and his co-author Sokolov is aided by AI. . . . I have seen any academic, at least not in economics, that has really pushed it to the limit.</p></blockquote>
<p>Horpedahl writes:</p>
<blockquote><p>Read any single paper, and it feels like just a normal academic paper, the kind of thing that an academic might work on for a few months.</p></blockquote>
<p>I wanted to see if I shared that judgment so I clicked through to <a href="https://papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=1700407">the list of Polson&#8217;s papers on SSRN</a> and looked for something interesting . . . ok, <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6278378">here&#8217;s something</a>.  It&#8217;s called <a href="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/ssrn-6278378.pdf">Theories of Human Connection</a>, and . . . ulp!  It&#8217;s 80 pages long.  The paper&#8217;s subtitle is &#8220;An Interdisciplinary Synthesis Across Economics, Psychology, Biology, Philosophy, Game Theory, and Spiritual Tradition.&#8221;</p>
<p>But let&#8217;s take a look.  The abstract on SSRN starts like this:</p>
<blockquote><p>Human connection is the most studied and least integrated phenomenon in the social sciences. Every discipline that examines intimate relationships captures something real that the others miss, yet no existing framework holds all dimensions in simultaneous view. This book synthesises fourteen thinkers across biology, psychology, economics, game theory, communication theory, existential philosophy, and spiritual tradition into a unified account. Morris established that the need for physical touch is an evolutionary drive as fundamental as hunger, with a biologically ordered sequence of escalating vulnerability whose disruption produces intensity without depth. Bowlby showed how early caregiving creates invisible templates governing adult intimacy, encoded in the nervous system before language exists. Becker revealed that partnerships generate value neither partner could produce alone, but assumed partners are interchangeable — an assumption Frankl demolishes by showing that the meaning generated by shared life cannot be transferred, and that meaning multiplies rather than adds to life satisfaction, explaining widespread disconnection in the wealthiest societies in history. Von Neumann&#8217;s game theory explains how mutual self-protective withdrawal, individually rational for each partner, produces the disconnection neither intended. Bateson identifies the communication structure that accelerates this collapse: contradictory demands that make any response wrong. Gottman&#8217;s laboratory models predict separation with over ninety percent accuracy from the ratio of positive to negative interactions. The Hindu philosophical tradition adds the final dimension: partnership as a laboratory in which selfishness and fear are progressively revealed and surrendered.</p>
<p>We argue that most relationship failures are not failures in one dimension but misidentifications of which dimension is actually in play, and that effective intervention requires the multi-dimensional map this synthesis provides.</p></blockquote>
<p>From the preface:</p>
<blockquote><p>The thinkers assembled here — Becker, Bateson, Von Neumann, Schelling, Keynes, Morris, Vaughan, Frankl, Maslow, Gottman, Yogananda, Vivekananda, Maharaj, Polson, Thomas, and Paltrow — did not, for the most part, know each other’s work. They worked in different centuries, different countries, different intellectual traditions. What unites them is that each identified a dimension of human connection that the others left in shadow.</p></blockquote>
<p>Gottman, huh?  The name rings a bell . . . <a href="https://statmodeling.stat.columbia.edu/2010/03/13/shooting_down_b/">He&#8217;s the guy who conned</a> Malcolm Gladwell and various media outlets&#8212;maybe he conned himself too&#8212;into believing that he could predict divorces with 94% accuracy.</p>
<p>I searched for Gottman in the document and found a whole chapter on that bullshit!  You can click through for yourself, it&#8217;s chapter 10.  Here&#8217;s a key bit:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-1024x242.png" alt="" width="584" height="138" class="alignnone size-large wp-image-54559" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-1024x242.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-300x71.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-768x181.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-1536x363.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59-500x118.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.16.59.png 1584w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>I wonder what prompts were used by Polson and his coauthor to write this article.  I guess the prompts did not include, &#8220;Evaluate implausible claims skeptically.&#8221;</p>
<p>They bring it all together in Chapter 13, &#8220;The Cumulative Model: A Unified Theory of Human Connection&#8221;:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-1024x781.png" alt="" width="584" height="445" class="alignnone size-large wp-image-54560" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-1024x781.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-300x229.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-768x586.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-1536x1172.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15-393x300.png 393w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.15.png 1594w" sizes="(max-width: 584px) 100vw, 584px" /><br />
<img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-1024x74.png" alt="" width="584" height="42" class="alignnone size-large wp-image-54561" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-1024x74.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-300x22.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-768x55.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-1536x110.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37-500x36.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.20.37.png 1586w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>But let&#8217;s not forget &#8220;The Master Equation&#8221; on page 75:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-1024x445.png" alt="" width="584" height="254" class="alignnone size-large wp-image-54562" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-1024x445.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-300x130.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-768x334.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-1536x668.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02-500x217.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-26-at-19.22.02.png 1596w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>Jesus Christ.  <a href="https://statmodeling.substack.com/p/cambridge-university-fraud-scandal">I&#8217;ve heard that Cambridge University has an open position in their school of education</a> . . . this kind of thing would fit in very well there, no?</p>
<p>In all seriousness, no, I don&#8217;t think this &#8220;feels like just a normal academic paper, the kind of thing that an academic might work on for a few months.&#8221;  At least, not the sort of thing a non-bullshitting academic might write.</p>
<p>I know Nick Polson&#8212;he&#8217;s a statistician, and he&#8217;s done lots of solid work over the years!  What happened here?</p>
<p>Here are a few possibilities:</p>
<p><strong>1.</strong>  It&#8217;s an experiment or a joke. But if it were a joke I&#8217;d think there&#8217;d be some internal clues, no?  It&#8217;s hard to imagine playing the whole thing straight.  And if it were an experiment, I&#8217;d expect they&#8217;d all read like straight-up statistics papers, nothing so obviously bogus as &#8220;Theories of Human Connection.&#8221;</p>
<p><strong>2.</strong>  Someone else is impersonating Polson.  Seems unlikely, but it&#8217;s possible.  It might not even be personal.  Maybe all this <em>is</em> an experiment, not by Nick but by someone else who programmed a chatbot to choose the name of a successful academic and then spew the internet with papers attributed to him.  If so, how horrible.</p>
<p><strong>3.</strong>  &#8220;Intellectual squatting.&#8221; That&#8217;s <a href="https://economistwritingeveryday.com/2026/08/22/intellectual-squatting/">the conjecture of</a> Michael Makowsky, who writes:</p>
<blockquote><p>The nice version is it’s putting out a series of half-baked papers in the hopes of establishing a property right to the underlying ideas at an earlier stage of the research process than previously possible. The less generous interpretation is it’s dumping a series of haystacks on the plains and laying claim to the needles probabilistically within each.</p></blockquote>
<p>I guess . . . but what does Nick ultimately get out of it?  Invitations to speak at more conferences??  I don&#8217;t get it.</p>
<p>Makowsky writes:</p>
<blockquote><p>Imagine you are a person who has highly esoteric, potentially important ideas every day. Many of those ideas you suspect, based on some combination of experience and ego, are new in at least one dimension. You would like to get credit for that newness. For being first. What’s the problem?</p>
<p>The problem is that scholarship remains more perspiration than inspiration. Having a new idea is great, but it takes years to work through the nuance in sufficient detail that you can convince your peers of the coherence and originality of the contribution. During the minutes each day you are not working on this singular project you have the inspiration for other ideas, sometimes multiple within a single day. How frustrating is the proposition that someone else gets credit for the originality of contribution just because they had time to reveal it to the world while you were embroiled in your investigation of what is only one of your many score ideas!?</p>
<p>Ah, but meta-level inspiration has struck you! What if you took each one of those ideas, spent an hour curating a series of prompts around it, and then let Chat GPT (or another LLM) fabricate an entire research paper around it? . . .</p></blockquote>
<p>But . . . that&#8217;s what blogging&#8217;s all about!  Often when I have an idea or a reaction, I blog it.  No need to pipe it through a chatbot; I&#8217;ll just save the cycles and post it right here.</p>
<p>To return to the 256 working papers, I can think of one more motivation:</p>
<p><strong>4.</strong>  Education.  This seems like the most plausible explanation to me.  Polson has had an active research and teaching career, and he&#8217;d like to share his insights with a broader audience than the readers of his published papers and the students at the University of Chicago business school.  And one way to reach people is . . . econ preprints!  So Nick picks 258 interesting topics, writes some prompts for each, and produces the articles.  I guess he&#8217;s programmed a bot to do this.  He just feeds it the prompts and the bot writes the paper and posts it directly to SSRN.</p>
<p>That could explain the mystery of how that ridiculous 80-page article with &#8220;The Cumulative Model: A Unified Theory of Human Connection&#8221; (shades of Stephen Wolfram!) ended up there.  Not only can&#8217;t you expect an author to write 258 articles of that length in less than a year, you can&#8217;t expect him to read all of them too.  The content of that bizarre article could be as much a surprise to Polson as it was to me.</p>
<p>This then raises a question:  setting aside the motivations of Polson (or his impersonator), do these 258 papers have any value?</p>
<p>It&#8217;s hard for me to answer this question, given that I&#8217;ve only looked at one of them.  My guess is that the net value of the papers is negative, in that the amount of time that people (including me) have wasted going through them outweighs any positive contributions that might have been there.</p>
<p><strong>My suggestion</strong></p>
<p>Here&#8217;s what Nick could do on this, which <em>could</em> have value:  Take these 258 prompts and write an article (himself, not using the chatbot) explaining why he thinks these ideas are important.  Aki and I wrote a paper a few years ago, <a href="https://sites.stat.columbia.edu/gelman/research/published/stat50.pdf">What are the most important statistical ideas of the past 50 years?</a>.  Nick could write something similar:  What are the 258 most important things in statistics to learn today?  Or something like that.  I&#8217;m not saying it would be easy&#8212;it would take more effort than programming a chatbot to spam SSRN&#8212;but valuable products often take work to produce. Nick has tenure and could set aside the time to do it.</p>
<p>Also I&#8217;d recommend withdrawing all those papers from SSRN.  Withdrawing 258 papers seems like a lot of work, but I&#8217;m sure he could easily program a bot to do the job.</p>
<p><strong>P.S.</strong>  There&#8217;s a further twist:  there are two accounts for Nicholas or Nick Polson at the University of Chicago business school; <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/#comment-2418234">see this comment thread</a>.  This would seem to be consistent with the &#8220;social experiment&#8221; hypothesis (if Nick decided to set up a separate account to play around with) or the &#8220;impersonation&#8221; hypothesis (if the bot that wrote and posted these papers was not created by Nick at all).  The whole thing remains a mystery to me.</p>
<p><strong>P.P.S.</strong>  OK, I did a bit more nosing around.</p>
<p>SSRN allows you to list the papers in time order.  If you go to Nick&#8217;s SSRN page linked from his website, you&#8217;ll see 16 papers, with the first (&#8220;The Impact of Jumps in Volatility and Returns&#8221;) being posted on 1 Jan 2001, then others through the next two decades, with the most recent being &#8220;Deep Learning in Characteristics-Sorted Factor Models,&#8221; posted on 23 Sep 2018 and last revised 26 Jun 2023.</p>
<p>If you go to the SSRN page with all the fake papers, it starts with &#8220;Kramnik vs Nakamura or Bayes vs p-value,&#8221; posted 7 Dec 2023.  It&#8217;s a badly written paper&#8212;I&#8217;m guessing not AI, just text by a non-English-speaking author that was not ever checked by native speakers before posting.  This rings a bell . . . I actually have a blog post on this paper, scheduled to appear next year.  Next on the list is a 25-page paper, &#8220;AI and Vivekananda,&#8221; posted 5 Mar 2024, then a gap of two years until another AI-related paper appeared on 9 Mar 2026, then on 11 Mar 2026 came the aforementioned &#8220;Theories of Human Connection.&#8221;</p>
<p>So, yes, Polson has two SSRN pages, but they have no overlap in time. He also has papers on Arxiv, including the <a href="https://statmodeling.stat.columbia.edu/2021/09/15/the-bayesian-cringe/">intriguingly-titled</a> &#8220;Bayes with No Shame: Admissibility Geometries of Predictive Inference,&#8221; dated 24 Aug 2026 . . . Hey, that&#8217;s just 3 days ago!  Oddly enough, I can&#8217;t find this one on SSRN.</p>
<p>But what about the article itself?  I don&#8217;t have the patience to read it, but I did catch that it mentions the martingale property, which is <a href="https://statmodeling.stat.columbia.edu/2026/06/16/the-new-york-knicks-and-the-martingale-property-of-calibrated-probability-forecasts/">one of my current interests</a>&#8212;that&#8217;s cool.  But, just flipping through, it looks much more substantive&#8212;much more like a real scientific paper&#8212;than that horrible &#8220;Theories of Human Connection&#8221; thing.  This could be a tribute to the power of modern chatbots to create something so convincing.</p>
<p><strong>P.P.P.S.</strong>  Update here:  <a href="https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/">The incredible shrinking SSRN page</a>.</p>
<p><strong>P.P.P.P.S.</strong> Vadim Sokolov, coauthor of several of the chatbot-assisted papers, <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/#comment-2418285">comments here</a>.  Assuming this comment itself is legitimate, those 258 papers were not an experiment or a hoax, nor was there any impersonator, nor was it intellectual squatting.  Rather, Polson and his coauthors just had 258 different things to say, and they felt the best way to do so was using a chatbot to write tens of thousands of words on these topics.</p>
<p>Also, according to Sokolov, they did not program a bot to do this.  He reports that at least one of the papers went through multiple rounds of review, so even if a chatbot was involved in that one, the human contribution was much more than inserting a prompt.  He says that some of the papers &#8220;are much newer and were developed with much more extensive AI assistance. Modern AI made it possible to develop and complete this material at a speed that would previously have been impossible.&#8221;</p>
<p>The fallacy in this statement by Sokolov is the idea that using a chatbot to turn a few hundred words of prompts into ten thousand words of text is a way to &#8220;complete this material.&#8221;  Nothing&#8217;s being &#8220;completed&#8221; here in any intellectual sense; the process just adds pages and pages of froth. I think that Josh&#8217;s <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/#comment-2418289">analogy</a> of this to the Sorcerer&#8217;s Apprentice is apt. </p>
<p><strong>P.P.P.P.P.S.</strong>  Polson appears to think he&#8217;s proved the Riemann hypothesis (see comments <a href="https://statmodeling.stat.columbia.edu/2026/08/27/258/#comment-2418246">here</a> and <a href="https://statmodeling.stat.columbia.edu/2026/08/27/the-incredible-disappearing-ssrn-page/#comment-2418288">here</a>).  I would think this would cause some concern among his collaborators.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/27/258/feed/</wfw:commentRss>
			<slash:comments>65</slash:comments>
		
		
			</item>
		<item>
		<title>Postdoc and doctoral student positions in Bayesian workflow at Aalto, Finland</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/27/postdoc-and-doctoral-student-positions-in-bayesian-workflow-at-aalto-finland/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/27/postdoc-and-doctoral-student-positions-in-bayesian-workflow-at-aalto-finland/#comments</comments>
		
		<dc:creator><![CDATA[Aki Vehtari]]></dc:creator>
		<pubDate>Thu, 27 Aug 2026 07:58:58 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Jobs]]></category>
		<category><![CDATA[Stan]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54566</guid>

					<description><![CDATA[This job ad is by Aki I&#8217;m looking for postdocs and doctoral students to work on Bayesian workflow. The candidates need to have knowledge of Bayesian inference and some experience with building models (for real applications, as part of methods &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/27/postdoc-and-doctoral-student-positions-in-bayesian-workflow-at-aalto-finland/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>This job ad is by Aki</p>
<p>I&#8217;m looking for postdocs and doctoral students to work on Bayesian workflow. The candidates need to have knowledge of Bayesian inference and some experience with building models (for real applications, as part of methods development, or as part of courses). Although we have published <a href="https://avehtari.github.io/Bayesian-Workflow/">Bayesian workflow book</a>, there is still a lot more to do. The focus in the group is in cross-validation, model checking and inference diagnostics (see <a href="https://users.aalto.fi/~ave/publications.html">my publication list</a>).</p>
<p>All positions are fully funded and the salaries at Aalto CS are 53k€-55k€ / year for postdocs and 40k€-45k€ / year for doctoral students. There are occupational healthcare and other benefits. Postdoc positions are typically offered for up to three years and doctoral student positions for four years. Starting dates are flexible and the details of each position will be agreed individually.</p>
<p>You can apply via <a href="https://www.ellisinstitute.fi/postdoc-and-phd-recruit-autumn-2026">joint ELLIS Institute Finland call</a> and pick me as your preferred supervisor.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/27/postdoc-and-doctoral-student-positions-in-bayesian-workflow-at-aalto-finland/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>(1) &#8220;Do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?, (2) Regarding Columbia University, &#8220;is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/26/do/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/26/do/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 26 Aug 2026 13:40:00 +0000</pubDate>
				<category><![CDATA[Decision Analysis]]></category>
		<category><![CDATA[Political Science]]></category>
		<category><![CDATA[Sociology]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53516</guid>

					<description><![CDATA[Luke Ford writes: [Regarding] the replication crisis, researcher degrees of freedom, and the gap between what statistical methods claim to establish and what they can actually support . . . Looking at the current state of the social sciences, do &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/26/do/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Luke Ford writes:</p>
<blockquote><p>[Regarding] the replication crisis, researcher degrees of freedom, and the gap between what statistical methods claim to establish and what they can actually support . . . Looking at the current state of the social sciences, do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?</p>
<p>And a question about Columbia specifically since you are there: the university has had a difficult two years in ways that have played out publicly. From your position as someone who thinks carefully about institutional incentives and measurement, what do you think the administration consistently misread, and is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?</p></blockquote>
<p>My reply:</p>
<p><strong>1.</strong>  I&#8217;m loath to give an answer about the changes in the culture of research because I have not studied this systematically.  My impression is that, yes, there&#8217;s more skepticism and less acceptance of noisy N=38 papers in psychology, etc., and less toleration for unfalsifiable evolutionary psychology and that sort of thing.  On the other hand, perhaps this has just shifted from the science establishment to social media.  Ten or fifteen years ago, there was a pipeline (partly abetted by Jeffrey Epstein) from researchers at top universities to publication in top journals to books, NPR, Ted, Gladwell, Freakonomics, etc., and lucrative speaking and consulting gigs.  So you get people like Marc Hauser or Albert-Laszlo Barabasi or Brian Wansink or Dan Ariely doing the basic research (such as it is), academic middlemen such as Steven Levitt and Cass Sunstein as promulgators, and the universities, journals, and prestige news media as part of this system (as for example here:  https://statmodeling.stat.columbia.edu/2023/08/31/the-variation-ignoring-junk-science-thats-promoted-by-association-for-psychological-science-and-related-academic-celebrities-its-like-a-poker-player-thinking-okay-if-push-all/).</p>
<p>Nowadays, though, social media runs on its own steam, and the models for academic junk science are researchers such as Andrew Huberman and Dr. Oz, who cut out the middleman and promote junk science directly, sell supplements, etc.  And social media is full of fake news and AI slop.  They don&#8217;t really need NPR, Ted, Gladwell, Freakonomics, etc., anymore; they can do it on their own.  So, in short, yes, I do have the impression that science has reformed from the bad old days of 2010-2015 (about which, see this article with Simine Vazire:  https://sites.stat.columbia.edu/gelman/research/published/jmmss-3062-gelman.pdf), but maybe the public intellectuals don&#8217;t need academic science anymore; they can just make up whatever they want on their own.</p>
<p><strong>2.</strong>  My take on Columbia is similar to my take on many institutions, which is that they have an executive function but minimal legislative or judicial functions; I discussed this here:  https://statmodeling.stat.columbia.edu/2018/01/19/lesson-charles-armstrong-plagiarism-scandal-separation-judicial-executive-functions/ and here:  https://statmodeling.stat.columbia.edu/2025/11/11/from-the-three-branches-of-government-to-the-bidirectional-nature-of-legal-reasoning-in-a-way-that-is-similar-to-how-statistics-works-and-should-work-in-the-real-world/.  As a result, their decisions are made on consequentialist rather than proceduralist gounds, and over and over again the administration takes the seemingly reasonable decision to cover up misdeeds.</p>
<p>Ford <a href="https://lukeford.net/blog/?p=178571">posted this discussion on his blog</a>.  It was kinda weird seeing myself discussed as a sociological object, but, fair enough, I&#8217;m a public figure, and people can say what they want as long as they don&#8217;t misrepresent my writings or claim that I said something I never said.</p>
<p>Ford&#8217;s assessment is accurate that I&#8217;m not very good at strategic behavior so often I don&#8217;t even try.  It&#8217;s similar to how I&#8217;m a bad negotiator so usually I&#8217;ll just try to make my goals clear and not try to optimize, following the &#8220;Getting to Yes&#8221; principle that the main thing getting in the way of smooth negotiation is ignorance of other people&#8217;s goals.  I think back to various successful and botched negotiations I&#8217;ve been involved with in the past, and almost always the problems come with struggles over details without there being clarity on the goals of the different parties.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/26/do/feed/</wfw:commentRss>
			<slash:comments>12</slash:comments>
		
		
			</item>
		<item>
		<title>Survey Statistics: more on SynthMargins and Bayes-Raking</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/25/survey-statistics-more-on-synthmargins-and-bayes-raking/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/25/survey-statistics-more-on-synthmargins-and-bayes-raking/#comments</comments>
		
		<dc:creator><![CDATA[shira]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 20:00:09 +0000</pubDate>
				<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54544</guid>

					<description><![CDATA[Last week we discussed SynthMargins, a method from the poster Modeling Complex Contingency Tables that uses partial information (margins) about poststratification variables. On theme for this series (&#8220;it is the people&#8221;), the authors commented thanking Andrew for introducing them, folks &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/25/survey-statistics-more-on-synthmargins-and-bayes-raking/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/">Last week</a> we discussed SynthMargins, a method from the poster <a href="https://www.dropbox.com/scl/fi/ypn038nijcaish4t53y9g/Kuriwaki_Shiro_Contingency-Tables-Shiro-Kuriwaki.pdf?rlkey=1pj94vlml5lzwb15libldy993&amp;dl=0" target="_blank" rel="noopener noreferrer">Modeling Complex Contingency Tables</a> that uses partial information (margins) about poststratification variables. On theme for this series (<a href="https://statmodeling.stat.columbia.edu/2025/06/01/survey-statistics-it-is-the-people/">&#8220;it is the people&#8221;</a>), the <a href="https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/#comment-2417843">authors commented</a> thanking Andrew for introducing them, folks from different fields with a shared goal: <a href="https://mgoplerud.com/">Max Goplerud</a>, <a href="https://www.shirokuriwaki.com/">Shiro Kuriwaki</a>, Jens Wiederspohn, Adam Conner-Sax, and <a href="https://www.simonsfoundation.org/people/philip-greengard/">Philip Greengard</a>.</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54549" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_6_2026_Helene_clearing-scaled.jpg" alt="" width="353" height="267" /></p>
<p><a href="https://github.com/bob-carpenter">Bob Carpenter</a> <a href="https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/#comment-2417824">shared a Stan example</a> to get a flat prior over tables that match specified margins. I think this prior would imply a prior on what the authors call alpha0, the covariance coefficients for the target geography. The method as described in the <a href="https://www.dropbox.com/scl/fi/ypn038nijcaish4t53y9g/Kuriwaki_Shiro_Contingency-Tables-Shiro-Kuriwaki.pdf?rlkey=1pj94vlml5lzwb15libldy993&amp;dl=0" target="_blank" rel="noopener noreferrer">Modeling Complex Contingency Tables</a> poster seems to focus on point estimates:</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54546" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/the_method_Modeling_Complex_Contingency_Tables.png" alt="" width="516" height="171" /></p>
<p>In other news, Shiro <a href="https://x.com/shirokuriwaki/status/2090101778616303866?s=20">responded on Twitter</a> to one of my questions about their Application 1 (ACS): &#8220;the density plot shown there is an empirical density of 1700+ TREs, where each error is for a non-Southern county.&#8221;</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54545" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/ACS_application_1.png" alt="" width="510" height="192" /></p>
<p>I have 2 remaining questions:</p>
<ol>
<li>What are the covariates w in this ACS example ?</li>
<li>Suppose you also have survey data in Palo Alto. Would this be added to the training tables ?</li>
</ol>
<p>Their Application 2 asks if lower postratification table reconstruction error improves the downstream MRP:</p>
<p class="p1"><img loading="lazy" decoding="async" class="alignnone wp-image-54547" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/application_2_Modeling_Complex_Contingency_Tables.png" alt="" width="509" height="313" /></p>
<p><a href="https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/#comment-2417849">In the comments last week, Shiro</a> cited related work by <a href="http://doi.org/10.1093/jssam/smaa008">Si and Zhou (2021)</a> who propose a method called Bayes-Raking to incorporate known margins into modeling. They found Bayes-Raking was similar to raking in the overall mean but outperformed raking for subgroups:</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54551" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_table_1.png" alt="" width="403" height="128" /></p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54550" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_figure_2.png" alt="" width="427" height="463" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_figure_2.png 1238w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_figure_2-276x300.png 276w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_figure_2-943x1024.png 943w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/bayes_raking_figure_2-768x834.png 768w" sizes="(max-width: 427px) 100vw, 427px" /></p>
<p>It would be interesting to directly compare Bayes-Raking to the poster&#8217;s SynthMargins !</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/25/survey-statistics-more-on-synthmargins-and-bayes-raking/feed/</wfw:commentRss>
			<slash:comments>9</slash:comments>
		
		
			</item>
		<item>
		<title>Bayesian Workflow free pdf!</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/25/bayesian-workflow-free-pdf/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/25/bayesian-workflow-free-pdf/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 13:00:56 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54552</guid>

					<description><![CDATA[Our wonderful new Bayesian Workflow book is now available as a free pdf! Just go the link&#8212;it&#8217;s right there! I recommend getting the hard copy too because you&#8217;ll want to be able to read it while working on the computer, &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/25/bayesian-workflow-free-pdf/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://sites.stat.columbia.edu/gelman/workflow-book/"><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/04/9780367490188_cover.jpg" alt="" width="400" /></a></p>
<p>Our wonderful new <a href="https://sites.stat.columbia.edu/gelman/workflow-book/">Bayesian Workflow book</a> is now available as a free pdf!  Just go the link&#8212;it&#8217;s right there!</p>
<p>I recommend getting the hard copy too because you&#8217;ll want to be able to read it while working on the computer, and the cost of the book is trivial compared to the benefit from faster learning that you will get by being able look at the book without taking up valuable screen real estate; also you can see connections when flipping through the pages that might not be apparent by viewing one page at a time on a screen.</p>
<p>Conversely, if you have the hard copy, you should still download the pdf because it fixes <a href="https://avehtari.github.io/Bayesian-Workflow/errata.html">a bunch of minor errors</a> that we caught after the book went to press.  Also in the printed version we accidentally repeated some of the exercises in chapters 2 and 3. For the pdf we fixed this.</p>
<p>Regarding the content, as I wrote <a href="https://statmodeling.stat.columbia.edu/2026/07/16/reviews-of-our-bayesian-workflow-book-from-bin-yu-david-spiegelhalter-brad-efron-christian-robert-and-rohan-alexander/">last month</a>, with Bayesian Data Analysis, the big steps forward were:</p>
<ul>
<li>Going beyond Bayesian inference to also consider Bayesian model building (as a researcher, you construct the model, it isn&#8217;t just given to you as in a textbook), model checking (breaking through the absolutely horrible attitude, common to Bayesians in the early 1990s, that the model was &#8220;subjective&#8221; and thus should not be checked), and model improvement (continuous model expansion, not the misguided idea of assigning posterior probabilities).</li>
<li>Going beyond simple conjugate models. BDA had lots of hierarchical models, also lots of computational tools so that you could fit the models you want by putting them together from understandable components. And I like how we had a clear separation between modeling and computing. The model comes first, then you figure out how to compute it. Or you set up a model that works within your computational constraints.</li>
<li>A Bayesian approach to sampling and causal inference. This was Rubin&#8217;s framework in which unobserved units in the population and unobserved causal outcomes are treated as missing data and are part of a joint probability model. We worked this out in chapter 7 of BDA (which became chapter 8 in the third edition of the book).</li>
<li>Lots of live examples. Not just &#8220;real-data examples,&#8221; but problems we&#8217;d directly worked on. This motivated us and I think it gave our readers a sense of how Bayesian methods worked not just in theory but in applied problems.</li>
<li>A pragmatic view of probability as a measurable quantity. That&#8217;s right there in chapter 1. Bayesian methods are not the product of a philosophical stance; they&#8217;re a way to connect models and data using probability.</li>
</ul>
<p>I could go on and on, but for that I can refer you to the <a href="https://sites.stat.columbia.edu/gelman/book/">Bayesian Data Analysis book</a>.</p>
<p>And these are the key innovations of Bayesian Workflow:</p>
<ul>
<li>Going beyond Bayesian data analysis (model building, inference, model checking, and model expansion) to consider the larger process of statistical modeling, including comparisons of multiple models fit to a single dataset.</li>
<li>A fuller use of informative priors. This is a big deal. In BDA we still had a bit of the <a href="https://statmodeling.stat.columbia.edu/2021/09/15/the-bayesian-cringe/">Bayesian cringe</a> going on. One reason we&#8217;ve moved toward stronger priors is that the replication crisis has taught us that the amount of prior information available in any given problem is often approximately the same as the information coming from an experiment (<a href="https://sites.stat.columbia.edu/gelman/research/published/default_prior_zwet.pdf">see here</a>, for example). Informative priors also fit our increased focus on generative modeling, and we&#8217;re doing a lot more prior predictive checking to understand the implications of our models.</li>
<li>More integration between modeling, data analysis, and computing. One way to see this is that the <a href="https://sites.stat.columbia.edu/gelman/workflow-book/">Bayesian Workflow webpage</a> has the code to run all our examples. We also have lots of code snippets in the text as a way of demonstrating the way in which coding is central to our statistical workflow.</li>
<li>Lots more live examples. It&#8217;s been 30 years since BDA first came out. One reason that Bayesian Workflow has 11 authors is that different collaborators worked on different examples (but the three principal authors read through the entire book, so the general approach should remain coherent).</li>
<li>Simulation-based experimentation. This is something my colleagues have been doing more and more over the years. At its most basic, simulation-based experimentation provides a best-case baseline for statistical methods: if you can&#8217;t recover your quantities of interest with sufficient accuracy under ideal conditions (when your data are simulated from the model you&#8217;re fitting), then you know you&#8217;re in trouble. And often this is the case! Beyond that, we can simulate from one model and fit another, and see what happens. Simulation experiments aren&#8217;t always so easy to construct, as they involve specifying the entire data-generation process. But we think this is effort worth expending, as it involves thinking about the problem you&#8217;re working on.</li>
</ul>
<p>I could go on and on, but for that I can refer you to the <a href="https://sites.stat.columbia.edu/gelman/workflow-book/">Bayesian Workflow book</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/25/bayesian-workflow-free-pdf/feed/</wfw:commentRss>
			<slash:comments>20</slash:comments>
		
		
			</item>
		<item>
		<title>What contributions can academic statisticians make to sports analytics?  (a discussion related to the launch of the new open-access Journal of Statistics and Data Science in Sports)</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/24/journal-of-statistics-and-data-science-in-sports/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/24/journal-of-statistics-and-data-science-in-sports/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Mon, 24 Aug 2026 13:40:19 +0000</pubDate>
				<category><![CDATA[Sports]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54534</guid>

					<description><![CDATA[Related to our recent post on the absurdity of open access fees, somebody recently informed me that a group of statisticians that work in sports decided they were sick and tired of this with the Journal of Quantitative Analysis in &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/24/journal-of-statistics-and-data-science-in-sports/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Related to <a href="https://statmodeling.stat.columbia.edu/2026/08/20/my-answer-is-no/">our recent post on the absurdity of open access fees</a>, somebody recently informed me that a group of statisticians that work in sports decided they were sick and tired of this with the Journal of Quantitative Analysis in Sports and launched a new open access journal, the <a href="https://jsds-sports.github.io">Journal of Statistics and Data Science in Sports</a>:</p>
<blockquote><p>JSDSS was founded on three core principles.</p>
<p>First, our commitment to open access is more than just lip service. At JSDSS, open access means free to read and free to publish—no exceptions. JSDSS is a Diamond Open Access journal. Research supported by academic institutions, public funding, or personal effort should not require a payment to reach the audience it deserves.</p>
<p>Second, reproducibility is essential and not an afterthought. The credibility of sports analytics research depends on our ability to verify, replicate, and build upon earlier work. JSDSS will actively incentivize transparency in data, code, and methodology, and work that meets our reproducibility standards will receive a special designation recognizing this commitment.</p>
<p>Third, we believe sport is a rich and underutilized laboratory for statistical and data science innovation. From player evaluation and in-game strategy to league design and fan engagement, sports data present compelling, real-world problems that demand rigorous and creative analytical thinking. We intend for JSDSS to be the definitive venue for this work.</p></blockquote>
<p>Cool!  I should send them something.  We think about <a href="https://statmodeling.stat.columbia.edu/category/sports/">statistics and data science in sports</a> a lot around here.</p>
<p>My only concern is that I get the impression that the cutting-edge work on sports analytics is happening outside of academia.  So I hope that this new journal can get useful contributions from people who are working in sports analytics who have material they can share without compromising their competitive advantage.</p>
<p><strong>What contributions can academic statisticians make to sports analytics?</strong></p>
<p>Or we could flip it around and ask, What contributions can academic statisticians (like me!) make to sports analytics?  Here are a few things:</p>
<p>&#8211; Developing general methods that can then be used in sports analytics (<a href="https://sites.stat.columbia.edu/gelman/research/published/stacking_paper_discussion_rejoinder.pdf">as here</a>);</p>
<p>&#8211; Writing textbooks explaining general methods that can then be used in sports analytics (<a href="https://sites.stat.columbia.edu/gelman/workflow-book/">as here</a>);</p>
<p>&#8211; Teaching students who can then work in sports analytics, or consulting on sports analytics projects, which can be thought of as a form of intense teaching;</p>
<p>&#8211; Doing work in sports analytics which, although it is not cutting edge, can still give valuable insights (<a href="https://avehtari.github.io/Bayesian-Workflow/golf/golf.html">as here</a>);</p>
<p>&#8211; Evaluation and criticisms of published work in sports-related topics (<a href="https://statmodeling.stat.columbia.edu/2016/06/28/khkhkj/">as here</a>);</p>
<p>&#8211; Contributing to the sports analytics community (<a href="https://statmodeling.stat.columbia.edu/2026/01/11/the-mets-are-hiring-2/">as here</a>, and indeed as in the present post);</p>
<p>&#8211; Collaboration on sports strategy, or sports medicine, or the sociology of sports, or various other places where sports links up with academic research;</p>
<p>&#8211; Clearing up confusion on topics related to statistics and sports (<a href="https://statmodeling.stat.columbia.edu/2015/07/09/hey-guess-what-there-really-is-a-hot-hand/">as here</a>), along with <a href="https://statmodeling.stat.columbia.edu/2024/03/15/hot-hand-the-controversy-that-shouldnt-be-and-thinking-more-about-what-makes-something-into-a-controversy/">social-sciency thinking</a> about how these misconceptions persist;</p>
<p>&#8211; Sports-related research where it can be helpful to have an outside perspective, something available to academics who aren&#8217;t on a deadline (<a href="https://statmodeling.stat.columbia.edu/2026/06/16/the-new-york-knicks-and-the-martingale-property-of-calibrated-probability-forecasts/">as here</a>).</p>
<p>There are probably some more things I didn&#8217;t think to include on this list.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/24/journal-of-statistics-and-data-science-in-sports/feed/</wfw:commentRss>
			<slash:comments>5</slash:comments>
		
		
			</item>
		<item>
		<title>Head to head on 125 St:  The Jamaican beef patty battle!</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/23/head-to-head-on-125-st-the-jamaican-beef-patty-battle/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/23/head-to-head-on-125-st-the-jamaican-beef-patty-battle/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 23 Aug 2026 22:19:49 +0000</pubDate>
				<category><![CDATA[Decision Analysis]]></category>
		<category><![CDATA[Public Health]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54539</guid>

					<description><![CDATA[My new posts here have a one-year waiting list, but I&#8217;m bumping this one up because it&#8217;s important. OK, we went on over and did a head-to-head Jamaican beef patty taste-off. As the above wrappers indicate, we compared the spicy &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/23/head-to-head-on-125-st-the-jamaican-beef-patty-battle/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_4594-768x1024.jpeg" alt="" width="300" /><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_4593-768x1024.jpeg" alt="" width="300" /></p>
<p>My new posts here have a <a href="https://statmodeling.stat.columbia.edu/2026/08/15/inbox-zero-bloglag-365/">one-year waiting list</a>, but I&#8217;m bumping this one up because it&#8217;s important.</p>
<p>OK, we went on over and did a head-to-head <a href="https://statmodeling.stat.columbia.edu/2026/08/17/jamaican-me-crazy-yet-again/">Jamaican beef patty taste-off</a>.  As the above wrappers indicate, we compared the spicy beef.</p>
<p>None of this randomization, tea-tasting crap, we just kept taking bites of each patty until they were done.</p>
<p>Both were good, but the classic Golden Krust was definitely better than the newcomer, Juici Patties.  The Golden Krust patty had a more delicious filling which was also better integrated with the crust.  By comparison, the Juici Patty had a bit too much crust and was too empty inside&#8212;it didn&#8217;t work as well as a whole.  Also the Golden Krust patty was $4.05 and the Juici was $4.25.</p>
<p>In summary:</p>
<p>1 Golden Krust patty > 1 Juici patty >>>> 1/1381 of <a href="https://statmodeling.stat.columbia.edu/2026/06/19/gray-davis-grover-norquist-and-a-rabbi-walk-into-a-conference-and-get-no-press-coverage/">a conference featuring Grover Norquist, Gray Davis, and a rabbi</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/23/head-to-head-on-125-st-the-jamaican-beef-patty-battle/feed/</wfw:commentRss>
			<slash:comments>7</slash:comments>
		
		
			</item>
		<item>
		<title>To the three different people who sent me chatbot emails today</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/23/to-the-three-people-who-sent-me-chatbot-emails-today/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/23/to-the-three-people-who-sent-me-chatbot-emails-today/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 23 Aug 2026 20:45:36 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54542</guid>

					<description><![CDATA[Just send me your goddamn prompt.]]></description>
										<content:encoded><![CDATA[<p><a href="https://statmodeling.stat.columbia.edu/2026/08/09/people-keep-sending-me-ai-slop-that-they-want-me-to-post-on-the-blog/">Just send me your goddamn prompt.</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/23/to-the-three-people-who-sent-me-chatbot-emails-today/feed/</wfw:commentRss>
			<slash:comments>11</slash:comments>
		
		
			</item>
		<item>
		<title>Failing upward, Norwegian style</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/23/failing-upward-norwegian-style/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/23/failing-upward-norwegian-style/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 23 Aug 2026 13:10:59 +0000</pubDate>
				<category><![CDATA[Political Science]]></category>
		<category><![CDATA[Public Health]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52376</guid>

					<description><![CDATA[Wendy Moore, in a review of a book by Oliver Basciano, writes: Leprosy is caused by a bacterium, Mycobacterium leprae, first identified by a Norwegian doctor, Gerhard Armauer Hansen, in 1873 . . . Convinced, wrongly, that leprosy was hereditary, &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/23/failing-upward-norwegian-style/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Wendy Moore, in a review of a book by Oliver Basciano, <a href="https://www.the-tls.com/science-technology/medicine/outcast-leprosy-oliver-basciano-book-review-wendy-moore">writes</a>:</p>
<blockquote><p>Leprosy is caused by a bacterium, Mycobacterium leprae, first identified by a Norwegian doctor, Gerhard Armauer Hansen, in 1873 . . .  Convinced, wrongly, that leprosy was hereditary, he spurred Norway to introduce the Seclusion of Lepers Act (1885), which enabled the authorities to remove people from their families to isolated leprosaria.</p></blockquote>
<p>OK, fine, everybody makes mistakes, better to err of the side of caution bla bla blah.</p>
<p>But then comes this stunner:</p>
<blockquote><p>After Hansen injected a virulent strain of the disease into the eye of one patient, Kari Nielsdatter Spidsøen, without her consent, she took him to court. He was stripped of his hospital post in 1880, but continued to oversee Norway’s leprosy policy.</p></blockquote>
<p>Whaaaa?</p>
<p>Further research (i.e., I went to the Hansen&#8217;s wikipedia page) yielded this:</p>
<blockquote><p>Hansen had attempted to infect at least one female patient with the nodular form of leprosy without consent, and although no damage was caused, the case ended up in court and Hansen lost his post at the hospital.</p></blockquote>
<p>&#8220;No damage was caused,&#8221; huh?  He just injected her in the eye, that&#8217;s all.  No harm, no foul, I guess.</p>
<p>Just amazing that he continued to run government policy after that.  Kinda reminds me of how <a href="https://statmodeling.substack.com/p/was-admiral-poindexter-a-terrorist">noted terrorist</a> John Poindexter was tasked by the U.S. government to run a terrorism prediction market.  I guess it makes sense&#8212;he was a true expert on the topic.</p>
<p>Also reminds me of that unfortunate psychology researcher who keeps coauthoring fraudulent research papers, got disciplined by MIT, and then left for a prestigious chair at Duke University, ran an advice column in a major newspaper, and had a TV show made about his research.  Actually two TV shows:  one is a highly critical documentary and another is a fictional show where a character based on him is the hero.</p>
<p>Or that political figure, I can&#8217;t remember his name now, who keeps citing discredited, fraudulent, fake, and racist research claims, and at the time of this writing he remains in charge of health policy in a major industrialized country.</p>
<p>Some people end up in prison for minor crimes.  Other people do things like inject people in the eye with a virulent strain of leprosy, and they get to make government policy.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/23/failing-upward-norwegian-style/feed/</wfw:commentRss>
			<slash:comments>13</slash:comments>
		
		
			</item>
		<item>
		<title>This is one of the worst scientific papers I&#8217;ve ever seen.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/22/this-is-one-of-the-worst-scientific-papers-ive-ever-seen/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/22/this-is-one-of-the-worst-scientific-papers-ive-ever-seen/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sat, 22 Aug 2026 13:32:42 +0000</pubDate>
				<category><![CDATA[Sports]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52904</guid>

					<description><![CDATA[A frequent commenter pointed me to this paper, &#8220;Sport and longevity: an observational study of international athletes.&#8221; All I can say is . . . Wow! This paper is an absolute clinic in bad quantitative social science research. I&#8217;ll leave &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/22/this-is-one-of-the-worst-scientific-papers-ive-ever-seen/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-1024x45.png" alt="" width="584" height="26" class="alignnone size-large wp-image-52905" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-1024x45.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-300x13.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-768x34.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-1536x67.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-2048x89.png 2048w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.38.11-500x22.png 500w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>A frequent commenter pointed me to <a href="https://link.springer.com/article/10.1007/s11357-024-01307-9">this paper</a>, &#8220;Sport and longevity: an observational study of international athletes.&#8221;</p>
<p>All I can say is . . . Wow!  This paper is an absolute clinic in bad quantitative social science research.</p>
<p>I&#8217;ll leave it as an exercise for the reader to count up all the problems.</p>
<p>The key takeaway:  For God&#8217;s sake don&#8217;t play volleyball.  It&#8217;ll reduce your life span by 5 years, and the result is statistically significant, with a p-value of 2.4e-11.</p>
<p>I used to play some volleyball.  Our team made the B-league playoffs in the MIT intramural league one year.  Now I&#8217;m scared.  Who knows what all that jumping did to our fragile young bodies.</p>
<p>Also this:</p>
<blockquote><p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.42.06-1024x170.png" alt="" width="450" /></p></blockquote>
<p>Your tax dollars at work!</p>
<p>Too bad the journal doesn&#8217;t seem to publish its reviews.  It would be hilarious to see the referee reports for this one.</p>
<p><strong>P.S.</strong>  To those of you who think I&#8217;m being mean here, &#8220;punching down,&#8221; etc., let me just say a few things.</p>
<p>1.  It&#8217;s not personal.  I know nothing about the authors of this paper and I purposely did not include their names in the post.  They may be wonderful people, indeed they may do wonderful research in other areas.</p>
<p>2.  As noted above, public funds were spent on this project.  And the paper was published, i.e. it&#8217;s there for anyone to read.  If you don&#8217;t want to be criticized, don&#8217;t publish.</p>
<p>3.  Unsupported claims get out there and they don&#8217;t go away.  For example here&#8217;s what happened when I googled *pole vaulting longevity*:</p>
<blockquote><p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-1024x411.png" alt="" width="584" height="234" class="alignnone size-large wp-image-52907" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-1024x411.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-300x120.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-768x308.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-1536x616.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-2048x821.png 2048w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/12/Screenshot-2025-12-08-at-15.49.21-500x200.png 500w" sizes="(max-width: 584px) 100vw, 584px" /></p></blockquote>
<p>This is just one silly example.  But, yeah, I do think we should be bothered by broadcasts of unsupported scientific claims, whether they&#8217;re about volleyball, himmicanes, faith healing, air rage, governors&#8217; lifespans, or anything else.  It&#8217;s bad science and it&#8217;s part of our culture of B.S.  Even if the authors of this particular paper are completely sincere in their efforts.  Remember, <a href="https://sites.stat.columbia.edu/gelman/research/published/ChanceEthics14.pdf">honesty and transparency are not enough</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/22/this-is-one-of-the-worst-scientific-papers-ive-ever-seen/feed/</wfw:commentRss>
			<slash:comments>36</slash:comments>
		
		
			</item>
		<item>
		<title>Here are the talks from StanCon 2026!</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/21/here-are-the-talks-from-stancon-2026/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/21/here-are-the-talks-from-stancon-2026/#respond</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 21 Aug 2026 13:25:35 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Stan]]></category>
		<category><![CDATA[Statistical Computing]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54529</guid>

					<description><![CDATA[StanCon 2026 just happened! And here are the talks: Matthew Kay Adrian Seyboldt Paul-Christian Bürkner Charles Margossian Javier Enrique Aguilar Sean Pinkney Nikolas Siccha Anna Dreber Kaitlyn Johnson Jonas Wallin Pranav Sanke Anna Elisabeth Riha Colling Cademartori Tim M. Szweczyk &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/21/here-are-the-talks-from-stancon-2026/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://www.stancon2026.org/">StanCon 2026</a> just happened!</p>
<p>And <a href="https://www.youtube.com/playlist?list=PLe7iuYj_G4ds">here are the talks:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-1024x810.png" alt="" width="584" height="462" class="alignnone size-large wp-image-54530" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-1024x810.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-300x237.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-768x607.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-1536x1214.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17-379x300.png 379w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.27.17.png 1986w" sizes="(max-width: 584px) 100vw, 584px" /></a></p>
<p>Matthew Kay<br />
Adrian Seyboldt<br />
Paul-Christian Bürkner<br />
Charles Margossian<br />
Javier Enrique Aguilar<br />
Sean Pinkney<br />
Nikolas Siccha<br />
Anna Dreber<br />
Kaitlyn Johnson<br />
Jonas Wallin<br />
Pranav Sanke<br />
Anna Elisabeth Riha<br />
Colling Cademartori<br />
Tim M. Szweczyk<br />
Bob Carpenter<br />
Chandler Ross<br />
Nils Rudi<br />
Fredrik Ronquist<br />
Jakob Torgander<br />
Aleksi Lahtinen<br />
Ville Laitenen<br />
Zeno Romero<br />
Soham Mukherjee<br />
Steve Bronder and Brian Ward<br />
John Ashley Burgoyne</p>
<p>Lots of great stuff here.  Check out the titles of the talks&#8212;all sorts of different topics.</p>
<p>Time to start planning for StanCon 2027.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/21/here-are-the-talks-from-stancon-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>My answer is No.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/20/my-answer-is-no/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/20/my-answer-is-no/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 21 Aug 2026 00:50:14 +0000</pubDate>
				<category><![CDATA[Decision Analysis]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54531</guid>

					<description><![CDATA[It would be great to have 35% more citations and over 5 times as many downloads, but I&#8217;d rather have 2,740 Jamaican beef patties. P.S. Here&#8217;s the article in question: The ladder of abstraction in statistical graphics. I absolutely love &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/20/my-answer-is-no/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.48.10-805x1024.png" alt="" width="584" height="743" class="alignnone size-large wp-image-54532" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.48.10-805x1024.png 805w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.48.10-236x300.png 236w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.48.10-768x977.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-20-at-20.48.10.png 1144w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>It would be great to have 35% more citations and over 5 times as many downloads, but I&#8217;d rather have <a href="https://statmodeling.stat.columbia.edu/2022/02/16/hey-i-got-an-exclusive-invitation-to-this-off-the-record-conference-but-i-think-ill-take-1907-jamaican-beef-patties-instead/">2,740 Jamaican beef patties</a>.</p>
<p><strong>P.S.</strong>  Here&#8217;s the article in question:  <a href="https://sites.stat.columbia.edu/gelman/research/published/ladder.pdf">The ladder of abstraction in statistical graphics</a>.  I absolutely love this paper.  It&#8217;s based on an idea I&#8217;ve had for awhile that I spoke on a few years ago in Ron Yurko&#8217;s statistical graphics class at CMU.  Then I wrote it up and submitted it to the journal, which gave some useful comments, and I enlisted Kaiser Fung to get it over the finish line.  Enjoy.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/20/my-answer-is-no/feed/</wfw:commentRss>
			<slash:comments>15</slash:comments>
		
		
			</item>
		<item>
		<title>One night in Uzbekistan:  Why was this one data point so influential, and what should these researchers had done ahead of time to see this?</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/20/we-couldnt-reproduce-their-findings-and-realized-that-it-was-all-driven-by-weird-data-from-uzbekistan/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/20/we-couldnt-reproduce-their-findings-and-realized-that-it-was-all-driven-by-weird-data-from-uzbekistan/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 13:20:30 +0000</pubDate>
				<category><![CDATA[Miscellaneous Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53503</guid>

					<description><![CDATA[Sol Hsiang writes: We have a comment coming out in Nature next week that is going to cause the retraction of a high-profile paper by Kotz et al. from last year (the second most cited climate paper in the news &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/20/we-couldnt-reproduce-their-findings-and-realized-that-it-was-all-driven-by-weird-data-from-uzbekistan/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img loading="lazy" decoding="async" class="alignnone size-large wp-image-54358" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-1024x317.png" alt="" width="584" height="181" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-1024x317.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-300x93.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-768x238.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-1536x475.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-2048x633.png 2048w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-04-at-16.40.02-500x155.png 500w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>Sol Hsiang writes:</p>
<blockquote><p>We have <a href="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/04/Bearpark_2025_MA.pdf">a comment coming out in Nature</a> next week that is going to cause the retraction of a <a href="https://www.nature.com/articles/s41586-024-07219-0">high-profile paper</a> by Kotz et al. from last year (the second <a href="https://www.carbonbrief.org/analysis-the-climate-papers-most-featured-in-the-media-in-2024/">most cited</a> climate paper in the news in 2024). </p>
<p>Basically, Kotz et al claimed that climate change was already costing the world economy a huge amount and would cost 300% of what prior estimates claimed (which was already large). This result got enormous attention in Europe, in particular.</p>
<p>We couldn’t reproduce their findings and realized that it was all driven by weird data from Uzbekistan. If you remove Uzbekistan from their data set, the result falls apart. The costs are still large, but not the extreme numbers that made headlines around the world.</p>
<p>One reason we think this is important is because these data were previously being used by central banks around the world to run stress tests for the effects of climate change.</p></blockquote>
<p>Here&#8217;s the retraction note, in full:</p>
<blockquote><p>The authors have retracted this paper for the following reasons: post-publication, the results were found to be sensitive to the removal of one country, Uzbekistan, where inaccuracies were noted in the underlying economic data for the period 1995–1999. Furthermore, spatial auto-correlation was argued to be relevant for the uncertainty ranges. The authors corrected the data from Uzbekistan for 1995–1999 and controlled for data source transitions and higher-order trends as present in the Uzbekistan data. They also accounted for spatial auto-correlation. These changes led to discrepancies in the estimates for climate damages by mid-century, with an increased uncertainty range (from 11–29% to 6–31%) and a lower probability of damages diverging across emission scenarios by 2050 (from 99% to 90%).</p>
<p>The authors acknowledge that these changes are too substantial for a correction, leading to the retraction of the paper. An updated version of the paper with these changes, which has yet to undergo peer review, is publicly available with continued open access to its data and methodology (https://doi.org/10.5281/zenodo.15984134). The authors intend to submit a revised version of the paper for peer review. If and when published, this retraction note will be updated to include a link to the new publication. The authors appreciate the corrective role of the global scientific community and thank Thomas Bearpark, Dylan Hogan, Solomon Hsiang and Christof Schötz for bringing these issues to their attention. All authors agree to this retraction.</p></blockquote>
<p>Good for them. And <a href="https://retractionwatch.com/2025/12/03/authors-retract-nature-paper-projecting-high-costs-of-climate-change/">here&#8217;s the story</a> in Retraction Watch.</p>
<p><strong>How science advances when data and methods are open</strong></p>
<p>Jonathan Falk independently pointed me to this story and wrote:</p>
<blockquote><p>Imagine how uphill it would have been without access to the original data/methods.</p></blockquote>
<p>Good point!</p>
<p>He also pointed to <a href="https://www.nytimes.com/2025/12/03/business/economy/study-climate-damage-retracted.html">this news article</a> which summarized the story:</p>
<blockquote><p>If Uzbekistan were excluded . . . the damages would look similar to earlier research. Instead of a 62 percent decline in economic output by 2100 in a world where carbon emissions continue unabated, global output would be reduced by 23 percent. . . .</p></blockquote>
<p>Wait&#8212;the estimate declines by almost a factor of 3 after removing just one data point? Uzbekistan&#8217;s not a tiny country but it&#8217;s not huge either (population 40 million); it doesn&#8217;t seem like its data should have so much influence as all that.</p>
<p>I went back to the original paper and it has some scatterplots, but (a) it&#8217;s hard to see that any one point would be so influential, and (b) the countries aren&#8217;t labeled so I don&#8217;t see which one is Uzbekistan.</p>
<p><strong>A question of influence</strong></p>
<p>What happened with the data? Is there some sort of scatterplot that would&#8217;ve indicated a concern?</p>
<p>To put it another way, if the data from a single medium-sized country could have that much of an impact on the findings, that would&#8217;ve been worth reporting from the get-go in the original paper. Even had there not been any data problems, we&#8217;d want to know that the results were so sensitive to one data point.</p>
<p>So the meta-question is: What data analysis should&#8217;ve been done originally, either to flag the problem with Uzbekistan&#8217;s data, or at least to reveal the extreme sensitivity of the headline results to that one data point?</p>
<p>I posed this question to Hsiang, who responded:</p>
<blockquote><p>We noticed this issue because we were looking at several papers and running some basic diagnostics on all of them. One thing we were doing was just dropping one country at a time and rerunning the models to make sure things weren’t being driving by a single country. We were surprised that this turned up. There are many issues with this paper conceptually, but it’s not even really possible to discuss any of them until you deal with the UZB issue. We had a lot of dialogue with the authors, and it turned out that they really hadn’t run much quality control on the more granular data. When we traced back this issue, it seemed like their research assistants had faithfully converted some numbers from a PDF document, but those numbers were just implausible.</p>
<p>There is <a href="https://www.nature.com/articles/s41597-023-02323-8/figures/7">a scatterplot</a> in their data paper that is supposed to provide technical validation of their data set.  We wanted to see why Uzbekistan didn’t jump out, so we reproduced it in our comment (Extended Data Fig 1). It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers (see <a href="https://www.nature.com/articles/s41586-025-09320-4/figures/2">red boxes in our version</a>). This seemed indicative of a different issue, which is why we documented it in the comment.</p></blockquote>
<p>I guess those graphs should be on the log scale?</p>
<p>It still seems crazy that the data from a single mid-sized country could have such a big effect of a global estimate.  That&#8217;s something that the original researchers should&#8217;ve been aware of, and what it suggests to me is that there is a larger methodological problem that this didn&#8217;t get looked at automatically during the research process.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/20/we-couldnt-reproduce-their-findings-and-realized-that-it-was-all-driven-by-weird-data-from-uzbekistan/feed/</wfw:commentRss>
			<slash:comments>15</slash:comments>
		
		
			</item>
		<item>
		<title>Why I am so against volunteering in academia (ecology-edition)</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/19/why-i-am-so-against-volunteering-in-academia-ecology-edition/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/19/why-i-am-so-against-volunteering-in-academia-ecology-edition/#comments</comments>
		
		<dc:creator><![CDATA[Lizzie]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 20:37:21 +0000</pubDate>
				<category><![CDATA[Miscellaneous Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54523</guid>

					<description><![CDATA[This post is by Lizzie. The photo is from this summer and included for no other reason than because I like a photo with a post.  A few months ago I wrote a post where I compared the toxic culture &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/19/why-i-am-so-against-volunteering-in-academia-ecology-edition/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm.jpg"><img loading="lazy" decoding="async" class="alignnone size-large wp-image-54522" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm-1024x768.jpg" alt="" width="584" height="438" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm-1024x768.jpg 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm-300x225.jpg 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm-768x576.jpg 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm-400x300.jpg 400w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/IMG_20260713_173841185_HDRsm.jpg 1138w" sizes="(max-width: 584px) 100vw, 584px" /></a></p>
<p><em>This post is by Lizzie. The photo is from this summer and included for no other reason than because I like a photo with a post. </em></p>
<p>A few months ago I wrote a post where I compared the toxic culture of the restaurant industry to the toxic culture of some parts of my academic world in ecology. One point of similarity was how you often need to &#8216;volunteer&#8217; your time to get a foot in the prestigious door. This led to <a href="https://statmodeling.stat.columbia.edu/2026/03/31/black-and-white-gray-and-in-between-what-color-is-the-media/#comment-2412703" target="_blank" rel="noopener">a query by Phil about why I am against volunteering</a>. Thanks to Phil for asking the query, and others for already effectively giving my general answer, but I will write a full post here since I think it is a good question.</p>
<p>So, why I am so against volunteering in academia? And here I am focusing on my part of academia, which is ecology. In short: I find volunteering in ecology in my world to be an exclusionary practice practiced by those who talk endlessly about inclusion. And, as is often the case in my life, this sort of hypocrisy drives me nuts and I can just never get over it.</p>
<p>And for the longer version&#8230;.</p>
<p>In my field (ecology) and related fields (evolution, conservation), there is a pervasive assumption that it is a-okay to hire/have &#8216;volunteers&#8217; to get your work done. I say hire, because these positions are advertised all the time with a list of qualifications you&#8217;ll need. Here&#8217;s one:</p>
<ul>
<li>BSc degree in a wildlife, environmental or conservation topic or in the process of completing one.</li>
<li>Intermediate level in English and Spanish (Oral and Written).</li>
<li>Knowledge in wildlife monitoring surveys. Previous research experience in any capacity is a plus.</li>
<li>Physically fit and able to work long hours in a difficult and harsh environment.</li>
<li>Good team member with excellent communication skills; able to live and work with a multicultural team.</li>
<li>Able to live in basic living conditions and tropical rainforest conservation campus.</li>
<li>Hard working and passionate with a desire to learn and improve; willing to put the hard work in to go the extra mile for conservation efforts and personal career development.</li>
<li>Excellent computer skills with a confidence in all Microsoft and Google programs.</li>
</ul>
<p>Conservation organizations or conservation-related research positions seem to feel especially allowed to do this, with the argument that their mission somehow absolves them of basic labor laws. The ones that maybe annoy me the most are the ones for ornithology work (that&#8217;s a fancy word for the study of birds) where you also need to pay your way to some (often tropical) locale and maybe even your housing and food so you can help slosh around in the jungle and search for birds. If you don&#8217;t believe me, most days you can set the Job Type to Volunteer/Training and search for &#8216;bird&#8217; <a href="https://jobs.rwfm.tamu.edu/" target="_blank" rel="noopener">here</a> and you&#8217;ll likely find something like this somewhere:</p>
<blockquote><p>This is an unpaid position and interns are responsible to cover food and accommodation at the biological station [in Peru/Belize/etc.].</p></blockquote>
<p>I learned about all this when I was on a &#8216;field crew&#8217; long ago (that&#8217;s a term in ecology for a lot of people working together on some project where they go outside every day to collect data from the natural world). I was doing my PhD on bugs, but everyone else was studying the birds that eat the bugs and they explained to me that usually to get a paying bird-crew job you must first pay your way (housing etc. wherever the field crew is based) to be trained in point counts (standing in one place for a set amount of time and listing all the birds you hear) and then maybe someone pays your room and board and you learn to do &#8216;nest-finding&#8217; (self-evident definition) and, as you get more and more training, you might get a small allowance (it&#8217;s an &#8216;allowance&#8217; because it is generally way below minimum wage if you do out the per hour rate).</p>
<p>And then maybe you get into grad school to study birds. Lucky you!</p>
<p>And who knows how marine biology works (&#8220;Junior scientists in marine mammalogy are expected to have at least one or two unpaid research experiences to qualify for a graduate programme&#8221; according to a <a href="https://www.nature.com/articles/d41586-020-02758-8" target="_blank" rel="noopener"><em>Nature Careers</em> piece in 2020</a>). Effectively, as best I can tell, the more coveted and sexy the position, the more it depends on extreme levels of unpaid work. I have never seen the data but when I look around at my ornithology and marine biology colleagues I wonder if the t-test on their socioeconomic status before they entered the ornithology/marine biology is higher than other subdisciplines within ecology.</p>
<p>And then the exact same disciples publish articles about how we need more diversity in the sciences. It blows my mind.</p>
<p>For decades people have been pointing out the problem:</p>
<p>Whitaker 2003 <a href="https://www.researchgate.net/publication/227987509_The_Use_of_Full-Time_Volunteers_and_Interns_by_Natural-Resource_Professionals" target="_blank" rel="noopener">The Use of Full-Time Volunteers and Interns by Natural-Resource Professionals </a>(Conservation Biology, Vol. 17, No. 1) explained that most often</p>
<blockquote><p>full-time [unpaid] positions are filled by aspiring natural-resource professionals (e.g., students or recent graduates) who need workplace experience if they are to advance to graduate school or more lucrative jobs. Consequently, employers often consider work experience adequate compensation for wage shortfalls.</p></blockquote>
<p>&#8230; and continues:</p>
<blockquote><p>I believe that our widespread use of volunteers and interns to compensate for budget shortfalls does our profession more harm than good. In many cases this approach is in conflict with labor law, hinders the development of new professionals, undermines our profession&#8217;s credibility, and is an impediment to achieving our conservation goals.</p></blockquote>
<p>He then argues that they these positions cause personal hardship, exclude certain groups, and does not meet societal norms:</p>
<blockquote><p>Our perception of the importance and urgency of our work as conservationists does not elevate us above the societal values that led to [labor laws].&#8217;), undervalues conservation and those working in this and related areas.</p>
<p>Our current strategy of cutting wages does not cut costs; rather, it transfers them directly to the lowest tier of professionals in the form of financial and emotional hardship.</p></blockquote>
<p>Despite this article from 20+ years ago my discipline continues to do this, now at the same time they decry the need for greater diversity in the field. Most articles on how to increase diversity preach for wider acceptance, statements of inclusivity, but not much action on this practice. That said, some articles acknowledge this as a problem we should maybe change if we want to increase diversity. They all cite Fournier &amp; Bond 2015 <a href="https://wildlife.onlinelibrary.wiley.com/doi/full/10.1002/wsb.603" target="_blank" rel="noopener">Volunteer Field Technicians Are Bad for Wildlife Ecology</a> (Wildlife Society Bulletin 39(4):819–821, which is good but I think the fact that they don&#8217;t cite Whitaker means they did not try very hard to look into this issue), which to me gives the 3 reasons folks usually give for not paying people who work for them:</p>
<ol>
<li>I don&#8217;t have enough money to pay them.</li>
<li>It&#8217;s what has always happened and &#8230;</li>
<li>pointing out other worse-treated/compensated workers.</li>
</ol>
<p>Fournier &amp; Bond 2015 also makes the general argument of why we should not have unpaid interns, with the Biden quote:</p>
<blockquote><p>Don’t tell me what you value. Show me your budget, and I’ll tell you what you value.</p></blockquote>
<p>I find (1) pretty common, but (2) is also pervasive. I have tried talking to my colleagues who are very well funded about this and they tell me that this as just &#8216;how science is done.&#8217; They often explain to me that a student gets a letter of reference from them out of it, so what&#8217;s the problem?</p>
<p>I think once it&#8217;s engrained that you can get free grunt work, it&#8217;s hard to get people to give it up.</p>
<p>And folks are trained in it early. Graduate students are often sent out to get others to help with their grunt work. I see posts on <a href="https://esa.org/membership/ecolog/" target="_blank" rel="noopener">ECOLOG</a> and flyers in my hallway each year explaining who will be needed for a full-time unpaid position to collect soils, or leaf samples or pipette all day long.</p>
<p>What do I think everyone should do? I suggest they start by what I have managed to do: I just don&#8217;t allow &#8216;volunteers&#8217; in my lab and I tell everyone why I don&#8217;t allow volunteers. I pay everyone who works in my lab (including undergraduates) and I make sure part of their paid word is training (how to use Git R, sometimes Stan). Is it expensive to pay every undergrad in my lab minimum wage (or better) given my super low budget each year? Hell yes. But that is not an excuse to me &#8212; if you cannot afford to get the work done without volunteers, then you cannot afford to get the work done.</p>
<p>Everyone is clearly not going to just do this though, so we could all benefit from some top-down help I suspect. Granting agencies could ask about how much a lab relies on volunteers and discourage it through various mechanisms. Instead, they often encourage it. In Canada, NSERC grades me on how many students I have &#8216;trained&#8217; (AKA: have passed through my lab) so finding undergraduate volunteers to stock my lab would certainly help the numbers game. Further, NSERC cares about how much you have volunteered in MSc/PhD fellowships, effectively encouraging students to get on board with this practice.</p>
<p>I think NSERC means to support short-term volunteering, which gets back to Phil&#8217;s question of &#8216;when is this okay?&#8217; In part he seems to ask if it is okay to take a minimum wage job on a rec crew or such. To which I say, of course! We&#8217;re all welcome to take low-paying jobs if we can get them if you ask me, especially as they are often the gateway to a career change and I support career changes. So the question is more: when is okay to work for free? I work on community-science projects, such as the <a href="https://usanpn.org/" target="_blank" rel="noopener">USA-NPN</a>, where data are entirely collected by volunteers and I see how volunteering in this way or in other ways that help you connect with something beyond your job, give back to community or provide aid are important for a multitude of reasons.</p>
<p>So my focus here is really on full-time or similar positions. Whitaker 2003 makes the distinction between &#8216;part-time or short-term service&#8217; where you do not have to forego outside opportunities for paid work vs. &#8220;other volunteer and internship positions require that individuals live in remote areas and work&gt;=40 hours per week, effectively denying them the opportunity to earn outside wages&#8221; and I think this is very good perspective &#8212; as it can be extended also cover the undergraduates working in my lab 10 hours/week during term (they do not forego getting part-time wages, which many students need to get through university).</p>
<p>But we should still be thoughtful of the diversity problems volunteering often leads to. Community scientists are often retired and in a good socioeconomic position, and this extends through many other places. If the National Parks Service relies on volunteers to help repair trails etc., I am sure they will have a bias in who signs up for those opportunities and that may not be the best thing for the NPS.</p>
<p>And while I am back on my diversity angle, if you are thinking it is okay for granting agencies like NSERC to ask about volunteering with something along the lines of &#8216;but don&#8217;t worry! If your life is too hard to volunteer, you can just explain that now and we&#8217;ll value it,&#8217; then try harder.</p>
<p>Try harder to think through what type of person is the one who will know they just need to profligate themselves and tell all their horrible life problems so the nice middle- and upper-class people on the committee can give them credit for their difficulties. Maybe you get a small slice of people who that works for, but more I think you get several predictable outcomes. You get people <a href="https://www.pbs.org/newshour/show/miracle-children-explores-admissions-scandal-that-exposed-inequalities-in-education" target="_blank" rel="noopener">lying</a> about it. And you get people who are horrified and scarred by having to do it (IMHO it&#8217;s a revolting ask if you back up and think about), or just don&#8217;t do it. So while I am on the topic of suggesting we all stop asking people to volunteer to do our work, we should also stop asking people to tell us how hard their lives were.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/19/why-i-am-so-against-volunteering-in-academia-ecology-edition/feed/</wfw:commentRss>
			<slash:comments>114</slash:comments>
		
		
			</item>
		<item>
		<title>What does it mean when different articles published by different people on the same topic have nearly identical titles?  (&#8220;A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study&#8221;)</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/19/wh/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/19/wh/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 13:50:34 +0000</pubDate>
				<category><![CDATA[Miscellaneous Science]]></category>
		<category><![CDATA[Public Health]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53493</guid>

					<description><![CDATA[Sanjeev Sripathi writes: I came across a firm that seems to be offering a miracle product &#8211; https://letsmoderate.com/products/sugar-slayer . Their claim is that ingesting it hammers your post-meal blood sugar spike down by 40%, and they have multiple studies to &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/19/wh/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Sanjeev Sripathi writes:</p>
<blockquote><p>I came across a firm that seems to be offering a miracle product &#8211; https://letsmoderate.com/products/sugar-slayer . Their claim is that ingesting it hammers your post-meal blood sugar spike down by 40%, and they have multiple studies to back it up: https://letsmoderate.com/pages/science</p>
<p>What stood out to me as weird:</p>
<p>(1) These three articles were published at different times by different people but follow the same format for the abstract and text and arrive at nearly the same very large effect size for the same product (the one being sold): https://www.mdpi.com/2072-6643/16/14/2237 , https://link.springer.com/article/10.1186/s41110-024-00275-6 , https://link.springer.com/article/10.1186/s41110-024-00294-3</p>
<p>(2) The other listed articles seem to have been fetched by doing a search for &#8216;mulberry extract&#8217;, which is the critical active ingredient. Whatever I can access does show a notable attenuation on blood glucose and studies exist all the way from 2007 to 2022, thought the impact varies a lot.</p>
<p>(3) These 3 studies are eerily similar to the first 3 I&#8217;ve mentioned here: https://www.liebertpub.com/doi/10.1089/jmf.2014.3160, https://link.springer.com/article/10.1186/s12986-021-00571-2, https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0172239#authcontrib . But they&#8217;re from 2007, 2015 and 2024, across 3 sets of researchers, which just seems very odd. Is this how everyone is meant to name their papers? i.e. is there a social pressure sitting outside these groups pushing them into common convention?</p>
<p>I&#8217;m sure you&#8217;d identify more elements if you took a look at it. I&#8217;m just unclear if I&#8217;m jumping at shadows.</p></blockquote>
<p>I was curious so I clicked on the link.  From the letsmoderate site:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-1024x303.png" alt="" width="584" height="173" class="alignnone size-large wp-image-53494" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-1024x303.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-300x89.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-768x227.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-1536x455.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-2048x606.png 2048w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2026-03-31-at-19.59.41-500x148.png 500w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>Here are the titles of the first three articles listed in the above email:</p>
<blockquote><p>A Novel Nutraceutical Supplement Lowers Postprandial Glucose and Insulin Levels upon a Carbohydrate-Rich Meal or Sucrose Drink Intake in Healthy Individuals—A Randomized, Placebo-Controlled, Crossover Feeding Study</p>
<p>A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study</p>
<p>A ready-to-mix nutraceutical supplement (GLUBLOC) lowers postprandial blood glucose levels in healthy individuals — A randomised, placebo-controlled, crossover study</p></blockquote>
<p>All the above authors are from India but with no overlap in the author list.</p>
<p>And here are the titles for the other three:</p>
<blockquote><p>Mulberry Leaf Extract Improves Postprandial Glucose Response in Prediabetic Subjects: A Randomized, Double-Blind Placebo-Controlled Trial</p>
<p>Mulberry leaf extract improves glycaemic response and insulaemic response to sucrose in healthy subjects: results of a randomized, double blind, placebo-controlled study</p>
<p>Mulberry-extract improves glucose tolerance and decreases insulin concentrations in normoglycaemic adults: Results of a randomised double-blind placebo-controlled study</p></blockquote>
<p>The first one&#8217;s from Korea, and the second and third are from England, with some overlap in the author list.</p>
<p>I agree that it&#8217;s weird that the papers have nearly identical trials.  But I don&#8217;t know how things go in the medical literature.  Maybe this is standard practice?</p>
<p>I&#8217;m not planning to fork over 799 rupees for 30 tablets of Sugar Slayer.  One of the article says, &#8220;Alkaloid- and polyphenol-rich white mulberry leaf and apple peel extracts have been shown to have potential glucose-lowering effects, benefitting the control of postprandial blood glucose levels,&#8221; so maybe I&#8217;ll just eat an apple.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/19/wh/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>Update on a regression discontinuity dispute:  some asynchronous collaboration</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/18/update-on-a-regression-discontinuity-dispute-some-asynchronous-collaboration/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/18/update-on-a-regression-discontinuity-dispute-some-asynchronous-collaboration/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 02:01:31 +0000</pubDate>
				<category><![CDATA[Causal Inference]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54520</guid>

					<description><![CDATA[Anjali Thomas writes: I am writing to share a paper which is a re-examination of my 2018 AJPS article entitled “Targeting Ordinary Voters or Political Elites”  which was previously discussed on this blog here. The paper, written in the spirit of &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/18/update-on-a-regression-discontinuity-dispute-some-asynchronous-collaboration/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Anjali Thomas writes:</p>
<blockquote>
<div>I am writing to share a paper which is a re-examination of <a href="https://urldefense.com/v3/__https://onlinelibrary.wiley.com/doi/abs/10.1111/ajps.1237)__;!!BDUfV1Et5lrpZQ!Uo5vtbeXnwmRjubOoRP3eU1G5iz-H6j7z-C9U36SA5CfB1pm8VL13m3x3x2WddzgXTW2tsPguvL1_xccuFaSkGXC6Co$" data-outlook-id="5a489930-5ca8-43a0-a10d-453a1bf6c5a9">my 2018 AJPS article entitled “Targeting Ordinary Voters or Political Elites”</a>  which was previously discussed on this blog <a href="https://statmodeling.stat.columbia.edu/2018/08/02/38160/" data-outlook-id="e581908f-f7fe-4b00-8dc3-00ae5d2e1e99">here.</a></div>
<p>The paper, written in the spirit of <a href="https://sites.stat.columbia.edu/gelman/research/published/causal_paths_3.pdf">Gelman (2022)</a>, conducts a thorough re-analysis of my earlier work in light of recent developments in regression discontinuity designs (RDD). It also directly addresses specific critiques of the article raised both in subsequent academic literature and <a href="https://statmodeling.stat.columbia.edu/2018/08/02/38160/" data-outlook-id="903990c7-be46-4f06-b312-81194bc43755">in previous comments on this blog</a>.</p>
<div>A full response to each critique raised on this blog appears in Section A.3 on page 48 of <a href="https://urldefense.com/v3/__https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598__;!!BDUfV1Et5lrpZQ!Uo5vtbeXnwmRjubOoRP3eU1G5iz-H6j7z-C9U36SA5CfB1pm8VL13m3x3x2WddzgXTW2tsPguvL1_xccuFaSD1rmCB0$" data-outlook-id="dddee0ef-515d-4ba9-8057-959cafb58406">the paper</a>. Among other things, the blog critiqued the use of the global fourth order polynomial, and commenters highlighted that it appeared to be picking up noise in the data rather than a true relationship. While the original article did present results in the SI showing robustness to a local-linear specification with alternative bandwidths, the current<a href="https://urldefense.com/v3/__https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598__;!!BDUfV1Et5lrpZQ!Uo5vtbeXnwmRjubOoRP3eU1G5iz-H6j7z-C9U36SA5CfB1pm8VL13m3x3x2WddzgXTW2tsPguvL1_xccuFaSD1rmCB0$" data-outlook-id="8a929b31-c90b-45ff-912b-6276997d0f13">paper</a><a href="https://urldefense.com/v3/__https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598__;!!BDUfV1Et5lrpZQ!Uo5vtbeXnwmRjubOoRP3eU1G5iz-H6j7z-C9U36SA5CfB1pm8VL13m3x3x2WddzgXTW2tsPguvL1_xccuFaSD1rmCB0$" data-outlook-id="0ea3f367-78c5-4025-b91f-88763c1f9e46"> </a>significantly extends these checks and presents new results showing:</div>
<ul>
<li>
<p role="presentation"><b>Polynomial &amp; Bandwidth Robustness:</b> Shows stability across lower-order polynomials, alternative bandwidths, kernel choices, and a donut-hole approach.</p>
</li>
<li>
<p role="presentation"><b>Noise Reduction via Aggregation:</b> Aggregates data to the level of the running variable to reduce noise, presenting new scatter plots where the visual discontinuity persists at this level of aggregation.</p>
</li>
<li>
<p role="presentation"><b>Spatial Structure:</b> Demonstrates that reported balance/placebo issues in recent re-analyses stem from ignoring within-constituency clustering and predictive controls.</p>
</li>
</ul>
<div>A link to the paper is here (<a href="https://urldefense.com/v3/__https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598__;!!BDUfV1Et5lrpZQ!Uo5vtbeXnwmRjubOoRP3eU1G5iz-H6j7z-C9U36SA5CfB1pm8VL13m3x3x2WddzgXTW2tsPguvL1_xccuFaSD1rmCB0$" data-outlook-id="b346133b-b3a7-4b86-acb2-cf384886df02">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598</a>) and the abstract is below:</div>
<blockquote><p>This note reexamines Thomas (2018), which advances and tests the &#8216;elite cooperation logic&#8217; whereby national politicians target resources along partisan lines to win over the cooperation of co-partisan state legislators in implementing development projects. Consistent with this logic, a close election regression discontinuity design (RDD) uses project-level data to show that national legislators in North India allocate systematically higher public works expenditures to constituencies of co-partisan state legislators in the period after a state election. Responding to critiques in subsequent literature of the RDD approach used, this note shows that both the evidence of covariate imbalance reported in Bicalho et al. (2026) and the high proportion of significant placebo estimates reported in Albada (2025) are artefacts of ignoring within-constituency clustering and omitting predictive controls. Meanwhile, this note presents re-analyses confirming that the core findings in Thomas (2018) are robust to covariate inclusion using either constituency-level clustering or constituency-level aggregation. Acknowledging the problems related to global fourth order polynomials (Gelman and Imbens, 2019; Albada, 2025), the results in Thomas (2018) are also shown to be robust to lower-order polynomials, alternative bandwidths, alternative kernel choices, the donut hole approach, and an inference procedure that adjusts for worst-case bias (Stommes et al., 2023). The note re-establishes the credibility of the substantive findings in Thomas (2018) and highlights the importance of accounting for spatial clustering in both estimation and diagnostic checks in RDDs.</p></blockquote>
</blockquote>
<p>Anjali took my Bayesian statistics class back in 2006!  It&#8217;s great to see what former students are doing, and I love seeing this sort of asynchronous collaboration.</p>
<p>Regarding the regression discontinuity issues, I do not think it makes sense to fit unrelated curves on the two sides of the cutoff.  So I prefer the versions that fit a single curve plus discontinuity, not two curves.</p>
<p>The other thing I always recommend (see <a href="https://sites.stat.columbia.edu/gelman/research/published/JCRE-2025-12-Gelman_and_Imbens.pdf">this recent paper with Imbens</a> is that regressions include not just the forcing variable but also other pre-treatment predictors.  This should help with both bias and efficiency.</p>
<p>All the discussion of bandwidth, functional forms, etc., can be wasted if the fitted model makes no sense or if it&#8217;s a distraction from including additional pre-treatment predictors.</p>
<p>In any case, it&#8217;s great to see this sort of open exploration.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/18/update-on-a-regression-discontinuity-dispute-some-asynchronous-collaboration/feed/</wfw:commentRss>
			<slash:comments>16</slash:comments>
		
		
			</item>
		<item>
		<title>Survey Statistics: Modeling Complex Contingency Tables</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/#comments</comments>
		
		<dc:creator><![CDATA[shira]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 20:00:31 +0000</pubDate>
				<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54514</guid>

					<description><![CDATA[Andrew looped me into an email thread with folks working on poststratification with partial population information. (He knew I&#8217;d be interested, see &#8220;poststratification without population level information&#8221;.) They pointed to work by Max Goplerud, Shiro Kuriwaki, Jens Wiederspohn, Adam Conner-Sax, &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Andrew looped me into an email thread with folks working on poststratification with partial population information. (He knew I&#8217;d be interested, see <a href="https://statmodeling.stat.columbia.edu/2026/07/21/survey-statistics-poststratification-without-population-level-information/">&#8220;poststratification without population level information&#8221;</a>.)</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54516" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_on_pack_view_2-scaled.jpg" alt="" width="358" height="271" /></p>
<p>They pointed to work by <a href="https://mgoplerud.com/">Max Goplerud</a>, <a href="https://www.shirokuriwaki.com/">Shiro Kuriwaki</a>, Jens Wiederspohn, Adam Conner-Sax, and <a href="https://www.simonsfoundation.org/people/philip-greengard/">Philip Greengard</a>, which was recently presented as poster at <a href="https://polmeth.msu.edu/program">polmeth</a>: <a href="https://www.dropbox.com/scl/fi/ypn038nijcaish4t53y9g/Kuriwaki_Shiro_Contingency-Tables-Shiro-Kuriwaki.pdf?rlkey=1pj94vlml5lzwb15libldy993&amp;dl=0" target="_blank" rel="noopener noreferrer">Modeling Complex Contingency Tables</a>.</p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-54515" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster.png" alt="" width="1728" height="1288" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster.png 1728w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster-300x224.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster-1024x763.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster-768x572.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster-1536x1145.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Goplerud_et_al_2026_poster-402x300.png 402w" sizes="(max-width: 1728px) 100vw, 1728px" /></p>
<p>I wish I got to attend, this is cool ! I&#8217;ve got questions about Application 1 (ACS):</p>
<ol>
<li>What are the covariates w?</li>
<li>Suppose you also have survey data in Palo Alto. Would this be added to the training tables ?</li>
<li>How is the table reconstruction error random ? (I see a density plot.)</li>
</ol>
<p>The poster gives an answer to Andrew&#8217;s question in <a href="https://statmodeling.stat.columbia.edu/2022/01/23/mister-p-when-you-dont-have-the-full-poststratification-table-you-only-have-margins/">&#8220;Mister P when you don’t have the full poststratification table, you only have margins&#8221;</a> and <a href="https://statmodeling.stat.columbia.edu/2023/12/31/the-continuing-challenge-of-poststratification-when-we-dont-have-full-joint-data-on-the-population/">&#8220;The continuing challenge of poststratification when we don’t have full joint data on the population&#8221;</a>:</p>
<blockquote>
<p>I’d recommend first imputing a full poststrat table &#8230; But then the question is how to do this.</p>
</blockquote>
<p>In Application 2, they ask if a better poststratification table reduces error in estimating the outcome using <a href="https://statmodeling.stat.columbia.edu/2025/06/24/survey-statistics-poststratification/">Multilevel Regression and Poststratification (MRP)</a>. Presumably this depends on the distribution of the outcome given the poststratification variables. In <a href="https://statmodeling.stat.columbia.edu/2026/07/07/survey-statistics-toy-example-for-energy-balancing-weights/">&#8220;toy example for energy balancing weights&#8221;</a> I wrote that raking does well &#8220;when Y | X1, X2 is additive&#8221;.</p>
<p>I&#8217;m excited to read the paper and learn more !</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/18/survey-statistics-modeling-complex-contingency-tables/feed/</wfw:commentRss>
			<slash:comments>16</slash:comments>
		
		
			</item>
		<item>
		<title>The improvement in political analysis in the past 25 years, as demonstrated by excellent demonstrations of statistical workflow from Elliott Morris, Nate Silver, and Eli Mckown-Dawson</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/18/modern-day-political-analysis-excellent-demonstrations-of-statistical-workflow-from-elliott-morris-nate-silver-and-eli-mckown-dawson/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/18/modern-day-political-analysis-excellent-demonstrations-of-statistical-workflow-from-elliott-morris-nate-silver-and-eli-mckown-dawson/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 13:37:40 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54500</guid>

					<description><![CDATA[As with baseball, football, and basketball (and I&#8217;m sure other sports too), the standard of political analytics is just so much higher than it was, decades ago. I was talking with Gustavo just the other day about Red State Blue &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/18/modern-day-political-analysis-excellent-demonstrations-of-statistical-workflow-from-elliott-morris-nate-silver-and-eli-mckown-dawson/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>As with baseball, football, and basketball (and I&#8217;m sure other sports too), the standard of political analytics is just so much higher than it was, decades ago.</p>
<p>I was talking with Gustavo just the other day about Red State Blue State, and how that work was motivated by confusion following the 2000 election emanating from pundits of the left, right, and center.  Back then I felt the compulsion to write a whole damn book to explain what was really going on.  I even came up with an entirely new (to the best of my knowledge) concept, &#8220;second-order availability bias,&#8221; <a href="https://statmodeling.stat.columbia.edu/2005/10/17/red_and_blue_vo/">to explain how the journalists</a> could&#8217;ve gotten things so wrong.</p>
<p>The concept of &#8220;second-order availability bias&#8221; never caught on, to say the least:  it appears only once in the easily-accessed published literature:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-1024x283.png" alt="" width="584" height="161" class="alignnone size-large wp-image-54501" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-1024x283.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-300x83.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-768x213.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-1536x425.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57-500x138.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-17.44.57.png 1626w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>So maybe it&#8217;s not such a useful psychological concept.  What&#8217;s relevant here, though, is that the pundits were getting it so wrong, and with such a consensus, that I felt the need to refute them.</p>
<p>Nowadays, things are different.  There aren&#8217;t so many all-purpose pundits like David Brooks&#8212;people who know essentially nothing and have no real interest in learning but present themselves as infallible experts&#8212;, and those who remain don&#8217;t have such a platform.  Also, with partisan polarization, commentary has become fragmented, and we rarely see much of a consensus among pundits across the political polarization.  It&#8217;s just not on the table.</p>
<p>Meanwhile, political analytics has become more and more impressive, with important contributions being made by academics, journalists, and political professionals.  There&#8217;s still disagreement (<a href="https://statmodeling.stat.columbia.edu/2025/04/16/did-non-voters-really-flip-republican-in-2024/">as here</a>) and some difficulties of communication (<a href="https://statmodeling.stat.columbia.edu/2024/01/14/and-while-i-dont-really-want-a-back-and-forth/">as here</a>), but in the past two decades the level has gone up so much:  the best analytics has become much more impressive, and what might be called replacement-level analytics has become much better too.  I don&#8217;t think my own sophistication has increased much at all, and that&#8217;s one reason why I&#8217;m now less likely to crunch the numbers myself (as I did in the <a href="https://statmodeling.stat.columbia.edu/2008/11/05/election-2008-what-really-happened/">early morning hours of 5 Nov 2008</a>) and more likely to just link to the analyses of others (as with <a href="https://statmodeling.stat.columbia.edu/2025/05/21/what-happened-in-2024/">Yair&#8217;s report on 2024</a>).</p>
<p>Just today I came across two excellent examples online from journalist colleagues of mine.</p>
<p><strong>Elliott Morris</strong>, &#8220;<a href="https://www.gelliottmorris.com/p/2026-08-13-wisconsin-poll-reanalysis">I re-analyzed the raw data from Wisconsin&#8217;s primary polls. Here&#8217;s what actually went wrong</a>,&#8221; which features this split-the-difference summary that warms <a href="https://statmodeling.stat.columbia.edu/2024/11/03/crude-bayesian-updating-from-iowa-election-poll/">my Bayesian heart</a>:<br />
• Most of the miss in polls in Wisconsin is attributable to faulty demographic targets (too many young people). This inflated Hong’s vote margin by somewhere between 5 and 10 points.<br />
• My best guess is that the race moved 6-10 points toward Crowley after pollsters released their final surveys.<br />
• Non-ignorable non-response within demographic categories likely further inflated Hong’s vote margin by 2-5 points.<br />
Morris goes through lots of details too.  I haven&#8217;t tried to check any of this, but it seems reasonable.  <a href="https://archive.nytimes.com/campaignstops.blogs.nytimes.com/2011/11/29/why-are-primaries-hard-to-predict/">We&#8217;ve been saying for a long time</a> that primary elections are hard to predict, but some polls are off by much worse than others, and it&#8217;s instructive to look into exactly how this can happen.</p>
<p>Beyond the details and the direct interest of this post to political organizations and pollsters, I appreciate Elliott&#8217;s work here because he goes beyond statistical generalities (&#8220;regression to the mean,&#8221; &#8220;sometimes you get a draw from the tail of the distribution,&#8221; etc.).  This is an important statistical point:  the &#8220;error term&#8221; is only an error term until you drill down, look at more data, and figure out what is going on.  It&#8217;s so common for researchers to just take their numbers and not think about where they came from (<a href="https://statmodeling.stat.columbia.edu/2018/08/02/38160/">as here</a>)&#8212;and, indeed, academics and pundits alike can be rewarded for that sort of asinine don&#8217;t-look-carefully-at-the-data attitude.  So it&#8217;s good to see Elliott demonstrating how it&#8217;s possible to do better&#8212;if you&#8217;re willing to put in the work.</p>
<p><strong>Nate Silver and Eli Mckown-Dawson</strong>, &#8220;<a href="https://www.natesilver.net/p/nate-silver-2026-midterm-election-polls-model">FLIPR 2026 midterm election forecast</a>,&#8221; which leads Nate to summarize that &#8220;[Michigan Senate candidate] El-Sayed would be an underdog in an election held today and is an underdog in our &#8220;Lite&#8221; (polls-only) version. The fancy versions look at the fundamentals and are more convinced he&#8217;ll come back.&#8221;</p>
<p>What I really like about this is how <a href="https://statmodeling.stat.columbia.edu/2026/05/29/15-new-articles-on-statistical-workflow/">&#8220;workflow&#8221;</a> it feels.  What I&#8217;m talking about here is how they fit two different models that are doing two different things, they learn something from the comparison, and then they track this back to their data.  This sort of thing isn&#8217;t in the textbooks (well, it wasn&#8217;t <a href="https://sites.stat.columbia.edu/gelman/workflow-book/">until now</a>) but it&#8217;s so important to good applied statistical work.  So I love to see it here.</p>
<p>The point about these two posts, one by Morris and one by Silver and Mckown-Dawson, is not that they&#8217;re so amazing.  I mean, yeah, they&#8217;re great, but the real point is how professional they are.</p>
<p>I remember Bill James once wrote, in reaction to the unexpected playoff heroics of Bucky Dent or Ray Knight or somebody like that, that, sure, it&#8217;s cool when someone steps up and does the unexpected, but what&#8217;s more impressive are those Eddie Murray types who can consistently deliver the expected.  Because then you can design a game plan around them and not just have to hope for a miracle.</p>
<p>Morris, Silver, Mckown-Dawson, and others doing what&#8217;s now expected, doing it well, and demonstrating modern principles of statistical workflow . . . That&#8217;s impressive.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/18/modern-day-political-analysis-excellent-demonstrations-of-statistical-workflow-from-elliott-morris-nate-silver-and-eli-mckown-dawson/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>Jamaican me crazy yet again</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/17/jamaican-me-crazy-yet-again/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/17/jamaican-me-crazy-yet-again/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 01:08:24 +0000</pubDate>
				<category><![CDATA[Public Health]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54509</guid>

					<description><![CDATA[Regular blog readers will remember Golden Krust, source of our standard unit of currency. (Sorry, Peter Thiel, we don&#8217;t take bitcoin.) The above picture is kinda blurry but it gives you a sense of the neighborhood. But here&#8217;s the big &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/17/jamaican-me-crazy-yet-again/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-17-at-21.01.19.png" alt="" width="350" /></p>
<p>Regular blog readers will remember <a href="https://statmodeling.stat.columbia.edu/2011/01/12/picking_pennies/">Golden Krust</a>, source of our standard <a href="https://statmodeling.stat.columbia.edu/2022/02/16/hey-i-got-an-exclusive-invitation-to-this-off-the-record-conference-but-i-think-ill-take-1907-jamaican-beef-patties-instead/">unit of currency</a>.  (Sorry, Peter Thiel, we don&#8217;t take <a href="https://statmodeling.stat.columbia.edu/2023/09/15/omid-malekan-on-why-crypto-is-not-a-scam/">bitcoin</a>.)  The above picture is kinda blurry but it gives you a sense of the neighborhood.</p>
<p>But here&#8217;s the big news!  A new Jamaican beef patty place has appeared, right across the street:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-17-at-21.01.56.png" alt="" width="350" /></p>
<p>Zooming in, it looks like it&#8217;s opening in two days:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-17-at-21.02.17.png" alt="" width="350" /></p>
<p>I&#8217;ll be back soon with a full report (but nothing more on <a href="https://statmodeling.stat.columbia.edu/2022/04/21/jamaican-me-crazy-the-return-of-the-overestimated-effect-of-early-childhood-intervention/">this story</a>).</p>
<p><strong>P.S.</strong>  We did a taste test!  <a href="https://statmodeling.stat.columbia.edu/2026/08/23/head-to-head-on-125-st-the-jamaican-beef-patty-battle/">The story is here</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/17/jamaican-me-crazy-yet-again/feed/</wfw:commentRss>
			<slash:comments>5</slash:comments>
		
		
			</item>
		<item>
		<title>Frequentism for Bayesians:  He wants to teach frequentist methods to engineering students with a strong Bayesian background</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/17/he-wants-to-teach-frequentist-methods-to-engineering-students-with-a-strong-bayesian-background/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/17/he-wants-to-teach-frequentist-methods-to-engineering-students-with-a-strong-bayesian-background/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Mon, 17 Aug 2026 13:05:34 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54311</guid>

					<description><![CDATA[Beyond the teaching question, this is an interesting topic on its own: thinking about classical statistical ideas of point estimation, hypothesis testing, and uncertainty quantification, but taking Bayesian methods as a starting point. The idea is that, instead of structuring &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/17/he-wants-to-teach-frequentist-methods-to-engineering-students-with-a-strong-bayesian-background/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Beyond the teaching question, this is an interesting topic on its own: thinking about classical statistical ideas of point estimation, hypothesis testing, and uncertainty quantification, but taking Bayesian methods as a starting point.</p>
<p>The idea is that, instead of structuring this based on the <em>methods</em> of maximum likelihood estimation, null hypothesis significance testing, and confidence intervals, you start with the <em>goals</em> of estimating parameters, making predictions, comparing and evaluating models (i.e., alternative explanations of the world), and summarizing and working with uncertainty.</p>
<p>If you&#8217;re interested in using Bayesian methods, we have two books (Bayesian Data Analysis and Bayesian Workflow) on the topic. But with these methods under our belt, we can now go back to the motivating questions&#8211;the more general statistical goals, which exist without reference to any particular models or any particular Bayesian or frequentist methods&#8211;and consider them from scratch.</p>
<p>This seems important.</p>
<p>My discussion here is motivated by this question sent in by Opher Donchin:</p>
<blockquote><p>Here&#8217;s one for your blog (although I&#8217;m happy to have your take as well).</p>
<p>I&#8217;m transitioning our undergrad biomedical engineering course this year from a standard frequentist syllabus to a Bayesian approach. We are mostly following the first 6 chapter of <a href="https://www.oreilly.com/library/view/bayesian-analysis-with/9781805127161/">Bayesian Analysis with Python</a> by Osvaldo Martin. In order to get the department to agree to this, I had to promise to teach basic frequentist methods.</p>
<p>Thus, I need to teach frequentist methods to students with a strong Bayesian background. There doesn&#8217;t seem to be much available on this. There is of course material comparing the two approaches, but I mean specific material designed to explain the frequentist approach to someone who knows the Bayesian approach.</p>
<p>I&#8217;d love to know if anyone knows of resources or has experience of insight or advice.</p></blockquote>
<p>I replied that I&#8217;ll see what the blog commenters suggest, but in the meantime, I recommend chapter 4 of <a href="https://sites.stat.columbia.edu/gelman/book/">Bayesian Data Analysis</a> as a start.</p>
<p>Donchin responded:</p>
<blockquote><p>Yes. Chapter 4 is very good on the principles involved. Much of it addresses the Bayesian alternatives to frequentist procedures or the Bayesian perspective on them.</p>
<p>I&#8217;m wondering about something more concrete, aimed at a less sophisticated audience. That is, my students will know how to build models, how to interpret the posterior samples, the basics of Bayesian workflow and also model comparison. On the other hand, they will have no knowledge of confidence intervals, maximum likelihood estimates, hypothesis testing, or multiple comparison procedures. Reasonably, my department demands that they be able to read the biomedical literature where such terms are widespread.</p>
<p>I want to give them an understanding of frequentist procedures without getting bogged down in frequentist justifications.</p>
<p>To take an example, I want to explain what an F test calculates when understood within a Bayesian framework. To that end, I can show students a Bayesian model of a normal distribution of group means and a normal likelihood within each group. Then, I can work through what a Bayesian would need to calculate on that model to produce an F statistic and an F test. It has something to do with summaries of posterior distributions of ratios of variances.</p>
<p>What I&#8217;m hoping for is somewhere where such questions are worked through in detail so that it could be used for developing lectures.</p>
<p>Of course, chapter 4 of BDA is 10 years old at this point. I imagine some of the ideas may have developed since then.</p></blockquote>
<p>Ahhh, good point! Chapter 4 of BDA is for statisticians who already know the classical methods and want to understand how these can be understood in light of Bayesian principles and adapted within a Bayesian workflow. But it&#8217;s really a completely different task to explain classical methods to students who haven&#8217;t already learned them. The idea would be to retcon ideas of classical statistics from a Bayesian angle.</p>
<p>This would be worth doing.</p>
<p>In the meantime, I recommend . . . chapter 4 of <a href="https://sites.stat.columbia.edu/gelman/regression/">Regression and Other Stories</a>, where we go through basic principles of point estimation, hypothesis testing, and uncertainty quantification from an applied perspectives. Also, if you flip through that book, you&#8217;ll see other places where we discuss classical procedures from first principles. We don&#8217;t have any F tests or multiple comparisons adjustments because I can&#8217;t bring myself to care about those things, but a lot else is there, so you might be able to put together much of what you need from that book.</p>
<p>In response to that recommendation, Donchin wrote:</p>
<blockquote>
<div class="elementToProof">I agree that Chapter 4 of Regression and Other Stories touches on many of the important points, but it is not sufficient for my needs.</div>
<div class="elementToProof"></div>
<div class="elementToProof">I am trying to teach a Bayesian-first undergraduate statistics course to biomedical engineers that also gives a background in frequentist approaches allowing  them to function effectively in environments that require them to use or understand frequentist stats.</div>
<div class="elementToProof"></div>
<div class="elementToProof">This means that the frequentist-realted topics we cover include:</div>
<div class="elementToProof"></div>
<ul>
<li>Estimation: MLE, standard errors, confidence intervals</li>
<li>Hypothesis testing: p-values, Type I/II errors, power</li>
<li>Proportion tests and t-tests (independent and paired)</li>
<li>Effect sizes; multiple comparisons (briefly)</li>
<li>Linear and multiple regression: least squares, coefficient tests, R2, F-tests</li>
<li>ANOVA: categorical predictors, interactions, sums of squares, effect sizes</li>
<li>Pearson correlation and inference</li>
<li>Repeated-measures and mixed-effects models</li>
<li>
<div role="presentation">Model comparison: AIC and cross-validation</div>
</li>
<li>
<div role="presentation">Replication crisis / open science / pre-registration (not exactly frequentist, but still)</div>
<div role="presentation"></div>
</li>
</ul>
<div class="elementToProof">You can see the syllabus and the lectures in the student-facing version of the course repo at:<a href="https://urldefense.com/v3/__https://github.com/opherdonchin/StatisticsCourse_36714361__;!!BDUfV1Et5lrpZQ!T8Hjt4t-h_pLF97gv5dXGteiDd_1XRpZesy0qLVh6B33zBDqrf96HmRLtY2XNeCTwpeQ108QN1ESZxYqybrpoA$">https://github.com/opherdonchin/StatisticsCourse_36714361</a></div>
<div class="elementToProof"></div>
<div class="elementToProof">Any further thoughts you might have would be great to hear.</div>
<div class="elementToProof"></div>
<div class="elementToProof">Also, I would be happy to get this some visibility, in hopes that other people would be interested in providing feedback, using some of the material, or just doing it better.</div>
</blockquote>
<p>I don&#8217;t think I could bring myself to teach a lot of the above topics, except in an &#8220;inoculation&#8221; sort of way, but I recognize that many students will need it, so if anyone has some good suggestions for Donchin, just leave them here in the comments!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/17/he-wants-to-teach-frequentist-methods-to-engineering-students-with-a-strong-bayesian-background/feed/</wfw:commentRss>
			<slash:comments>48</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;How Can America Be So Miserable When It’s So Rich?&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/16/how-can-america-be-so-miserable-when-its-so-rich/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/16/how-can-america-be-so-miserable-when-its-so-rich/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 16 Aug 2026 13:03:10 +0000</pubDate>
				<category><![CDATA[Economics]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53469</guid>

					<description><![CDATA[The above is the title of a recent a newspaper op-ed (pdf here) by columnist David French. He&#8217;s asking a good question! My quick take is that there are four things going on: 1. People want improvement. If you&#8217;re at &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/16/how-can-america-be-so-miserable-when-its-so-rich/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>The above is the title of a recent <a href="https://www.nytimes.com/2026/03/26/opinion/economy-attitudes-republicans-democrats.html">a newspaper op-ed</a> (pdf <a href="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Opinion-How-Can-America-Be-So-Miserable-When-Its-So-Rich-The-New-York-Times.pdf">here</a>) by columnist David French.  He&#8217;s asking a good question!  My quick take is that there are four things going on:</p>
<p>1.  People want improvement.  If you&#8217;re at X, then it&#8217;s great to go to 1.2X, it&#8217;s kinda scary to stay at X, and it&#8217;s distressing to go to 0.9X.  It does seem that people react to the change in their economic situation rather than than the absolute level.</p>
<p>2.  People want more.  Housing is more expensive than it used to be, but part of that is that people want to live in bigger apartments and bigger houses.  Cars are much more reliable than ever, but nowadays lots of people want two or three cars in their household.</p>
<p>3.  Entitlement.  Some people are struggling and it makes sense for them to be distressed with the economy.  On the other hand, lots of people are sitting there with expensive houses all paid for and some money in the bank, but they&#8217;re not saying how amazingly wonderful things are, because they feel that they deserve every bit of it.  This averages out to unhappy in the population.</p>
<p>4.  Lack of security.  The concern that, things might be ok now, but you might have to be scrambling in the future.</p>
<p>This is not to deny that lots of people are struggling economically.  I&#8217;m just talking here about average poll findings.  The question is not why <em>some</em> or even <em>many</em> people are economically miserable, so much as why the average response to such questions is lower than in earlier decades when people had a lot less.  The point is that people are thinking about their future, they&#8217;re not comparing to how things were in past decades.</p>
<p>French gives some examples of how the competitive economy has distorted the upper middle class.  He even talks about <a href="https://statmodeling.stat.columbia.edu/2025/03/12/those-youth-sports-travel-teams/">youth sports travel teams</a>.  It&#8217;s hard for me to believe that the high cost of Disney tickets and travel team fees are driving the affordability crisis&#8211;neither of these is anything like the struggle to pay the rent and take care of kids or elderly relatives while commuting to some faraway job with rigid working conditions.  French also talks about the frustration of flying, but most people get on planes only rarely.</p>
<p>One other thing.  French writes:</p>
<blockquote><p>In this story, maybe the problem isn’t oligarchy. Elon Musk’s billions don’t tangibly change my life.</p></blockquote>
<p>But . . . Elon Musk&#8217;s billions <em>do</em> tangibly change your life!  Remember, they went in and screwed up the government.  Planes are crashing, people are getting shot on the street, your data are <a href="https://www.nytimes.com/2025/08/26/us/politics/doge-social-security-data.html">at risk</a>, they&#8217;re reducing childhood vaccinations and making it more difficult to develop future adult vaccines . . . so, yeah, you can blame those billions.  You can also blame the zillions of dollars spent by lobbyists to legalize gambling on people&#8217;s phones.</p>
<p>I agree that &#8220;oligarchy&#8221; is not the only problem&#8211;and, arguably, oligarchy does lots of good things too&#8212;but it&#8217;s naive to think that it&#8217;s not changing your life.  Even if you are not personally a gambling addict, you might have a loved one who is.  Even if the kids in your family are currently vaccinated, you or they could get killed by the next pandemic.  Etc.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/16/how-can-america-be-so-miserable-when-its-so-rich/feed/</wfw:commentRss>
			<slash:comments>108</slash:comments>
		
		
			</item>
		<item>
		<title>Inbox zero.  Bloglag 365.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/15/inbox-zero-bloglag-365/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/15/inbox-zero-bloglag-365/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sat, 15 Aug 2026 13:00:08 +0000</pubDate>
				<category><![CDATA[Decision Analysis]]></category>
		<category><![CDATA[Literature]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54495</guid>

					<description><![CDATA[With some sustained effort, I cleaned out my entire inbox! Now I&#8217;m ready to do research again. Actually, cleaning out the inbox entailed doing a lot of research and writing, lots of interesting things I&#8217;d put off for a while. &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/15/inbox-zero-bloglag-365/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-1024x444.png" alt="" width="584" height="253" class="alignnone size-large wp-image-54503" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-1024x444.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-300x130.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-768x333.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-1536x666.png 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21-500x217.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-14-at-18.57.21.png 1860w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>With some sustained effort, I cleaned out my entire inbox!  Now I&#8217;m ready to do research again.</p>
<p>Actually, cleaning out the inbox entailed doing a lot of research and writing, lots of interesting things I&#8217;d put off for a while.</p>
<p>One byproduct of all this work is that our blog queue is now a year long.  That&#8217;s right&#8212;the new posts are being scheduled for August, 2027.  Although I do reserve the right to post things earlier if it seems that they&#8217;re particularly relevant, my co-bloggers will continue to post whenever they want, and we sometimes post things early on <a href="https://statmodeling.substack.com">the newsletter</a>.  So you never know what&#8217;ll be coming.</p>
<p>It&#8217;s a funny thing, though. I&#8217;ve been writing so much for 2027, that now when I read something with a 2026 date, my first thought is, Hey, that happened last year!  I&#8217;m feeling like <a href="https://statmodeling.stat.columbia.edu/2015/01/20/another-benefit-bloglag/">Jones</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/15/inbox-zero-bloglag-365/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;Q collar&#8221; promoters have no shame, advertising their discredited research on social media using lying, cheating chatbots</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/14/q-collar-promoters-have-no-shame-advertising-their-discredited-research-on-social-media-using-lying-cheating-chatbots/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/14/q-collar-promoters-have-no-shame-advertising-their-discredited-research-on-social-media-using-lying-cheating-chatbots/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 14 Aug 2026 13:11:39 +0000</pubDate>
				<category><![CDATA[Public Health]]></category>
		<category><![CDATA[Sports]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54446</guid>

					<description><![CDATA[In a distressing update to an already distressing story, James Smoliga writes: Q-collar is ramping up its fall sports campaign. Lots of buzz on Facebook, with misleading comments from the company, new ways of spinning the FDA disclaimers, and plenty &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/14/q-collar-promoters-have-no-shame-advertising-their-discredited-research-on-social-media-using-lying-cheating-chatbots/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>In a distressing update to <a href="https://statmodeling.stat.columbia.edu/2026/08/13/we-should-probably-be-doing-more-coverage-of-science-adjacent-medical-scams-like-the-q-collar-and-less-on-repulsively-self-promoting-but-ultimately-less-harmful-academic-grifters/">an already distressing story</a>, James Smoliga writes:</p>
<blockquote><p>Q-collar is ramping up its fall sports campaign. Lots of buzz on Facebook, with misleading comments from the company, new ways of spinning the FDA disclaimers, and plenty of excited parents.</p>
<p>People are also now posting AI generated summaries which mention the research &#8211; but skip the Expressions of Concern, etc.</p>
<p>Here&#8217;s some examples, happy to provide more!</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot_20260806_233326_Facebook-473x1024.jpg" alt="" width="300" /></p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-473x1024.jpeg" alt="" width="300" /></p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot_20260806_232906_Facebook-473x1024.jpg" alt="" width="300" />
</p></blockquote>
<p>Wow, those chatbot-generated evasive responses (&#8220;You are absolutely right, and we never want to gloss over this&#8221;) . . . what a bunch of assholes!  This is the sort of dishonesty we could expect from an outlet like <a href="https://statmodeling.stat.columbia.edu/2010/05/12/alert_incompete/">Wolfram Research</a>.</p>
<p>Matthew Walker, Brian Wansink, Cass Sunstein, and various other professional-avoiders-of-inconvenient-facts could lean a lot from these guys on <a href="https://statmodeling.stat.columbia.edu/2019/01/18/ladder-responses-criticism-responsible-destructive/">how to pretend to respond to criticism</a> without directly addressing it.</p>
<p>As is so often the case, the threat from AI is not the AI doing malevolent things on its own; it&#8217;s bad actors using AI capabilities to manipulate humans.</p>
<p>The next step, I guess, would be to get some <a href="https://statmodeling.stat.columbia.edu/2025/06/20/nih-plan-to-remove-ideological-influence-from-science-how-does-this-fit-in-the-junk-science-being-promoted-by-the-u-s-dept-of-health-and-human-services/">official promotion</a> from the U.S. Department of Health and Human Services, by way of some strategic campaign contributions.  Maybe they could arrange some contract to put the Q collar on the updated Mount Rushmore, something like that?  Just spitballin here.</p>
<p><strong>P.S.</strong>  Just to remember:  Yes, real brains are at stake here.  <a href="https://www.chronicle.com/article/this-device-is-proven-to-protect-athletes-brains-the-science-is-under-fire">This is evil</a>.</p>
<p><strong>P.P.S.</strong>  I can see why there&#8217;s so little reaction to this story, as this sort of little violation of human dignity is nothing compared to the official legalization of corruption, as <a href="https://www.nytimes.com/2026/08/12/us/politics/treasury-scrutiny-shell-companies.html">in this story</a> that contains this Orwellian line from the Treasury secretary: &#8220;Treasury is eliminating a burdensome reporting requirement for millions of law-abiding business owners without compromising our national security.&#8221;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/14/q-collar-promoters-have-no-shame-advertising-their-discredited-research-on-social-media-using-lying-cheating-chatbots/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Answering some questions about statistics from a high school student</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/13/answering-some-questions-about-statistics-from-a-high-school-student/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/13/answering-some-questions-about-statistics-from-a-high-school-student/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 13 Aug 2026 22:45:06 +0000</pubDate>
				<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54494</guid>

					<description><![CDATA[This came in the mail: My name is ** and I am a junior at ** High School. I am writing to you as a part of a summer assignment for my upcoming research class where I had to choose &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/13/answering-some-questions-about-statistics-from-a-high-school-student/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>This came in the mail:</p>
<blockquote><p>My name is ** and I am a junior at ** High School. I am writing to you as a part of a summer assignment for my upcoming research class where I had to choose an expert in the math field to contact. I chose you considering your work and investment in the statistics field and how you have used statistics in the political landscape. I found it very interesting and I was wondering a few things,<br />
• What got you into math and statistics?<br />
• What advice would you give to students interested in statistics?<br />
• What is the biggest challenge when working with statistics?<br />
• What is a common mistake people make when looking at statistics?<br />
• Where would be good resources to learn more?<br />
I appreciate your time and I hope to hear from you soon.</p></blockquote>
<p>Here were my responses:</p>
<blockquote><p>
• What got you into math and statistics?<br />
See here:  https://statmodeling.stat.columbia.edu/2010/11/02/fragment_of_sta/</p>
<p>• What advice would you give to students interested in statistics?<br />
In addition to statistics, you should have some applied field you are interested in.  This could be biology, economics, sociology, medicine, business, communications, history, . . . just about anything.  Then when you learn statistical ideas, consider how they apply to this field, and work on a project applying statistics in that way.</p>
<p>• What is the biggest challenge when working with statistics?<br />
Measurement; see here:  https://statmodeling.stat.columbia.edu/2015/04/28/whats-important-thing-statistics-thats-not-textbooks/</p>
<p>• What is a common mistake people make when looking at statistics?<br />
Looking at the numbers and not checking where they came from.  For example, see here:  https://statmodeling.stat.columbia.edu/2017/01/02/bogus-north-korea/</p>
<p>• Where would be good resources to learn more?<br />
I recommend our blog:  https://statmodeling.stat.columbia.edu/ and our book, Active Statistics, which is full of stories:  https://sites.stat.columbia.edu/gelman/active-statistics/<br />
Also I like Frederick Mosteller&#8217;s classic book from 1965, Fifty Challenging Problems in Probability.  I was a teaching assistant for Mosteller for the last course he ever taught!</p></blockquote>
<p>Since I was answering the questions anyway, I thought I&#8217;d post it all here as others might be interested too.  I think that all of this is reasonable advice.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/13/answering-some-questions-about-statistics-from-a-high-school-student/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>7 1/2 Schools&#8212;mistakes were made</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/13/7-1-2-schools-mistakes-were-made/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/13/7-1-2-schools-mistakes-were-made/#comments</comments>
		
		<dc:creator><![CDATA[Bob Carpenter]]></dc:creator>
		<pubDate>Thu, 13 Aug 2026 19:00:06 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Stan]]></category>
		<category><![CDATA[Statistical Computing]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54470</guid>

					<description><![CDATA[Mitzi Morris sends along the paper trail (I love archaic idioms) for her data spelunking in advance of her all-day introductory Stan tutorial next week at StanCon (we&#8217;ll see you in Uppsala). She was trying to reproduce the results from &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/13/7-1-2-schools-mistakes-were-made/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://github.com/mitzimorris">Mitzi Morris</a> sends along the <a href="https://idioms.thefreedictionary.com/paper+trail">paper trail</a> (I love archaic idioms) for her data spelunking in advance of her all-day introductory Stan tutorial next week at StanCon (we&#8217;ll see you in Uppsala).  She was trying to reproduce the results from my decade-old <a href="https://mc-stan.org/learn-stan/case-studies/pool-binary-trials.html">case study on hierarchical modeling</a> when she ran into a couple discrepancies.</p>
<p>The case study was designed to provide a Bayesian replication of the results in Efron and Morris&#8217;s five decade old paper on hierarchical modeling (aka Stein&#8217;s estimator, aka population regularization, aka &#8220;empirical&#8221; Bayes), which is still under a paywall courtesy of our &#8220;friends&#8221; at the American Statistical Association.  If your organization isn&#8217;t paying ASA for access to a paper that an academic donated for free 50 years ago, I&#8217;ll leave you to find your own pirated copy in good conscience, or you can follow the link and let Google <a href="https://hinative.com/questions/21450050">hoist the Jolly Roger</a> for you (now that&#8217;s an even more obscure and archaic reference).</p>
<ul>
<li>Bradley Efron and Carl Morris. 1975.  <a href="https://scholar.google.com/scholar?cluster=6208261065379644186&#038;hl=en&#038;as_sdt=0,33">Data Analysis Using Stein’s Estimator and Its Generalizations</a>. JASA 70(350).
</ul>
<p>My case study&#8217;s been out for ten years.  Mitzi found a problem when matching the data provided in the R package <code>pscl</code> against that in Efron and Morris&#8217;s paper.  In particular, the data for the player named &#8220;Williams&#8221; was wrong.  Efron and Morris manually transcribed the data from a newspaper with the goal of finding a bunch of players with the exact same number of at bats on a given day (45, it turns out).  They did so accurately.</p>
<p>Sadly, I imported an error in the 18-item data set from the R package. Here&#8217;s an edited report from Claude after it tracked down what went wrong.</p>
<ol>
<li> The 1975 paper is correct, although it only gives the surname “Williams.”
<p></p>
<li> The R package <code>pscl</code> expanded that name incorrectly to Billy Williams, while retaining Walt’s team and remaining-season data. Published <a href="https://okumuralab.org/~okumura/stat/efronmorris.html"><code>pscl</code> data</a>.
<p></p>
<li>The <a href="https://beanmachine.org/docs/overview/tutorials/Hierarchical_modeling/HierarchicalModeling/">Bean Machine tutorial</a> () contains the exact mixed row found in your CSV:
<pre>
  Billy,Williams,...,591,0.3299,636,205,0.251
</pre>
<p></p>
<li>That row combines Billy Williams’s at-bats and hits with Walt Williams’s .251 average. Thus, ID 15 in your data should be corrected throughout to:
<pre>
  Walt,Williams,45,10,0.222,270,0.2556,315,79,0.251
</pre>
</ol>
<p>But wait, that&#8217;s not all.  There was another typo in the Bean Machine tutorial.  They mistranscibed <a href="https://en.wikipedia.org/wiki/Ron_Swoboda">Ron Swoboda</a>&#8216;s surname as &#8220;Swaboda.&#8221;  I should have caught that as I was a huge baseball fan in the early 1970s, am an excellent speller, and grew up around a lot of Poles and Czechs in Detroit and its western suburbs.  Mitzi then had the following conversation with Claude after she noticed the typo.  This may be the first time I&#8217;ve seen Claude crack a spontaneous joke.</p>
<blockquote><p>
<I>Mitzi</I>: did that [the typo] come also from the R file?</p>
<p><I>Claude</I>: If you mean the Swaboda misspelling: no. The pscl R dataset correctly uses Ron Swoboda. The misspelling appears in the Bean Machine tutorial<br />
(https://beanmachine.org/docs/overview/tutorials/Hierarchical_modeling/HierarchicalModeling/), which contains Ron,Swaboda.</p>
<p><I>Mitzi</I>: are you sure?  I&#8217;ve heard it both ways.</p>
<p><I>Claude</I>: Yes—despite what Shawn Spencer might say, it is definitively Ron Swoboda, with an o.</p>
<p><I>Mitzi</I>: who&#8217;s Shawn Spencer?  in 2017, Bob Carpenter published this dataset in a Stan case study.</p>
<p><I>Claude</I>: Shawn Spencer is the fictional detective in the TV comedy <i><a href="https://en.wikipedia.org/wiki/Psych">Psych</a></i>, played by James Roday Rodriguez. “I’ve heard it both ways” is his recurring response when someone corrects him—hence my reference.
</p></blockquote>
<p>Mitzi and I use &#8220;I&#8217;ve heard it both ways&#8221; all the time, having watched umpteen seasons of <I>Psych</I>.  I would recommend if you want a lighthearted detective story that doesn&#8217;t take itself too seriously with lovable characters and you can put up with the ubiquitous mid-00s background sexism and racism.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/13/7-1-2-schools-mistakes-were-made/feed/</wfw:commentRss>
			<slash:comments>21</slash:comments>
		
		
			</item>
		<item>
		<title>Lots of detail on the continuing &#8220;Q collar&#8221; scandal and why we should probably be doing more coverage of that sort of science-adjacent medical scams and less on repulsively self-promoting but ultimately less harmful academic grifters.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/13/we-should-probably-be-doing-more-coverage-of-science-adjacent-medical-scams-like-the-q-collar-and-less-on-repulsively-self-promoting-but-ultimately-less-harmful-academic-grifters/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/13/we-should-probably-be-doing-more-coverage-of-science-adjacent-medical-scams-like-the-q-collar-and-less-on-repulsively-self-promoting-but-ultimately-less-harmful-academic-grifters/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 13 Aug 2026 13:53:48 +0000</pubDate>
				<category><![CDATA[Economics]]></category>
		<category><![CDATA[Miscellaneous Science]]></category>
		<category><![CDATA[Public Health]]></category>
		<category><![CDATA[Sports]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54378</guid>

					<description><![CDATA[The Q collar, I hope you don&#8217;t remember, is a neck accessory that is advertised as &#8220;proven&#8221; to protect athletes&#8217; brains, but the studies purporting to offer this proof are suspect. This is an interesting example of science-adjacent incompetence or &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/13/we-should-probably-be-doing-more-coverage-of-science-adjacent-medical-scams-like-the-q-collar-and-less-on-repulsively-self-promoting-but-ultimately-less-harmful-academic-grifters/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>The Q collar, I hope you don&#8217;t remember, is a neck accessory that is advertised as &#8220;proven&#8221; to protect athletes&#8217; brains, but the studies purporting to offer this proof <a href="https://statmodeling.stat.columbia.edu/2025/01/29/this-device-is-proven-to-protect-athletes-brains-the-science-is-under-fire/#comment-2391308">are suspect</a>.</p>
<p>This is an interesting example of science-adjacent incompetence or fraud or grifting or whatever you want to call it, as it overlays junk science with big business (according to journalist Stephanie Lee, the device costs at least $200 and &#8220;investment in the Q-collar has totaled more than $30 million&#8221;), the government (including the U.S. Army, a U.S. senator, and the Food and Drug Administration), news and social media, and some aspects of the scientific establishment (medical journals).</p>
<p>You can think of Q collar as a baby version of health scams promoted by wealthy and powerful medical-adjacent celebrities such as Dr. Oz, Andrew Huberman, and RFK Jr. It&#8217;s junk science, but junk science with a purpose, in this case not so much to promote an ambitious social climber (as with <a href="https://statmodeling.substack.com/p/why-do-david-agus-and-jason-arday">Agus and Arday</a>), it seems to be more about the money. And in that way it could be seen as a more typical (if less dramatic) example of a science scam. It doesn&#8217;t involve any plagiarized giraffes or implausible claims of ultra-marathoning, just the exploitation of people&#8217;s legitimate health concerns and the lure of the filthy lucre.  It could even be that the scammers believe their own claims (remember, the first step is to fool yourself; the rest then comes easy) and believe they&#8217;re doing well be doing good.</p>
<p>We should probably be doing more coverage of science-adjacent medical scams like the &#8220;Q collar&#8221; and less on the repulsively self-promoting but ultimately less harmful Aguses, Ardays, Wansinks, Hausers, and Arielys of the world who each in his way degrades the reputation of science but isn&#8217;t so closely locked in to a moneymaking scheme. (OK. I guess Wansink did ok as a consultant, and Agus and Ariely founded a few businesses which probably netted them a good chunk of change.  They might well be first-class wealthy, even if not in the private-jet stratum.)</p>
<p>Here&#8217;s a Q-collar update from science sleuth Mu Yang:</p>
<blockquote><p>Two new articles on Q collar came out today:</p>
<ul>
<li><a href="https://urldefense.com/v3/__https:/www.bmj.com/content/391/bmj.r2028.short?rss=1__;!!BDUfV1Et5lrpZQ!QCE1QRtBn953OPjYZyZqzc-xMg6kN0NfWrTQemkqbASJOm8CY7ohDXbxY29VxOMlBo6VYwbAJW-qLZy4CIzC1Epib40$">BMJ</a> (with a nice video clip) and <a href="https://urldefense.com/v3/__https:/www.bmj.com/content/391/bmj.r2028.short?rss=1__;!!BDUfV1Et5lrpZQ!QCE1QRtBn953OPjYZyZqzc-xMg6kN0NfWrTQemkqbASJOm8CY7ohDXbxY29VxOMlBo6VYwbAJW-qLZy4CIzC1Epib40$">WaPo</a>.</li>
<li>Here are our own takes on <a href="https://urldefense.com/v3/__https:/www.linkedin.com/posts/mu-yang-11229436_the-science-behind-this-brain-protection-activity-7384567659872350209-qCUC?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAeELJ4B1QYBbcltHKam2Rb6-9WyCPbaZGc__;!!BDUfV1Et5lrpZQ!QCE1QRtBn953OPjYZyZqzc-xMg6kN0NfWrTQemkqbASJOm8CY7ohDXbxY29VxOMlBo6VYwbAJW-qLZy4CIzCSM3qwVE$">LinkedIn</a> (mine) and <a href="https://urldefense.com/v3/__https:/beyondtheabstract.substack.com/p/error-correction-collapse-when-does__;!!BDUfV1Et5lrpZQ!QCE1QRtBn953OPjYZyZqzc-xMg6kN0NfWrTQemkqbASJOm8CY7ohDXbxY29VxOMlBo6VYwbAJW-qLZy4CIzC7V4Wy0M$">Substack</a> (Jamey&#8217;s).</li>
</ul>
<p style="font-weight: 400;">The duplicated tables are almost cute compared to the real issue yet to be addressed. &#8212; I mumbled about it in <a href="https://statmodeling.stat.columbia.edu/2025/01/29/this-device-is-proven-to-protect-athletes-brains-the-science-is-under-fire/">a comment to your post back inJanuary</a>.</p>
<p style="font-weight: 400;">Basically, the design is what one would use to teach “confounding factor” in Stats 101:</p>
<ul style="font-weight: 400;">
<li>5 studies (3 on Q collar and 2 on helmet type) reported data from subjects from the same cohort.</li>
<li>The subjects in the helmet studies are the controls from the Q collar studies.</li>
<li>Two types of helmets were worn by Q-collar-controls, and helmet type had significant impact on brain white matter parameters</li>
<li>What helmets Q collar subjects wore is unknown.</li>
<li>Q collar is said to affect the same brain white matter parameters as helmet type, in the same direction (the good one)</li>
</ul>
<p style="font-weight: 400;">I [Yang] wrote more about it <a href="https://urldefense.com/v3/__https://www.linkedin.com/posts/mu-yang-11229436_two-weeks-ago-our-bmj-article-how-an-fda-activity-7390444014061162497-EpiM?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAeELJ4B1QYBbcltHKam2Rb6-9WyCPbaZGc__;!!BDUfV1Et5lrpZQ!WuQnhKPnOp521I9ihS0N8QOt4IH8qfnVXU5BW7jokRngymPQk7DR0uJylras1UjFE39g0yD1ZeUb8DT9aJuK0tQiJXE$">here</a> and posted on <a href="https://urldefense.com/v3/__https://pubpeer.com/publications/A56C0ECFC2EF0644EE81821FEC1FC2__;!!BDUfV1Et5lrpZQ!WuQnhKPnOp521I9ihS0N8QOt4IH8qfnVXU5BW7jokRngymPQk7DR0uJylras1UjFE39g0yD1ZeUb8DT9aJuKUECc2nQ$">Pubpeer</a>.</p>
<p style="font-weight: 400;">Things are at a standstill. The authors have been contacted many times (and I made sure that they are notified about my Pubpeer posts), but remain reticent. It seems that FDA does not care.</p>
</blockquote>
<p>Awhile later, Yang added:</p>
<blockquote><p>Due to the incurable stubborness of some of us, we have been digging deeper in the Q collar junkyard long after <a href="https://www.chronicle.com/article/this-device-is-proven-to-protect-athletes-brains-the-science-is-under-fire">Stephanie&#8217;s article</a>, <a href="https://statmodeling.stat.columbia.edu/2025/01/29/this-device-is-proven-to-protect-athletes-brains-the-science-is-under-fire/">your blog post</a>, and the <a href="https://www.bmj.com/content/391/bmj.r2028.abstract">BMJ article</a>.</p>
<p>To put it extremely simply: The control subjects in the Q collar studies were described in two studies on helmet types &#8212;which were shown to have substantial impacts on the same brain imaging metrics as Q-collar does. Given controls and Q-collar subjects were students in the same schools, Q collar subjects likely also wore different helmets. BUT, this key confounding factor was never mentioned, quantified, controlled for, or discussed in Q collar papers. I commented on this <a href="https://pubpeer.com/publications/CD9DCC655CB74BCCCAC75D9E9A6506#1">here</a>. Following that, Dr. M. Shane Tutwiler, a quantitative researcher in the field of educational psychology in the University of Rhode Island, became interested in the case, and wrote <a href="https://pubpeer.com/publications/CD9DCC655CB74BCCCAC75D9E9A6506#2">a very comprehensive comment</a> on Pubpeer. Shane&#8217;s detailed analysis argues that the design and analysis of one of the key studies are grossly inadequate to support the claims of the study&#8212;which was crucial for the FDA clearance. </p>
<p>Before posting on PubPeer, Shane spoke to the authors directly to alert them of his concerns and request the data&#8212;the 2021 article has a data sharing agreement&#8212;but they denied his request because he was critical of their work. </p>
<p>This &#8220;silly&#8221; and &#8220;messy&#8221; case is looking much worse than previously thought. However, Dr. Smoliga and I are being stonewalled by SAGE &#8212; the six papers under EoC never moved forward. Dr. Tutwiler&#8217;s reports to Wiley are also not getting meaningful replies.</p>
<p>We are wondering if you are willing to take a look at the <a href="https://pubpeer.com/publications/CD9DCC655CB74BCCCAC75D9E9A6506">new thread of evidence</a> and weigh in.</p></blockquote>
<p>My quick answer is that the story seems clear enough already, even without this new thread!  I also noticed that one of the authors of one of the papers under discussion is certain &#8220;Jeffery N Epstein.&#8221;  <a href="https://statmodeling.stat.columbia.edu/2026/02/25/axel-f-meets-samuel-beckett-in-the-worlds-most-pointless-conversation/">Not the same guy</a>, and I&#8217;m guessing that my Columbia colleague Richard &#8220;Axel&#8221; Foley wouldn&#8217;t spend any time responding to his emails.</p>
<p>But I digress.</p>
<p>To return to this horrible &#8220;Q collar&#8221; story, Shane Tutwiler writes:</p>
<blockquote><p>I just wanted to close the loop, as it were, on the &#8220;Investigation&#8221; of the paper I evaluated on <a href="https://pubpeer.com/publications/CD9DCC655CB74BCCCAC75D9E9A6506#2">PubPeer</a>. The authors shared the data with the editors, and the editors were seemingly content that the paper was perfectly fine as-is (despite many of the issues I pointed out being raised in peer review and ignored by the authors and managing editor at the time). Wiley has closed the case (see below). </p>
<p>If this were a run of the mill journal article that nobody was going to read, I wouldn&#8217;t feel so upset. But this research was used by the FDA as part of the evidence chain to justify the device&#8217;s use in public. It&#8217;s one of those cases where questionable analyses can have real world consequences. </p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5-1024x390.png" alt="" width="584" height="222" class="alignnone size-large wp-image-54452" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5-1024x390.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5-300x114.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5-768x292.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5-500x190.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/image-5.png 1177w" sizes="(max-width: 584px) 100vw, 584px" /></p></blockquote>
<p>Without any additional knowledge of this particular case, just based on my general understanding of such things, my guess is that the journals&#8217; decision to do nothing here is motivated less by corruption and more by simple laziness.  If you retract an article, the authors can scream at you.  If you do nothing, people like Tutwiler and me will scream at you . . . but people like us are much less dangerous than authors who refuse to let go of bad research:  Those people have money and reputations on the line and can lash out or even sue!  So safer to just do nothing and pretend that all is well. The turtle strategy was <a href="https://statmodeling.substack.com/p/whos-the-most-pitiful-person-in-the">not so effective recently</a> with Cambridge University, but usually it works just fine, so I guess I can understand why the editors are doing this, even though I think it&#8217;s unethical behavior on their part.</p>
<p>And James Smoliga adds some additional context:</p>
<blockquote><p>The investigation we raised with the Journal of Neurotrauma has been ongoing for over two years now.  In the 10 months since the Expressions of Concern were issued, no further action has taken place.</p>
<p>Those six papers are separate from the one that Shane flagged to Wiley… And importantly, that study was the key one that resulted in the FDA’s authorization of the product, which further generated millions in venture capital investment, DoD funding, and consumer purchasing.</p>
<p>This is a mess at every level – improper statistics, inappropriate interpretation, sloppy errors, and allegations of p-hacking (an admission from a collaborator on their research, who doesn’t seem excited to publicly state this).</p>
<p>If you have any interest in writing more about it, or know of any journalists who may be interested in covering this story further, any help is appreciated!</p></blockquote>
<p>I know lots of journalists, but they usually don&#8217;t like covering this sort of story.  The outlet most likely to cover it might be Defector, but I don&#8217;t know anyone there.  Actually, it would fit wonderfully into their Only If You Get Caught podcast.  Maybe I can do the contacts-of-contacts thing and see if I can reach anyone there to make the suggestion.</p>
<p>The New York Times already had <a href="https://www.nytimes.com/2022/12/19/sports/concussion-q-collar.html">this article in 2022</a> that was critical of the Q collar, so I don&#8217;t see them running more stories on it.  Once you get to the point of junk-science-is-refuted-but-still-keeps-on-making-money, you&#8217;re entering &#8220;dog bites man&#8221; territory, no?</p>
<p>Maybe you could convince noted science skeptics Sean Carroll and Steven Levitt to cover this story . . . ha ha, <a href="https://statmodeling.stat.columbia.edu/2025/01/27/does-anyone-actually-expect-meaningful-insight-to-come-from-a-study-like-this/">just kidding</a>!  More seriously, maybe <a href="https://statmodeling.stat.columbia.edu/2026/05/17/if-books-could-kill-podcast/">If Books Could Kill</a> would be interested, but I don&#8217;t know them either.</p>
<p><strong>P.S.</strong>  <a href="https://statmodeling.stat.columbia.edu/2026/08/14/q-collar-promoters-have-no-shame-advertising-their-discredited-research-on-social-media-using-lying-cheating-chatbots/">More here</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/13/we-should-probably-be-doing-more-coverage-of-science-adjacent-medical-scams-like-the-q-collar-and-less-on-repulsively-self-promoting-but-ultimately-less-harmful-academic-grifters/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>Bayesian computation meeting in France in May, 2027</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/12/bayesian-computation-meeting-in-france-in-may-2027/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/12/bayesian-computation-meeting-in-france-in-may-2027/#respond</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 12 Aug 2026 21:46:08 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Statistical Computing]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54483</guid>

					<description><![CDATA[OK, this one&#8217;s time constrained so I&#8217;ll post it right away, not on the usual lag. Christian Robert points to this conference announcement: We invite proposals for the BayesComp 2027 mirror event in Aussois, France, that will take place on &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/12/bayesian-computation-meeting-in-france-in-may-2027/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28-1024x249.png" alt="" width="584" height="142" class="alignnone size-large wp-image-54484" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28-1024x249.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28-300x73.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28-768x187.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28-500x122.png 500w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Screenshot-2026-08-12-at-17.43.28.png 1266w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>OK, this one&#8217;s time constrained so I&#8217;ll post it right away, not on the usual lag.</p>
<p>Christian Robert points to <a href="https://docs.google.com/forms/d/e/1FAIpQLSfK1wAmKMCRnLz57-gCzbFMuQWwsfb5HvhWYtLRkLLzvKedMg/viewform">this conference announcement</a>:</p>
<blockquote><p>We invite proposals for the BayesComp 2027 mirror event in Aussois, France, that will take place on 17-21 May, 2027. The event will complement the main BayesComp 2027 meeting in Texas and provide a European forum for researchers in Bayesian computation and related areas. Plenary sessions will be broadcasted from Texas, while talks in Aussois will be recorded.</p>
<p>Please submit one form per proposed contribution. Proposals may be for contributed talks, or poster presentations. The programme committee will review submissions based on scientific quality, timeliness, coherence, and expected appeal to the BayesComp community. Unless the submission mentions otherwise, talk proposals that are rejected will be automatically added as posters.</p>
<p>Details about the venue (CAES Paul Langevin), potential registration fees, and technical details are being finalized and will be announced on the event website.</p></blockquote>
<p>Looks like fun!  And Bayesian computation is important.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/12/bayesian-computation-meeting-in-france-in-may-2027/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Who were &#8220;The Makers of Public Policy&#8221; in 1965?</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/12/who-were-the-makers-of-public-policy-in-1965/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/12/who-were-the-makers-of-public-policy-in-1965/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 12 Aug 2026 13:03:04 +0000</pubDate>
				<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52832</guid>

					<description><![CDATA[I was in the Playroom the other day and noticed this book on the shelf, The Makers of Public Policy, by R. J. Monsen and M. W. Cannon, copyright 1965. Here&#8217;s the table of contents: It&#8217;s interesting to see here &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/12/who-were-the-makers-of-public-policy-in-1965/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>I was in the Playroom the other day and noticed this book on the shelf, The Makers of Public Policy, by R. J. Monsen and M. W. Cannon, copyright 1965.  Here&#8217;s the table of contents:</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-1024x1017.png" alt="" width="584" height="580" class="alignnone size-large wp-image-52826" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-1024x1017.png 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-300x298.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-150x150.png 150w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-768x763.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36-302x300.png 302w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.41.36.png 1196w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00-644x1024.png" alt="" width="584" height="929" class="alignnone size-large wp-image-52827" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00-644x1024.png 644w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00-189x300.png 189w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00-768x1222.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00-965x1536.png 965w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-25-at-17.42.00.png 1188w" sizes="(max-width: 584px) 100vw, 584px" /></p>
<p>It&#8217;s interesting to see here what is and isn&#8217;t included.</p>
<p>1.  The section on Formal Power Groups contains one chapter on business, one on agriculture, and <em>two</em> on labor.  Organized labor is so much less now that I don&#8217;t think it would merit even one chapter.</p>
<p>2.  There&#8217;s a chapter on &#8220;negroes.&#8221; I guess makes sense from the standpoint on 1965 that there are no chapters on other social outgroups such as Latinos, women, gays, involuntary celibates, immigrants, disabled people, etc.</p>
<p>3.  There&#8217;s no chapter at all on organized religion, which is a notable missing power group now but even more so for the 1960s.</p>
<p>4.  There&#8217;s a chapter on &#8220;intellectuals&#8221; but nothing on political ideology.</p>
<p>5.  There is no chapter regarding the organized Left, but I guess it&#8217;s covered in the penumbra of &#8220;labor&#8221; and &#8220;intellectuals.&#8221;</p>
<p>6.  There&#8217;s no chapter regarding the organized Right, and that&#8217;s more of a loss.  The organized Right has been very influential in American politics and policy ever since the late 1970s, but it was a big deal even as of 1965.  This book does have one section on &#8220;the white reaction&#8221; in the &#8220;Negroes&#8221; chapter, but that&#8217;s hardly enough to cover it.</p>
<p>7.  There&#8217;s no chapter on the ultra-rich, which in 1965 would pretty much fall in the &#8220;business&#8221; category but now surely deserves its own chapter.</p>
<p>8.  There&#8217;s no chapter on science and technology.  The book is called The Makers of Public Policy, and my point here is that scientists and technologists make public policy even without (or in addition to) traditional lobbying, on issues ranging from vaccines and DNA testing to social media to climate change.</p>
<p>So I think it&#8217;s instructive to look at the topics covered in this sixty-year-old book, partly to see what they missed even then and partly to reflect on how much things have changed.</p>
<p>Also, this won&#8217;t be news to many of you but it&#8217;s still interesting to see:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2025/11/Screenshot-2025-11-26-at-16.03.38.png" alt="" width="400" /></p>
<p>And here&#8217;s the first page of the index:</p>
<p><img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/03/Screenshot-2025-11-26-at-16.06.22.png" alt="" width="500" /></p>
<p>Lots has changed in what&#8217;s considered important.</p>
<p>Also relevant to the question of how things have changed: <a href="https://statmodeling.stat.columbia.edu/2022/08/27/what-rule-did-louis-menand-use-to-select-what-went-into-book-on-u-s-cold-war-culture-of-1945-1965/">What rule did Louis Menand use to select what went into his book on U.S. “cold war” culture of 1945-1965?</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/12/who-were-the-makers-of-public-policy-in-1965/feed/</wfw:commentRss>
			<slash:comments>12</slash:comments>
		
		
			</item>
		<item>
		<title>Survey Statistics: wanting workflow</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/11/survey-statistics-wanting-workflow/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/11/survey-statistics-wanting-workflow/#comments</comments>
		
		<dc:creator><![CDATA[shira]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 20:00:48 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Public Health]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54461</guid>

					<description><![CDATA[Last week Andrew commented that we need a more transparent workflow for survey statistics. So I looked in the new Bayesian Workflow book: Chapter 19 &#8220;Building up to a hierarchical model: Coronavirus testing&#8221; is a case study about a 2020 &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/11/survey-statistics-wanting-workflow/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p><a href="https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/#comment-2417284">Last week Andrew commented</a> that we need a more transparent workflow for survey statistics. So I looked in the <strong>new <a href="https://statmodeling.stat.columbia.edu/2026/04/16/the-bayesian-workflow-book-is-coming/">Bayesian Workflow</a> book</strong>:</p>
<p><img loading="lazy" decoding="async" class="" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/04/9780367490188_cover.jpg" width="272" height="312" /><img loading="lazy" decoding="async" class="alignnone wp-image-54475" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_on_pack_view-scaled.jpg" alt="" width="227" height="299" /></p>
<p>Chapter 19 &#8220;Building up to a hierarchical model: Coronavirus testing&#8221; is a case study about a 2020 survey that tested n = 3330 residents of Santa Clara County, California for SARS-CoV-2 antibodies (Bendavid et al. <a href="https://www.medrxiv.org/content/10.1101/2020.04.14.20062463v1/">2020a</a>, <a href="https://www.medrxiv.org/content/10.1101/2020.04.14.20062463v2/">2020b</a>). y = 50 people tested positive.</p>
<p>To estimate population prevalence, they want to account for measurement error in the test (&#8220;<strong>Measurement</strong>&#8220;) and differences between sample and population (&#8220;<strong>Representation</strong>&#8220;), both sides of <a href="https://www.wiley.com/en-us/Survey+Methodology%2C+2nd+Edition-p-9780470465462">Groves et al.</a> Figure 2.5:</p>
<p><img loading="lazy" decoding="async" class="" src="https://pbs.twimg.com/media/DzqQv0bWoAA52V5?format=jpg&amp;name=4096x4096" alt="Image" width="465" height="617" /></p>
<p>The <strong>measurement error model</strong> includes <strong>specificity</strong> gamma = P[test negative | no disease] and <strong>sensitivity</strong> delta = P[test positive | disease], which take population prevalence pi to test positivity rate p:</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54471" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_measurement_error_model_19_2.png" alt="" width="435" height="70" /></p>
<p>Bendavid et al. (<a href="https://www.medrxiv.org/content/10.1101/2020.04.14.20062463v1/">2020a</a>) analyzed their data using gamma = 0.995 and delta = 0.80. But these aren&#8217;t known exactly, so <a href="https://sites.stat.columbia.edu/gelman/research/published/specificity.pdf">Gelman and Carpenter 2020</a> recommended using priors from previous studies to reflect uncertainty:</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54472" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_gelman_carpenter_priors.png" alt="" width="176" height="59" /></p>
<p>They extend these priors to account for multiple studies. They look at <a href="https://arxiv.org/abs/2107.14054">sensitivity to these priors</a> (overloaded term here: &#8220;sensitivity&#8221;). Cool stuff !</p>
<p>In this Survey Statistics series we hadn&#8217;t yet talked about models for measurement error in the outcome variable y. We&#8217;ve seen measurement error in an adjustment variable X, e.g. recalled vote (<a href="https://statmodeling.stat.columbia.edu/2025/12/23/survey-statistics-is-a-mismeasured-x-better-than-none-at-all/">&#8220;is a mismeasured X better than none at all ?&#8221;</a>, <a href="https://statmodeling.stat.columbia.edu/2025/12/30/survey-statistics-more-adventures-in-mismeasured-x/">&#8220;more adventures in mismeasured X&#8221;</a>, <a href="https://statmodeling.stat.columbia.edu/2026/02/10/survey-statistics-more-on-recalled-vote/">&#8220;more on recalled vote&#8221;</a>, <a href="https://statmodeling.stat.columbia.edu/2026/06/02/survey-statistics-it-is-still-the-people/">&#8220;it is (still) the people&#8221;</a>).</p>
<p>On the &#8220;<strong>Representation</strong>&#8221; side, Bendavid et al. (<a href="https://www.medrxiv.org/content/10.1101/2020.04.14.20062463v1/">2020a</a>, <a href="https://www.medrxiv.org/content/10.1101/2020.04.14.20062463v2/">2020b</a>) adjusted for differences between sample and population using weights and encountered the difficulty we saw last week: <strong>adjusting for lots of variables can lead to very large weights</strong>. Maybe modeling can help (see last week&#8217;s <a href="https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/">&#8220;structured MRP to smooth survey weights&#8221;</a>).</p>
<p>So <a href="https://sites.stat.columbia.edu/gelman/research/published/specificity.pdf">Gelman and Carpenter 2020</a> proposed replacing (19.2) above with</p>
<pre>y_i ~ bernoulli(p_i)
p_i = (1-gamma) (1 - pi_i) + delta pi_i</pre>
<p>and using <strong><a href="https://statmodeling.stat.columbia.edu/2025/06/24/survey-statistics-poststratification/">Multilevel Regression and Poststratification (MRP)</a></strong>:</p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-54473" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_gelman_carpenter_mrp.png" alt="" width="528" height="42" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_gelman_carpenter_mrp.png 528w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_gelman_carpenter_mrp-300x24.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/workflow_book_gelman_carpenter_mrp-500x40.png 500w" sizes="(max-width: 528px) 100vw, 528px" /></p>
<p>I thought this story sounded familiar and remembered Andrew talked about it at his <a href="https://gelman60.com/">Birthday workshop</a>&#8216;s closing talk (<a href="https://youtu.be/Oi7gaOLr3Z4?si=H-Fh6WoVtKgon_aN&amp;t=122">this part on YouTube, minute 2 to minute 8</a>).</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54474" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/youtube_birthday_closing_talk.png" alt="" width="492" height="335" /></p>
<p>Andrew said that when his friend tried the method from <a href="https://sites.stat.columbia.edu/gelman/research/published/specificity.pdf">Gelman and Carpenter 2020</a> it &#8220;crashed and burned&#8221; ! There is a new paper about this: <a href="https://sites.stat.columbia.edu/gelman/research/unpublished/bayes_bad.pdf">Kuh, Kennedy, Chen, and Gelman 2026</a>. They present &#8220;a statistical workflow for diagnosing unexpected results in complex models&#8221;. These authors have researched a lot about evaluating MRP models (see <a href="https://statmodeling.stat.columbia.edu/2026/06/09/survey-statistics-should-mrp-workflow-include-loco-cv/">&#8220;should MRP workflow include LOCO-CV ?&#8221;</a>). We will have to revisit their new paper later in this series.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/11/survey-statistics-wanting-workflow/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>Computer scientists today are in the position of economists in the early 2000s and Freudian psychiatrists in the 1950s</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/11/computer-scientists-today-are-like-economists-in-the-early-2000s-and-freudian-psychiatrists-in-the-1950s/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/11/computer-scientists-today-are-like-economists-in-the-early-2000s-and-freudian-psychiatrists-in-the-1950s/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 13:27:47 +0000</pubDate>
				<category><![CDATA[Economics]]></category>
		<category><![CDATA[Miscellaneous Science]]></category>
		<category><![CDATA[Sociology]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52811</guid>

					<description><![CDATA[Regular readers of this blog will trace my short career as a Freud expert to a post from 2012, Economics now = Freudian psychology in the 1950s: More on the incoherence of “economics exceptionalism”. But recently I thought we need &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/11/computer-scientists-today-are-like-economists-in-the-early-2000s-and-freudian-psychiatrists-in-the-1950s/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Regular readers of this blog will trace my short career as a <a href="https://statmodeling.stat.columbia.edu/2014/05/19/short-career-freud-expert/">Freud expert</a> to a post from 2012, <a href="https://statmodeling.stat.columbia.edu/2012/03/15/economics-now-freudian-psychology-in-the-1950s-more-on-the-incoherence-of-economics-exceptionalism/">Economics now = Freudian psychology in the 1950s: More on the incoherence of “economics exceptionalism”</a>.</p>
<p>But recently I thought we need to update this analogy.</p>
<p>Back in the early part of this century, economists were riding high:  they were the country&#8217;s all-purpose pundits, they had tons of influence but were <a href="https://statmodeling.stat.columbia.edu/2005/09/14/why_are_there_s/">lamenting</a> that they didn&#8217;t have enough, and they were going on and on about how special they were, most amusingly in the self-contradictory argument that they were different because they &#8220;assume everyone is fundamentally alike; we believe circumstances, not culture, drive people’s decisions.&#8221;  I&#8217;m still not sure what is the difference between &#8220;circumstance&#8221; and &#8220;culture&#8221; except that maybe talking about the former is associated with overconfidence.</p>
<p>Nowadays, though, economics is just one more social science.  OK, I don&#8217;t want to overstate things.  I assume they still get paid more than sociologists and political scientists, and, yeah, there&#8217;s a Council of Economic Advisers but no Council of Sociology Advisers.  Still, I think that economics has lost some of its standing in the past twenty years, partly as a result of the crash of 2008 and its aftermath (political polarization, Brexit, etc.) and partly just the natural ebb and flow of influence, the inevitable cycle of hype and disappointment.  Econ hero Steven Levitt was supplanted by data analyst Nate Silver (who identifies as a poker player, not an economist), and we&#8217;re not hearing from economists so much anymore, except to hear them fighting in vain against tariffs.</p>
<p>There&#8217;s a new hegemonic science in town, and it&#8217;s information science, or computer science.</p>
<p>Computer scientists currently have a lot of prestige. Fair enough: they’ve earned it through all the amazing things they’ve built.  Indeed, it&#8217;s worth comparing to the earlier alpha-dog social scientists.  Freudian psychiatry and neoclassical economics were powerful, all-explaining theories that addressed people&#8217;s concerns about mental health, happiness, prosperity, and future prospects.  They had big theories, which is great.  (The theories were not falsifiable, but that can be fine, as the function of such all-encompassing theories is not to make predictions or to explain the world but rather to supply <a href="https://statmodeling.stat.columbia.edu/2019/06/28/racism-is-a-framework-not-a-theory/">a framework</a> by which the social world can be studied.)  Computer science is different: they didn&#8217;t develop a theory of society, but they built impressive tools that change how we live in the world.</p>
<p>Now I&#8217;d say that computer science today is like economics at the beginning of the century or  Freudian psychiatry in the 1950s in being at apex academic and social prestige and influence.  Computer scientists, or computer-science-associated businessmen, are the new gurus, in a way that wasn&#8217;t the case before.  Yes, Steve Jobs and Bill Gates were culture heroes back in the 80s and 90s, but nobody was particularly interested in what they had to say outside of their narrow technological realms.  Nowadays, many computer scientists and tech lords present themselves, and are often taken as, all-purpose pundits.  (Many of the tech lords are actually tech investors; they’ve absorbed the prestige of information science through the transitive property of money.) As with the economists and Freudians of past eras, they are presented as having some special authority derived in part from their almost inhuman hyper-rationality, a willingness to tell uncomfortable truths.</p>
<p>But then you get the same problem we had before, which is when the gurus and their hangers-on start to believe their own hype.</p>
<p>One of the benefits of living a long time is that you get to see the aftermath of the hype wave.  Back when academic economists were on the top of the world, they were <a href="https://statmodeling.stat.columbia.edu/2005/09/14/why_are_there_s/">complaining that they didn&#8217;t have</a> even more power and influence than they already did.  Now that economists are just one more group&#8212;more influential than political scientists or sociologists to be sure, but no longer apex academics&#8212;they&#8217;re no longer doing that.  It&#8217;s when a group is at its most overvalued that <a href="https://statmodeling.stat.columbia.edu/2025/04/29/they-had-it-all-but-they-wanted-more-left-wing-radicals-in-the-1960s-and-right-wingers-now/">it wants even more</a>, and we&#8217;re seeing that with the tech industry (and, to extent, their academic allies) today.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/11/computer-scientists-today-are-like-economists-in-the-early-2000s-and-freudian-psychiatrists-in-the-1950s/feed/</wfw:commentRss>
			<slash:comments>44</slash:comments>
		
		
			</item>
		<item>
		<title>I recommend this book on Applied Regression and Causal Inference</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/10/i-recommend-this-book-on-applied-regression-and-causal-inference/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/10/i-recommend-this-book-on-applied-regression-and-causal-inference/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Mon, 10 Aug 2026 13:41:53 +0000</pubDate>
				<category><![CDATA[Causal Inference]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52335</guid>

					<description><![CDATA[This post offers two recommendations and a story. First, the recommendations. • Here&#8217;s the book referred to in the title of the post. I highly recommend it! My only regret is that we didn&#8217;t call it Applied Regression and Causal &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/10/i-recommend-this-book-on-applied-regression-and-causal-inference/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>This post offers two recommendations and a story.</p>
<p>First, the recommendations.</p>
<p>• <a href="https://sites.stat.columbia.edu/gelman/regression/">Here&#8217;s the book</a> referred to in the title of the post.  I highly recommend it!  My only regret is that we didn&#8217;t call it Applied Regression and Causal Inference, cos that&#8217;s what it&#8217;s about.</p>
<p>• I also recommend our book, <a href="https://sites.stat.columbia.edu/gelman/active-statistics/">Active Statistics</a>, which has literally hundreds of stories, class-participation activities, computer demonstrations, and discussion problems.  A great teaching resource and also just a good read.  Really!</p>
<p>Now, the story.</p>
<p>This came in the email one day:</p>
<blockquote><p>Dear Professor Andrew Gelman,</p>
<p>I hope this message finds you well. I’m part of the Cambridge University Press sales support team and I work with your local Cambridge Sales Representative ** **@cambridge.org</p>
<p>We wanted to follow up regarding your course(s) for Fall 2025 and ensure you’ve had a chance to review Regression and Other Stories, 1 ed by Andrew Gelman/Jennifer Hill/Aki Vehtari as a potential option for your course materials.</p>
<p>If you’ve decided to adopt this title, I’d be happy to provide access to any available instructor resources. If you&#8217;ve chosen a different title or are still considering your options, I’d appreciate a quick update so I can ensure our records are accurate.</p>
<p>Also, please don’t forget to check with your bookstore about any Equitable or Inclusive Access programs that may be available. These programs can provide your students with more affordable digital access to their course materials.</p>
<p>I am sincerely grateful for your attention to this matter; I look forward to the opportunity to hear from you at your earliest convenience.</p>
<p>Cordially,</p>
<p>**<br />
Sales Support Assistant<br />
Higher Education Sales, Americas<br />
**</p></blockquote>
<p>On one hand, hey, I&#8217;m thrilled that they&#8217;re promoting our book.  On the other hand, it seems like they&#8217;re kinda flying blind if they&#8217;re sending this promotional message to the author of the book.</p>
<p><img loading="lazy" decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/download.png" alt="" width="180" height="237" class="alignnone size-full wp-image-54457" /> <img decoding="async" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/active_statistic.jpg" alt="" width="170" /></p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/10/i-recommend-this-book-on-applied-regression-and-causal-inference/feed/</wfw:commentRss>
			<slash:comments>5</slash:comments>
		
		
			</item>
		<item>
		<title>People keep sending me AI slop that they want me to post on the blog.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/09/people-keep-sending-me-ai-slop-that-they-want-me-to-post-on-the-blog/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/09/people-keep-sending-me-ai-slop-that-they-want-me-to-post-on-the-blog/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Mon, 10 Aug 2026 00:25:53 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54451</guid>

					<description><![CDATA[I don&#8217;t get it. Instead of sending me the chatbot output, just send me your goddamn prompt. Cut out the middleman.]]></description>
										<content:encoded><![CDATA[<p>I don&#8217;t get it.  Instead of sending me the chatbot output, just send me your goddamn prompt.  Cut out the middleman.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/09/people-keep-sending-me-ai-slop-that-they-want-me-to-post-on-the-blog/feed/</wfw:commentRss>
			<slash:comments>22</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;Sane-washing&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/09/sane-washing/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/09/sane-washing/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 09 Aug 2026 13:46:23 +0000</pubDate>
				<category><![CDATA[Political Science]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=50964</guid>

					<description><![CDATA[After quoting Donald Trump at a political rally with Robert Kennedy Jr., That is why today I am repeating my pledge to establish a panel of top experts, working with Bobby, to investigate what is causing the decades-long increase in &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/09/sane-washing/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>After quoting Donald Trump at a political rally with Robert Kennedy Jr.,</p>
<blockquote><p>That is why today I am repeating my pledge to establish a panel of top experts, working with Bobby, to investigate what is causing the decades-long increase in chronic health problems and childhood diseases, including auto-immune disorders, autism, obesity, infertility, and more.</p></blockquote>
<p><a href="https://observationalepidemiology.blogspot.com/2024/08/trump-kennedy-and-autism-all-news-thats.html">Palko writes</a>:</p>
<blockquote><p>Coming at this from a conspiracy theorist/pseudoscience crank worldview, the one word that should set off immediate alarm bells is &#8220;autism.&#8221; . . . If you had to list the remaining terms in order of tinfoil hat appeal, they would be infertility, autoimmune disorders, and at a distant fourth, obesity. A reporter covering the speech might also feel compelled to add that the causes which RFK Junior has proposed in the past include not just vaccines but also chemtrails, Wi-Fi, and 5G.</p></blockquote>
<p>Sounds about right:  the candidate who endorses the ludicrous Pizzagate conspiracy theory, tried to overturn an election, and is allied with child-murder-denialist Alex Jones is teaming up with a public figure whose big political issue is opposing vaccines.</p>
<p>Palko continues:</p>
<blockquote><p>Here&#8217;s how the NYT covered that quote in the third paragraph from the end of the article on Kennedy endorsing Trump:</p>
<blockquote><p>Earlier in the day in Phoenix, at his speech announcing the suspension of his campaign, Mr. Kennedy said Mr. Trump had offered him a role in a second Trump administration, dealing with health care and food and drug policy. In Glendale, Mr. Trump said that, if elected to a second term, a panel of experts “working with Bobby” would investigate obesity rates and other chronic health issues in the United States.</p></blockquote>
</blockquote>
<p>I agree with Palko that this is absolutely ridiculous in the context of the actual quote from Trump.  The Times is &#8220;sane-washing&#8221; Trump and Kennedy, working really hard to take dangerous and extreme positions and treat them as reasonable and normal.</p>
<p>I guess this doesn&#8217;t matter to political junkies for whom Kennedy Jr. is associated with anti-vaxx conspiracy theories.  To many normies, though, I&#8217;m guessing that he&#8217;s perceived more as a generic Kennedy heir&#8212;maybe they don&#8217;t know that he makes Dr. Oz and Andrew Huberman look like purveyors of solid medical advice.</p>
<p>This seems like<br />
<a href="https://statmodeling.stat.columbia.edu/2020/03/06/junk-science-then-and-now/">another example of</a> how junk science and conspiracy theories have moved from the periphery of our culture to the mainstream, and I agree with Palko that it was bad journalism for the Times to report on that speech and sane-wash it.</p>
<p>Newspapers make mistakes all the time; the reasons for posting this one are:  (1) this was the first time I&#8217;d heard of &#8220;sane-washing,&#8221; and (2) it&#8217;s scary to see pseudoscience escape the narrow world of Lancet/PNAS/Ted/NPR/Gladwell/Freakonomics and potentially become government policy.</p>
<p>I do think it&#8217;s possible that the <a href="https://statmodeling.stat.columbia.edu/2024/10/18/its-a-very-short-jump-from-believing-kale-smoothies-are-a-cure-for-cancer-to-denying-the-holocaust-happened/">mainstreaming</a> of pseudoscience into elite discourse through the Lancet/PNAS/Ted/NPR/Gladwell/Freakonomics nexus has made it easier for pseudoscience like election denial and vaccine denial to get elite traction (even if it&#8217;s different subset of elites).</p>
<p>On the other hand, <a href="https://statmodeling.stat.columbia.edu/2022/09/11/merchants-of-doubt-operating-in-organized-science/">Merchants of Doubt</a>, etc:  <a href="https://statmodeling.stat.columbia.edu/2012/09/02/cigarettes/">Cigarette companies</a> and fossil fuel companies have been working for many decades to blur the lines between science and propaganda, so maybe this would all be happening even without all the junk science that entered the mainstream in the replication-crisis era.  I don&#8217;t know.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/09/sane-washing/feed/</wfw:commentRss>
			<slash:comments>51</slash:comments>
		
		
			</item>
		<item>
		<title>“Denial, when you are not part of it, is actually a terrifying thing. One watches one’s fellow humans doing things that will damage themselves, while being wholly unable to help.”</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/08/denial-when-you-are-not-part-of-it-is-actually-a-terrifying-thing-one-watches-ones-fellow-humans-doing-things-that-will-damage-themselves-while-being-wholly-unable-to-help/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/08/denial-when-you-are-not-part-of-it-is-actually-a-terrifying-thing-one-watches-ones-fellow-humans-doing-things-that-will-damage-themselves-while-being-wholly-unable-to-help/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sat, 08 Aug 2026 13:48:35 +0000</pubDate>
				<category><![CDATA[Sociology]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=52450</guid>

					<description><![CDATA[Picking up a thread from awhile ago, another item from my review of that Daniel Davies book: Davies writes, “Denial, when you are not part of it, is actually a terrifying thing. One watches one’s fellow humans doing things that &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/08/denial-when-you-are-not-part-of-it-is-actually-a-terrifying-thing-one-watches-ones-fellow-humans-doing-things-that-will-damage-themselves-while-being-wholly-unable-to-help/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Picking up a thread from awhile ago, <a href="https://statmodeling.stat.columbia.edu/2023/07/07/cheating-in-science-sports-journalism-business-and-art-how-do-they-differ/">another item from my review</a> of that Daniel Davies book:</p>
<p>Davies writes, “Denial, when you are not part of it, is actually a terrifying thing. One watches one’s fellow humans doing things that will damage themselves, while being wholly unable to help.”</p>
<p>I agree. This is how I felt when corresponding with the ovulation-and-clothing researchers and with the elections-and-lifespan researchers, and other examples too. The people on the other side of these discussions seemed perfectly sincere; they just couldn’t consider the possibility they might be on the wrong track. (You could say the same about me, except: (1) I did consider the possibility I could be wrong in these cases, and (2) there were statistical arguments on my side; these weren’t just matters of opinion.) Anyway, setting aside if I was right or wrong in these disputes, the denial (as I perceived it) really upset me.  Seeing young researchers, with their whole careers ahead of them, turning their eyes away from reality, it just makes me want to cry.</p>
<p>I don’t think graduate students are well trained in handling mistakes&#8212;indeed, they&#8217;re <a href="https://statmodeling.stat.columbia.edu/2018/01/13/solution-puzzle-scientists-typically-respond-legitimate-scientific-criticism-angry-defensive-closed-non-scientific-way/">effectively trained to respond</a> to criticism by deflecting it, and then when they grow up and publish research, they can remain stuck in this attitude of denial.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/08/denial-when-you-are-not-part-of-it-is-actually-a-terrifying-thing-one-watches-ones-fellow-humans-doing-things-that-will-damage-themselves-while-being-wholly-unable-to-help/feed/</wfw:commentRss>
			<slash:comments>13</slash:comments>
		
		
			</item>
		<item>
		<title>How to avoid the &#8220;clean data, clean model&#8221; trap when teaching statistics and data science</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/07/how-to-avoid-the-clean-data-clean-model-trap-when-teaching-statistics-and-data-science/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/07/how-to-avoid-the-clean-data-clean-model-trap-when-teaching-statistics-and-data-science/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 07 Aug 2026 13:30:45 +0000</pubDate>
				<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53438</guid>

					<description><![CDATA[Parthsarthi Joshi writes: There is a stark problem that I have observed with how data science is taught in most books and tutorials &#8211; the concepts are taught on an individual level but fail to give a complete idea. The &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/07/how-to-avoid-the-clean-data-clean-model-trap-when-teaching-statistics-and-data-science/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Parthsarthi Joshi writes:</p>
<blockquote><p>There is a stark problem that I have observed with how data science is taught in most books and tutorials &#8211; the concepts are taught on an individual level but fail to give a complete idea. The so-called “real-world” projects are carefully orchestrated like an “act” already setup by the instructor. The students are not given the taste of the real world – where data is messy, models do not continuously improve, and assumptions are almost always broken. We teach them the concepts of statistics and ML but not the art of data science, which includes curiosity and hit-and-trial experiments.</p>
<p>I am writing a one-of-its-kind book on Data Science that would be completely based on projects. The data will be messy, not everything we do will improve the performance. The aim is to give a feel of uncertainty and challenges that promotes learning from failure and develops an analytical bent of mind.</p></blockquote>
<p>I agree with Joshi&#8217;s concerns, and it&#8217;s something that&#8217;s bothered me for a long time.  Textbooks and research articles are full of examples that work, they&#8217;re full of estimates that are a bit more than 2 standard errors from zero, everything&#8217;s a bit too clean.</p>
<p>I&#8217;ve tried to push against this in my own teaching and writing.  Here are some examples:</p>
<p>1.  Decision analysis textbooks often use examples that are trivial and artificial (a college student is trying to choose an apartment or dorm room, balancing the factors of cost, amenities, and commute time) or important but vague (a company is deciding where or whether to build a new plant, it&#8217;s an interesting problem of decision making under uncertainty but the details are not given).  A bit reason I wanted to do <a href="https://sites.stat.columbia.edu/gelman/research/published/lin.pdf">the full decision analysis</a> with Phil for the home radon problem was that I wanted a real problem, with real data, going from beginning to end.</p>
<p>2.  A few years ago there was a heart-stent study reported in the news.  The experiment yielded results that were not statistically significant (the p-value was 0.20) and it was inappropriately reported as a null effect.  We looked at the paper carefully, and a more reasonable analysis gave a p-value of 0.09&#8211;still not reaching conventional levels of significance.  <a href="https://sites.stat.columbia.edu/gelman/research/published/Stents_published.pdf">We wrote this up</a>, and it&#8217;s a good example where the result is not clean.  Again, something you don&#8217;t usually see much discussion of in textbooks, applied research papers, or articles on research methods.</p>
<p>3.  Our book <a href="https://sites.stat.columbia.edu/gelman/active-statistics/">Active Statistics</a> is full of stories with real-world complexity.</p>
<p>4.  Appendix A of <a href="https://sites.stat.columbia.edu/gelman/regression/">Regression and Other Stories</a>, on computing in R, has examples where you have to download and clean the data.  One problem with a lot of computer-friendly textbooks is that they&#8217;ll have clean data already loaded into packages so you can just click and do the analysis.  I like there to be a bit of a struggle, to give a sense of what data analysis really feels like.</p>
<p>We try also to convey a sense of real-world complexity in the worked examples and homework problems in our textbooks.</p>
<p>The flip side of this is that, when you&#8217;re using live problems in your teaching, it&#8217;s good to come to class prepared!  It&#8217;s fine to show students the steps of data cleaning and the frustration of messy data and models that don&#8217;t fit&#8211;but it&#8217;s best for you, the teacher, to encounter and resolve those problems on your own first, so that you don&#8217;t confuse the students even further.  In class it&#8217;s good to have a sense of what&#8217;s coming next, so that when something doesn&#8217;t work, you&#8217;re not surprised and you can contextualize it for the students right away.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/07/how-to-avoid-the-clean-data-clean-model-trap-when-teaching-statistics-and-data-science/feed/</wfw:commentRss>
			<slash:comments>14</slash:comments>
		
		
			</item>
		<item>
		<title>The blessing of dimensionality</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/06/the-blessing-of-dimensionality/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/06/the-blessing-of-dimensionality/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 06 Aug 2026 13:45:17 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53433</guid>

					<description><![CDATA[Jonathan &#8220;No Trump&#8221; Falk writes: Your post this week on Integration and Differentiation got me thinking, never a good sign. Adding to this, I only really learned in January what Attention means, the foundation of LLM. And what it means &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/06/the-blessing-of-dimensionality/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Jonathan <a href="https://sites.stat.columbia.edu/gelman/research/unpublished/notrump_falk_gelman_icml.pdf">&#8220;No Trump&#8221;</a> Falk writes:</p>
<blockquote><p>Your post this week <a href="https://statmodeling.stat.columbia.edu/2026/03/14/the-paradox-of-derivatives-and-integrals/">on Integration and Differentiation</a> got me thinking, never a good sign.  Adding to this, I only really learned in January what Attention means, the foundation of LLM.  And what it means is that in a vector space of 14,000 dimensions or so, you can express all manner of nuance&#8230; enough to start vitriolic arguments about the humanity of the output.</p>
<p>As I started thinking about this, and as your post crystallized, I have spent 50 years fighting the Curse of Dimensionality.  I know this curse in my marrow.  Brilliant inferences await me, but the space in which these insights are found is simply too vast to explore.  So we simplify, reducing the dimensionality to something that while still vast, is confined to a hyperplane where we can, like Plato, see the projections of truth, not the truth itself.</p>
<p>But then what LLMs and their generation have taught me is that nuance, which is really just the inverse of inference (in that it&#8217;s the vast set of all things consistent with some inference) has an amazing boon of dimensionality.  There appears to be no thought that can&#8217;t be described by a 14,000 dimension or so vector whose tuning has the huge advantage that 14,000-dimensional space is so empty that tiny nuances can be readily distinguished in such a space, so that you can hide uniqueness in the vastness of 14000-dimensional space that you couldn&#8217;t recover in a raw search in that same space.</p>
<p>I&#8217;m sure this inversion of the Curse of dimensional search into the Boon of nuance in high dimensions is not original to me, but I think it&#8217;s worth noting, and your post was the impetus.</p></blockquote>
<p>I replied by pointing to our of our very earliest blog posts, <a href="https://statmodeling.stat.columbia.edu/2004/10/27/the_blessing_of/">The blessing of dimensionality</a>, where I wrote:</p>
<blockquote><p>The phrase “curse of dimensionality” has many meanings (with 18800 references, it loses to “bayesian statistics” in a googlefight, but by less than a factor of 3). In numerical analysis it refers to the difficulty of performing high-dimensional numerical integrals.</p>
<p>But I am bothered when people apply the phrase “curse of dimensionality” to statistical inference.</p>
<p>In statistics, “curse of dimensionality” is often used to refer to the difficulty of fitting a model when many possible predictors are available. But this expression bothers me, because more predictors is more data, and it should not be a “curse” to have more data. Maybe in practice it’s a curse to have more data (just as, in practice, giving people too much good food can make them fat), but “curse” seems a little strong.</p>
<p>With multilevel modeling, there is no curse of dimensionality. When many measurements are taken on each observation, these measurements can themselves be grouped. Having more measurements in a group gives us more data to estimate group-level parameters (such as the standard deviation of the group effects and also coefficients for group-level predictors, if available).</p>
<p>In all the realistic “curse of dimensionality” problems I’ve seen, the dimensions–the predictors–have a structure. The data don’t sit in an abstract K-dimensional space; they are units with K measurements that have names, orderings, etc.</p>
<p>For example, Marina gave us an example in the seminar the other day where the predictors were the values of a spectrum at 100 different wavelengths. The 100 wavelengths are ordered. Certainly it is better to have 100 than 50, and it would be better to have 50 than 10. (This is not a criticism of Marina’s method, I’m just using it as a handy example.)</p>
<p>For an analogous problem: 20 years ago in Bayesian statistics, there was a lot of struggle to develop noninformative prior distributions for highly multivariate problems. Eventually this line of research dwindled because people realized that when many variables are floating around, they will be modeled hierarchically, so that the burden of noninformativity shifts to the far less numerous hyperparameters. And, in fact, when the number of variables in a a group is larger, these hyperparameters are easier to estimate.</p>
<p>I’m not saying the problem is trivial or even easy; there’s a lot of work to be done to spend this blessing wisely.
</p></blockquote>
<p>It&#8217;s been over 20 years so worth sharing the point again.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/06/the-blessing-of-dimensionality/feed/</wfw:commentRss>
			<slash:comments>7</slash:comments>
		
		
			</item>
		<item>
		<title>Manned Mars Mission Miscellanea</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/05/manned-mars-mission-miscellanea/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/05/manned-mars-mission-miscellanea/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 05 Aug 2026 13:05:40 +0000</pubDate>
				<category><![CDATA[Miscellaneous Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=51729</guid>

					<description><![CDATA[Maciej Cegłowski writes: Unlike the Moon, which hangs in the sky like a lonely grandparent waiting for someone to visit, Mars leads a rich orbital life of its own and is not always around to entertain the itinerant astronaut. There &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/05/manned-mars-mission-miscellanea/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Maciej Cegłowski <a href="https://idlewords.com/2025/02/the_shape_of_a_mars_mission.htm">writes</a>:</p>
<blockquote><p>Unlike the Moon, which hangs in the sky like a lonely grandparent waiting for someone to visit, Mars leads a rich orbital life of its own and is not always around to entertain the itinerant astronaut. There is just one brief window every 26 months when travel between our two planets is feasible, and this constraint of orbital mechanics is so fundamental that we’ve known since Lindbergh crossed the Atlantic what a mission to Mars must look like. . . .</p>
<p>We shouldn’t send human beings to Mars, at least not anytime soon. Landing on Mars with existing technology would be a destructive, wasteful stunt whose only legacy would be to ruin the greatest natural history experiment in the Solar System. It would no more open a new era of spaceflight than a Phoenician sailor crossing the Atlantic in 500 B.C. would have opened up the New World. And it wouldn’t even be that much fun. . . .</p>
<p>It wasn’t always like this. There was a time when going to Mars made sense, back when astronauts were a cheap and lightweight alternative to costly machinery, and the main concern about finding life on Mars was whether all the trophy pelts could fit in the spacecraft. No one had been in space long enough to discover the degenerative effects of freefall, and it was widely accepted that not just exploration missions, but complicated instruments like space telescopes and weather satellites, were going to need a permanent crew.</p>
<p>But fifty years of progress in miniaturization and software changed the balance between robots and humans in space. Between 1960 and 2020, space probes improved by something like six orders of magnitude, while the technologies of long-duration spaceflight did not. Boiling the water out of urine still looks the same in 2023 as it did in 1960, or for that matter 1060. . . .</p>
<p>Mars is also not the planet we took it for. . . . The surface might be dry, but in most places there was water ice just underneath. Dynamic surface features hinted that water (or at least brine) was flowing to the surface from deep underground. . . . The news from the ground also got better. Arriving at Gale Crater in 2012, the Curiosity rover found itself looking at an ordinary lake bed, complete with organic sediment and odd stick-like structures that would be called fossils if we found them on Earth. The crater had been habitable for millions of years in the past, and something in it was still emitting methane at night. Over in its own crater, the Perseverance rover found complex organic molecules of indeterminate origin.</p>
<p>But the really exciting news for Mars was the discovery of unexpected life on Earth. . . . not just dozens of unsuspected microbial phyla, but two entire new branches of life . . . These new techniques confirmed that earth’s crust is inhabited to a depth of kilometers by a ‘deep biosphere’ of slow-living microbes nourished by geochemical processes and radioactive decay. . . . This underground ecology, which we have barely started to explore, might account for a third of the biomass on earth.</p>
<p>The fact that we failed to notice 99.999% of life on Earth until a few years ago is unsettling and has implications for Mars. The existence of a deep biosphere in particular narrows the habitability gap between our planets to the point where it probably doesn’t exist—there is likely at least one corner of Mars that an Earth organism could call home. . . . if our distant relatives are still alive in some deep Martian cave, then just about the worst way to go looking for them would be to land in a septic spacecraft.</p>
<p>But the fact that a Mars landing stopped making sense has not had the slightest impact on NASA’s plan to go there in a rocket-propelled terrarium.</p></blockquote>
<p>And more:</p>
<blockquote><p>The chief technical obstacle to a Mars landing is not propulsion, but a lack of reliable closed-loop life support. . . . The technology program required to close this gap would be remarkably circular, with no benefits outside the field of applied zero gravity zookeeping. The web of Rube Goldberg devices that recycles floating animal waste on the space station has already cost twice its weight in gold and there is little appetite for it here on Earth, where plants do a better job for free. I would compare keeping primates alive in spacecraft to trying to build a jet engine out of raisins. Both are colossal engineering problems, possibly the hardest ever attempted, but it does not follow that they are problems worth solving. In both cases, the difficulty flows from a very specific design constraint, and it’s worth revisiting that constraint one or ten times before starting to perform miracles of engineering. . . . The only way to explore Mars in our lifetime is to ditch the requirement that people accompany the machinery. . . . </p>
<p>In recent years, there’s been a remarkable division in space exploration. On one side of the divide are missions like Curiosity, James Webb, Gaia, or Euclid that are making new discoveries by the day. These projects have clearly defined goals and a formidable record of discovery.</p>
<p>On the other side, there is the International Space Station and the now twenty-year old effort to return Americans to the moon. These projects have no purpose other than perpetuating a human presence in space, and they eat through half the country’s space budget with nothing to show for it. Forget even Mars—we are further from landing on the Moon today than we were in 1965.</p>
<p>In going to Mars, we have a choice about which side of this ledger to be on.</p></blockquote>
<p>This all makes sense.  I&#8217;ve never thought much about this Mars mission thing because it&#8217;s always seemed like a bit of <a href="https://statmodeling.stat.columbia.edu/2015/12/15/mars-1-this-american-life-0/">a joke</a>.  But if powerful people are really gonna try to use this as pretext to take a big chunk out of our national resources, then, yeah, it&#8217;s good to have people like Cegłowski pushing back.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/05/manned-mars-mission-miscellanea/feed/</wfw:commentRss>
			<slash:comments>45</slash:comments>
		
		
			</item>
		<item>
		<title>Survey Statistics: structured MRP to smooth survey weights</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/#comments</comments>
		
		<dc:creator><![CDATA[shira]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 22:48:21 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Multilevel Modeling]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54340</guid>

					<description><![CDATA[Last week, Raphael K shared a concern: adjusting for lots of variables can lead to very large weights. So today let&#8217;s dive into Si et al. 2020, who saw this in constructing survey weights for the NYC Longitudinal Study of &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Last week, <a href="https://statmodeling.stat.columbia.edu/2026/07/28/survey-statistics-equivalent-models-equivalent-weights-locally/#comment-2417040">Raphael K shared a concern</a>: adjusting for lots of variables can lead to very large weights. So today let&#8217;s dive into <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a>, who saw this in constructing survey weights for the <a href="https://cprc.columbia.edu/content/new-york-city-longitudinal-survey-wellbeing">NYC Longitudinal Study of Wellbeing</a>.</p>
<p>To adjust for lots of variables, <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> turned to <a href="https://statmodeling.stat.columbia.edu/2018/05/19/regularized-prediction-poststratification-generalization-mister-p/"><strong>MRP</strong> (Multilevel Regression and Poststratification)</a> and <strong>equivalent weights</strong> based on these models (see <a href="https://statmodeling.stat.columbia.edu/2025/10/07/survey-statistics-struggles-with-equivalent-weights/">“struggles with equivalent weights”</a>, <a href="https://statmodeling.stat.columbia.edu/2025/11/04/survey-statistics-continued-struggles-with-equivalent-weights/#comment-2416983">continued struggles</a>, and <a href="https://statmodeling.stat.columbia.edu/2026/07/28/survey-statistics-equivalent-models-equivalent-weights-locally/">&#8220;equivalent models, equivalent weights (locally)&#8221;</a>).</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54382" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-scaled.jpg" alt="" width="349" height="263" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-scaled.jpg 2560w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-300x225.jpg 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-1024x768.jpg 1024w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-768x576.jpg 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-1536x1152.jpg 1536w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-2048x1536.jpg 2048w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Doobie_TN_AT_May_8_2026_view_higher_up-400x300.jpg 400w" sizes="(max-width: 349px) 100vw, 349px" /></p>
<p>In a simulation study they compare:</p>
<ul>
<li><strong>Ind-P</strong>: MRP with commonly-used Independent Normal priors</li>
<li><strong>Str-P</strong>: MRP with a structured prior, see below.</li>
<li><strong>Ind-W</strong>: equivalent weights version of Ind-P</li>
<li><strong>Str-W</strong>: equivalent weights version of Str-P</li>
<li><strong>Rake-W</strong>: classical raking weights, a type of calibrated weights</li>
<li><strong>PS-W</strong>: classical poststratification weights, another type of calibrated weights</li>
<li><strong>IP-W</strong>: inverse probability of selection weights</li>
</ul>
<p>They cover the <a href="https://statmodeling.stat.columbia.edu/2025/06/17/survey-statistics-3-flavors-of-survey-weights/">&#8220;3 flavors of survey weights&#8221;</a>: equivalent weights, calibrated weights, IP-W.</p>
<p>I won&#8217;t bury the lead, they found <strong>MRP performed best, then equivalent weights, then classical weights (calibrated or IP-W)</strong>. See their Figure 4.1 for the simulation scenario without terribly many empty poststratification cells:</p>
<p><img loading="lazy" decoding="async" class="alignnone size-full wp-image-54379" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_4_1_Si_et_al_2020.png" alt="" width="665" height="483" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_4_1_Si_et_al_2020.png 665w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_4_1_Si_et_al_2020-300x218.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_4_1_Si_et_al_2020-413x300.png 413w" sizes="(max-width: 665px) 100vw, 665px" /></p>
<p>With many empty poststratification cells, Str-P outperforms Ind-P. (They don&#8217;t redo Figure 4.1 for this scenario, which confused me a bit.) So what is this structure that helps ?</p>
<p>In <a href="https://statmodeling.stat.columbia.edu/2026/04/07/survey-statistics-improving-with-structure/">&#8220;improving with structure&#8221;</a> we saw that <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9203002/pdf/nihms-1811398.pdf">Gao et al. 2021</a> found it helpful to use the ordinal structure of variables like age. <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> use the <strong>interaction structure</strong>:</p>
<blockquote>
<p class="p1">We induce structured prior distributions to be able to handle deep interactions and account for their hierarchy structure, where the high-order interaction terms will be excluded if one of the corresponding main effects is not selected.</p>
</blockquote>
<p>I asked about sparse priors for MRP back in <a href="https://statmodeling.stat.columbia.edu/2025/07/01/survey-statistics-sparsified-mrp/">&#8220;Sparsified MRP&#8221;</a>. I didn&#8217;t remember that <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> had worked on this ! Ok so they write their structure more generally but I find it easier to read with a specific example. Consider just 2 variables from their motivating <a href="https://cprc.columbia.edu/content/new-york-city-longitudinal-survey-wellbeing">NYC Longitudinal Study of Wellbeing</a>: age (5 categories) and race (5 categories). Here&#8217;s how their Ind-P prior differs from the Str-P:</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54380" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Ind-P-vs-Str-P.png" alt="" width="573" height="275" /></p>
<p>(I had Claude type up my hand-drawn notes, <a href="https://statmodeling.stat.columbia.edu/2024/02/26/hand-drawn-statistical-workflow-at-nelson-mandela/">though I still share Brendan Leonard&#8217;s preference for hand-drawn materials</a>.)</p>
<p><a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> say these are <strong>similar to the Horseshoe prior</strong>. It differs in two ways, I think ? First, <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> have the selection at the batch level (e.g. age or race). The usual Horseshoe would have local scale lambdas for each age and race category. Second, the usual Horseshoe would use a half-Cauchy rather than half-Normal prior on these local scales.</p>
<p><img loading="lazy" decoding="async" class="alignnone wp-image-54381" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Horseshoe.png" alt="" width="723" height="133" /></p>
<p>Ok let&#8217;s get back to the original concern: adjusting for lots of variables can lead to very large weights. <a href="https://sites.stat.columbia.edu/gelman/research/published/survey_methodology.pdf">Si et al. 2020</a> show in Figure 5.1 that the equivalent weights based on this structured prior model look much less variable than calibration weights (I don&#8217;t see the IP-W weights in the figure itself): <img loading="lazy" decoding="async" class="alignnone size-full wp-image-54383" src="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_5_1_Si_et_al_2020.png" alt="" width="798" height="508" srcset="https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_5_1_Si_et_al_2020.png 798w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_5_1_Si_et_al_2020-300x191.png 300w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_5_1_Si_et_al_2020-768x489.png 768w, https://statmodeling.stat.columbia.edu/wp-content/uploads/2026/08/Figure_5_1_Si_et_al_2020-471x300.png 471w" sizes="(max-width: 798px) 100vw, 798px" /></p>
<p>&nbsp;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/04/survey-statistics-structured-mrp-to-smooth-survey-weights/feed/</wfw:commentRss>
			<slash:comments>4</slash:comments>
		
		
			</item>
		<item>
		<title>Walnutpie version 0.0.1 Released</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/04/walnutpie-version-0-0-1-released/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/04/walnutpie-version-0-0-1-released/#comments</comments>
		
		<dc:creator><![CDATA[Bob Carpenter]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 20:42:46 +0000</pubDate>
				<category><![CDATA[Stan]]></category>
		<category><![CDATA[Statistical Computing]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54372</guid>

					<description><![CDATA[We are happy to announce the official release of Walnutpie version 0.0.1. Walnutpie is an MCMC sampler for continuously differentiable densities coded in Python, accepting models coded in Stan, PyMC, NumPyro, JAX, and plain old Python. Walnutpie is not an &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/04/walnutpie-version-0-0-1-released/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>We are happy to announce the official release of <b>Walnutpie version 0.0.1</b>.</p>
<p>Walnutpie is an MCMC sampler for continuously differentiable densities coded in Python, accepting models coded in Stan, PyMC, NumPyro, JAX, and plain old Python.</p>
<p><b>Walnutpie is <i>not</i> an official Stan project</b></p>
<p>I thought this was worth saying up front. It may eventually migrate to Stan, but for now, we followed the Nutpie approach of building a standalone sampling package that worked with a range of packages for defining models.</p>
<p><b>R version</b></p>
<p>We plan to develop an R interface after we release version 1.0.0 of the Python interface.  So hopefully in 2026.</p>
<p><b>pip installable</b></p>
<p>Walnutpie is on <a href="https://pypi.org">PyPI</a>, so it&#8217;s <a href="https://pypi.org/project/pip/">pip installable</a>. The documentation includes information on getting started, running models, and posterior analysis.  We have not yet included case studies for modeling tools other than Stan and Python.</p>
<ul>
<li><a href="https://flatironinstitute.github.io/walnutpie/latest/install.html">Walnutpie documentation</a></li>
</ul>
<p><b>Stan through C++</b></p>
<p>Walnutpie runs Stan models through C++ using <a href="https://roualdes.us/bridgestan/latest/">BridgeStan</a>, so there is no Python dispatch overhead for Stan sampling. The basic architecture of the API is based on Adrian Seyboldt&#8217;s sampler <a href="https://pymc-devs.github.io/nutpie/">Nutpie</a>. We are working on doing that for other packages like NumPy and JAX to the extent that we can.</p>
<p><b>GitHub source</b></p>
<p>Development discussions and source code are managed through GitHub.</p>
<ul>
<li><a href="https://flatironinstitute.github.io/walnutpie/latest/">Walnutpie on GitHub</a></li>
</ul>
<p><b>Features of Walnutpie</b></p>
<p>We are almost ready to release the paper on arXiv with all of the gory pseudocode details of all of the algorithms used for Walnutpie. It will explain the following points in detail.</p>
<ol>
<li>
<p><b>Walnuts</b>: The underlying Hamiltonian Monte Carlo sampler is <a href="https://www.jmlr.org/papers/volume27/25-1452/25-1452.pdf">Walnuts</a>. Walnuts uses Nuts for choosing the number of steps per iteration. It further allows step sizes within the Hamiltonian dynamics simulation to be lowered when necessary to preserve simulation accuracy. This helps with robustness and with accuracy in multi-scale distributions (i.e., ones where the curvature as represented by the Hessian varies around the posterior). With a high tolerance threshold, Walnuts reverts to Nuts&#8217;s behavior.</p>
</li>
<li>
<p><b>Mass-matrix warmup</b>: The mass-matrix warmup strategy is an online form of Nutpie (links to: the <a href="https://arxiv.org/abs/2603.18845v1">paper</a> and <a href="https://pymc-devs.github.io/nutpie/">software</a>). Nutpie minimizes Fisher divergence by estimating the inverse mass matrix as the midpoint (in the appropriate manifold) between an estimate based on the variance of the draws and the covariance of the scores (gradients of the log density). The target is better than Nuts&#8217;s variance of draws in both convergence speed and sampling efficiency. Walnutpie only supports diagonal mass matrices (Nuts supports dense matrices and Nutpie supports low-rank plus diagonal and even more general normalizing flow approaches). The approach is online in the sense that it is not blocked like warmup in Nuts or Nutpie&mdash;it updates every iteration by exponentially discounting the past to mimic Stan&#8217;s exponentially increasing history sizes. We also borrow Nutpie&#8217;s mass matrix initialization based on a regularized outer product of gradients at the initial point.</p>
</li>
<li>
<p><b>Step-size warmup</b>: The step size adaptation strategy has not changed, but the underlying stochastic gradient descent algorithm is <a href="https://en.wikipedia.org/wiki/Stochastic_gradient_descent#Adam">Adam</a> rather than dual averaging. We found Adam to be faster to converge and much more stable. Matt Hoffman included a hack in the original Nuts approach to stabilize dual averaging, but even with that it is not as stable as Adam.</p>
</li>
<li>
<p><b>Concurrency and automatic stopping</b>: The underlying sampler is multi-threaded (using <a href="https://isocpp.org/wiki/faq/cpp11-library-concurrency">C++11 threads</a>) with shared data. On top of the multi-threading, we have layered a convergence monitor in a separate thread that communicates with the chains through lock-free, latest-only, single-producer/single-consumer (SPSC) buffers (specifically, a <a href="https://en.wikipedia.org/wiki/Multiple_buffering#Triple_buffering">triple buffer</a>). The monitor automatically stops warmup when the mass matrices and step sizes have converged within tolerance to their cross-chain averages. The monitor automatically stops sampling when a target (traditional, non-split, non-ranked) R-hat; threshold is satisfied for the unnormalized log density, which typically converges more slowly than any of the individual parameters. It can also be configured to run for a fixed number of warmup and/or sampling iterations. The link between the original R-hat and effective sample size makes this essentially an unscaled ESS target.</p>
</li>
<li>
<p><b>Ragged chain summaries</b>: Asynchronous concurrent execution of chains with automatic stopping produces chains of different lengths. Because <a href="https://www.arviz.org/en/latest/">ArviZ</a> does not accept ragged chain input of this kind, we have included posterior analysis tools for means, variances/standard deviations, quantiles, traditional R-hat, effective sample size, and Monte Carlo standard error that work with ragged chains.</p>
</li>
<li>
<p><b>C++20</b>: Walnutpie is implemented in <a href="https://cppreference.com/cpp/20">C++20</a>. As a programming language type fanatic (I&#8217;ve written two books with &ldquo;type&rdquo; in the title!), I don&#8217;t know how I survived without <a href="https://en.cppreference.com/cpp/language/constraints">C++ concepts</a> before C++20.</p>
</li>
<li>
<p><b>ctypes FFI</b>: The foreign function interface in Python uses <a href="https://docs.python.org/3/library/ctypes.html">ctypes</a> rather than a higher-level interface, which sidesteps the requirement of <a href="https://en.wikipedia.org/wiki/Application_binary_interface">ABI compatibility</a> of C++ binaries.</p>
</li>
</ol>
<p><b>Developers</b></p>
<p>We would also like to welcome new developers who may want to get involved. There are already a stack of improvements we&#8217;d like to make, which we have enumerated on the GitHub issues.</p>
<p><b>Sources of algorithms</b></p>
<p>The Walnuts algorithm was a joint effort among Nawaf Bou-Rabee, Sifan Liu, Tore Kleppe, and Milo Marsden. Nutpie was developed by Adrian Seyboldt. Nuts, in the form used currently in Stan, was originally developed by Matt Hoffman and Andrew Gelman, then improved with multinomial sampling and mass matrix adaptation by Michael Betancourt. We haven&#8217;t yet added cross-chain adaptation as developed by Ben Bales, but the pieces are all in place to do so.</p>
<p>Brian Ward and I wrote all of the version 0.0.1 code with design advice and code review from Steve Bronder. Claude (the LLM) helped with code review and testing, but we didn&#8217;t use it to write the actual code (not out of principle, but because Brian and I both prefer the control of doing things manually).</p>
<p><b>Feedback</b></p>
<p>We would very much appreciate any feedback people have, including feedback on the documentation and ease of use, the source code, and performance.</p>
<p>We are happy to get feedback through issues on GitHub, through replies to this post, or through mail to one of the developers.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/04/walnutpie-version-0-0-1-released/feed/</wfw:commentRss>
			<slash:comments>6</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;In that era, undergrad males at Madison and elsewhere, had to take ROTC classes, and he kept intentionally failing them because he was very leftwing politically.&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/04/in-that-era-undergrad-males-at-madison-and-elsewhere-had-to-take-rotc-classes-and-he-kept-intentionally-failing-them-because-he-was-very-leftwing-politically/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/04/in-that-era-undergrad-males-at-madison-and-elsewhere-had-to-take-rotc-classes-and-he-kept-intentionally-failing-them-because-he-was-very-leftwing-politically/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 18:10:33 +0000</pubDate>
				<category><![CDATA[Art]]></category>
		<category><![CDATA[Political Science]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54370</guid>

					<description><![CDATA[Paul Alper writes: Regarding your blog of today, believe it or not, I knew Marshall Brickman pretty well before he became famous. He was an undergrad at the University of Wisconsin in Madison when I was a grad student there. &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/04/in-that-era-undergrad-males-at-madison-and-elsewhere-had-to-take-rotc-classes-and-he-kept-intentionally-failing-them-because-he-was-very-leftwing-politically/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Paul Alper writes:</p>
<blockquote><p>Regarding <a href="https://statmodeling.stat.columbia.edu/2026/08/04/53426/">your blog of today</a>, believe it or not, I knew Marshall Brickman pretty well before he became famous.  He was an undergrad at the University of Wisconsin in Madison when I was a grad student there.   We belonged to the same very left-wing eating co-op, &#8220;The Green Lantern.&#8221;  In that era, undergrad males at Madison and elsewhere, had to take ROTC classes, and he kept intentionally failing them because he was very leftwing politically.</p>
<p>Inasmuch as I am a helpful sort, I offered to nominate him for &#8220;Military Ball King,&#8221; but he declined.  Marshall was a really gifted musician as well as a writer.  In addition, he looked strikingly like the actor Richard Carlson who played Herb Philbrick in the 1950s television series I Led 3 Lives.</p></blockquote>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/04/in-that-era-undergrad-males-at-madison-and-elsewhere-had-to-take-rotc-classes-and-he-kept-intentionally-failing-them-because-he-was-very-leftwing-politically/feed/</wfw:commentRss>
			<slash:comments>3</slash:comments>
		
		
			</item>
		<item>
		<title>Stand-up comedy&#8212;like teaching, and book writing&#8212;requires &#8220;a collaboration with the audience.&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/04/53426/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/04/53426/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 13:05:35 +0000</pubDate>
				<category><![CDATA[Literature]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53426</guid>

					<description><![CDATA[In our bathroom we have this book, &#8220;And here&#8217;s the kicker: Conversations with 21 top humor writers on their craft,&#8221; edited by Mike Sacks. It&#8217;s well suited for the throne, as you can dip in and read as little or &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/04/53426/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>In our bathroom we have this book, &#8220;And here&#8217;s the kicker:  Conversations with 21 top humor writers on their craft,&#8221; edited by Mike Sacks.  It&#8217;s well suited for the throne, as you can dip in and read as little or as much as you want.</p>
<p>The gimmick is that you&#8217;d buy the book if you wanted to enter the comedy field.  Which doesn&#8217;t seem very likely, but, then again, who reads books either?  So maybe I&#8217;m one of the few people reading this one for entertainment value alone.</p>
<p>One of the interviews is with Woody Allen collaborator Marshall Brickman, coauthor of the 1970s classics Sleeper, Annie Hall, and Manhattan.  Lots of interesting stuff in this interview, including this:</p>
<blockquote><p><em>Marshall Brickman:</em>  [Woody] found a whole new area of insight:  relationships of a certain kind, psychoanalysis, and the creation of the so-called loser&#8212;mostly with women.  To some extent, the lower-with-women character was someone Bob Hope would play, but in a much more general and mainstream way.  Woody&#8217;s character was more ethnically and culturally specific. </p>
<p><em>Mike Sacks:</em>  It takes true genius to develop a comic character like that.</p>
<p><em>Brickman:</em>  It does, but it also requires a collaboration with the audience. It&#8217;s the only way you can do it.  You have to get out there and do a variety of material. Over times, certain things, statistically, will continue to work, and other things will drop away, and the audience will tell you what seems correct for you&#8212;for what you project onstage as a personality.</p>
<p><em>Sacks:</em>  But even with that said, you can work for twenty years and never connect with the audience half as much as Woody Allen.</p>
<p><em>Brickman:</em>  That&#8217;s right.  That&#8217;s the genius.  Creating something that somehow resonates with an audience that strongly.</p></blockquote>
<p>If only Allen had died in 1986 (after releasing Hannah and Her Sisters) or maybe 1992 (after Husbands and Wives), just think how high his reputation would be.  Even setting aside his personal life, releasing a series of flawed projects takes a toll.  Not everyone&#8217;s a Scorsese or Spielberg who can keep it up into their old age.  If Woody had gone in 1986, we&#8217;d still be speculating about the amazing work he didn&#8217;t have the opportunity to do.</p>
<p>But the thing that really interested me in the above-quoted interview is what Brickman said about collaboration with the audience.</p>
<p>This is true of teaching too!  Teaching a class, even writing a book, is a collaboration with the audience.  A book doesn&#8217;t read itself.  With books, the challenge is that the audience doesn&#8217;t see it until it comes out.  To continue the analogy, a live class is like stand-up; a book is like a movie.</p>
<p>When writing books, we always have the audience in my mind.  The trouble is that the book might never reach the audience.  That&#8217;s what I think has happened with <a href="https://sites.stat.columbia.edu/gelman/active-statistics/">Active Statistics</a>.  I&#8217;d like to rearrange it and rerelease it as a book called Statistics Stories.  Or maybe, Important Concepts of Statistics in Story Form. Also I wish we&#8217;d given <a href="https://sites.stat.columbia.edu/gelman/regression/">Regression and Other Stories</a> the more serious and descriptive title, Applied Regression and Causal Inference.  Maybe we can rerelease it under that title.</p>
<p>Anyway, yeah, &#8220;collaboration with the audience&#8221;:  definitely.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/04/53426/feed/</wfw:commentRss>
			<slash:comments>10</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;Placebo tests deserve a model, not just a glance.&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/03/placebo-tests-deserve-a-model-not-just-a-glance/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/03/placebo-tests-deserve-a-model-not-just-a-glance/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Mon, 03 Aug 2026 13:30:44 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Causal Inference]]></category>
		<category><![CDATA[Economics]]></category>
		<category><![CDATA[Sociology]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54327</guid>

					<description><![CDATA[Miha Gazvoda shares this post with the above title and the subtitle, &#8220;Using Bayesian multilevel models to correct bias and calibrate uncertainty.&#8221; He&#8217;s using the chickens model from our Slamming the Sham paper in the more general setting of placebo &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/03/placebo-tests-deserve-a-model-not-just-a-glance/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Miha Gazvoda <a href="https://mihagazvoda.com/posts/placebo-tests/">shares this post</a> with the above title and the subtitle, &#8220;Using Bayesian multilevel models to correct bias and calibrate uncertainty.&#8221;</p>
<p>He&#8217;s using the chickens model from our <a href="https://sites.stat.columbia.edu/gelman/research/published/chickens.pdf">Slamming the Sham paper</a> in the more general setting of placebo control tests.</p>
<p>In econometrics, a &#8220;placebo control test&#8221; does not need to literally involve a placebo treatment; it more generally is used to describe a procedure in which the same statistical analysis that was used to estimate a causal effect is applied to a different dataset, or a different part of the existing dataset, in which the treatment did not occur.</p>
<p>For example, if you want to measure the effect of an intervention that occurred in 2021, you could repeat the analysis but using data from a different year.  Or if you want to measure the effect on a particular outcome, you could repeat the analysis but looking at a different outcome that should be unaffected by the treatment.</p>
<p>The idea is that, if your estimation method has an artifact or systematic bias, this should show up in the placebo analysis as well, indicating a problem.  Conversely, if the placebo analysis does <em>not</em> show an effect, this is taken as evidence that there is no artifact.</p>
<p>In his post, Gazvoda argues that, rather than using the result of the placebo check to make a go/no-go decision, it should be possible to partially adjust the treatment effect to account for the information in that supplementary analysis.</p>
<p>This makes sense to me, and of course I&#8217;m happy that he&#8217;s using our chickens model.</p>
<p>There&#8217;s a tricky thing going on here with the placebo check, which, interestingly, arose in the chicken example too, and that is that we usually don&#8217;t have any good theory for where the effect is coming from in the placebo control.  After all, if our causal identification is working as designed, we shouldn&#8217;t even need the placebo comparison, as we&#8217;re already getting an unbiased estimate of the treatment effect.  The placebo control is typically there to address unspecified concerns of bias.  And, indeed, in practice, researchers don&#8217;t always do placebo controls.  And when, as hoped, the placebo control shows no statistically significant effect, the inclination is to take that as a reassurance and move on, in the same way that is done with other robustness checks.</p>
<p>The lesson Gasvoda takes from the chickens example is that if replications are available, you can assess the evidence for the placebo adjustments being relevant:  you can fit a multilevel model to estimate how much adjustment needs to be done.</p>
<p>From a sociology-of-science point of view, it&#8217;s interesting to me that the conventions in biomedical statistics and econometrics go in opposite directions:</p>
<p>&#8211; In biomedical statistics, the default recommended behavior is to compute the difference in differences, taking the estimated effect from the main experiment and subtracting the estimate from the placebo experiment.  As we explain in the chickens paper, this correction has the disadvantage of doubling the variance of the estimate, a true statistical crime if, as is often the case, the effect of the placebo treatment is indistinguishable from zero.</p>
<p>&#8211; In econometrics, the default procedure, if you&#8217;re calling it &#8220;difference-in-differences estimation,&#8221; is the same as above.  But if you&#8217;re calling it &#8220;placebo control,&#8221; and the placebo estimate is not statistically significant from zero, the default is to ignore the placebo results entirely, not to adjust for them.</p>
<p>In general we recommend a partial adjustment, with the amount of adjustment depending on the problem at hand.  If there is internal replication, as in the chickens example, the appropriate adjustment factor can be estimated from the data.  If it&#8217;s a one-shot experiment, you&#8217;ll need to use prior information.  I don&#8217;t have any good examples demonstrating how to do that; it&#8217;s something we should do.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/03/placebo-tests-deserve-a-model-not-just-a-glance/feed/</wfw:commentRss>
			<slash:comments>1</slash:comments>
		
		
			</item>
		<item>
		<title>What gets you is not what you don&#8217;t know but what you don&#8217;t know you don&#8217;t know.</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/02/what-gets-you-is-not-what-you-dont-know-but-what-you-dont-know-you-dont-know/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/02/what-gets-you-is-not-what-you-dont-know-but-what-you-dont-know-you-dont-know/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sun, 02 Aug 2026 13:14:23 +0000</pubDate>
				<category><![CDATA[Causal Inference]]></category>
		<category><![CDATA[Economics]]></category>
		<category><![CDATA[Literature]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Sociology]]></category>
		<category><![CDATA[Zombies]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=53415</guid>

					<description><![CDATA[I was thinking about the above saying in the context of bad regression discontinuity analyses. Statistical methods can be characterized in terms of how they can go wrong. Some common modes of &#8220;how can things go wrong&#8221; include: &#8211; biased &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/02/what-gets-you-is-not-what-you-dont-know-but-what-you-dont-know-you-dont-know/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>I was thinking about the above saying in the context of <a href="https://statmodeling.stat.columbia.edu/2021/03/11/regression-discontinuity-analysis-is-often-a-disaster-so-what-should-you-do-instead-do-we-just-give-up-on-the-whole-natural-experiment-idea-heres-my-recommendation/">bad regression discontinuity analyses</a>.</p>
<p>Statistical methods can be characterized in terms of how they can go wrong.  Some common modes of &#8220;how can things go wrong&#8221; include:<br />
&#8211; biased measurements,<br />
&#8211; selection bias into the observed dataset,<br />
&#8211; differences between treatment and control groups,<br />
&#8211; nonstationarity or, in general, differences between sample and population of interest,<br />
&#8211; biased or noisy statistical estimates,<br />
&#8211; missing data,<br />
&#8211; problems with functional forms (for example, using a linear model for a nonlinear relation, or not including important interactions),<br />
&#8211; problems with error terms (dependence, choice of distribution, etc.),<br />
&#8211; unmodeled spillovers, hierarchical structure, or other violations of causal model assumptions,<br />
&#8211; researcher degrees of freedom and forking paths,<br />
and lots more!</p>
<p>Regression discontinuity is used for observational studies, and the #1 thing that can go wrong in an observational study is #3 on the above list:  differences between treatment and control groups.</p>
<p>The problem is that when researchers perform regression discontinuity analysis, they focus on the problem with the functional form of the expected outcome given the one predictor, often ignoring other potential predictors even if they are screaming to be included (as with age in <a href="https://sites.stat.columbia.edu/gelman/research/published/causal_paths_3.pdf">this example</a>).  There&#8217;s all this obsessing over the functional form (and I guess Guido and I <a href="https://sites.stat.columbia.edu/gelman/research/published/2018_gelman_jbes.pdf">contributed</a> to this) that&#8217;s kind of missing the point (as <a href="https://sites.stat.columbia.edu/gelman/research/published/JCRE-2025-12-Gelman_and_Imbens.pdf">we discuss</a> briefly here).</p>
<p>So, the problem is not that these users of regression discontinuity don&#8217;t know what to do with other predictors; it&#8217;s that they don&#8217;t know that they don&#8217;t know this.</p>
<p>As Tolstoy might have said had he been a data analyst, Models that fit are all alike; every poorly-fitting model is poorly fitting in its own way.</p>
<p><strong>P.S.</strong>  <a href="https://quoteinvestigator.com/2018/11/18/know-trouble/">Here&#8217;s</a> some background from the Quote Investigator on various version of the saying that I used as the title of this post.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/02/what-gets-you-is-not-what-you-dont-know-but-what-you-dont-know-you-dont-know/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
		<item>
		<title>Why quantitative understanding of effect sizes matters, even if all you care about is the presence of the effect</title>
		<link>https://statmodeling.stat.columbia.edu/2026/08/01/how-can-we-train-researchers-and-consumers-of-research-to-put-numbers-in-perspective-as-a-matter-of-course/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/08/01/how-can-we-train-researchers-and-consumers-of-research-to-put-numbers-in-perspective-as-a-matter-of-course/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Sat, 01 Aug 2026 13:38:12 +0000</pubDate>
				<category><![CDATA[Causal Inference]]></category>
		<category><![CDATA[Miscellaneous Science]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<category><![CDATA[Political Science]]></category>
		<category><![CDATA[Teaching]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54283</guid>

					<description><![CDATA[In reaction to my article with Andy King proposing post-publication review, Dan &#8220;Fast and Frugal&#8221; Goldstein writes: Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/08/01/how-can-we-train-researchers-and-consumers-of-research-to-put-numbers-in-perspective-as-a-matter-of-course/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>In reaction to <a href="https://sites.stat.columbia.edu/gelman/media/andrew_gelman_andy_king_chronicle.pdf">my article with Andy King proposing post-publication review</a>, Dan &#8220;Fast and Frugal&#8221; Goldstein writes:</p>
<blockquote><p>Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that term. It was something I coined as a grad student.</p>
<p>In your proposal, only hit papers get audited. It reminds me a bit of Mel Brooks&#8217; The Producers in which the protagonists use the logic &#8220;who would audit a flop?&#8221; and stay under the radar by intentionally producing a bad show.  Fraudsters have likely attempted to make their work seem worthy of publication while ensuring it doesn&#8217;t attract too much attention. Just like in The Producers, though, this sometimes backfires.</p></blockquote>
<p>I don&#8217;t know about that!  My impression with fraudsters is that they think that fraud is normal science, perhaps out of some mixture of bad education in research methods, a view that &#8220;everybody does it,&#8221; and a general lack of understanding of how non-cheaters (like you and me!) think.  Think of people like Wansink who gave general advice to to the world on how to p-hack, or Gino and Ariely, who published papers on dishonesty, or Mary Rosh, who surely believes that whatever shady statistical manipulations she does are nothing compared to the dastardly deeds done by the Democrats.</p>
<p>I&#8217;m sure there&#8217;s tons of below-the-radar cheating and bad science that we don&#8217;t hear about, but a fair number of prominent science fraudsters seem to enjoy the limelight.  One reason for this seemingly self-sabotaging behavior, I think, is that cheating enabled these people to attain great professional success for years.  They had no reason to think the juice would stop flowing.</p>
<p>To return to my proposal with Andy King:  I think it&#8217;s ok that only the hit papers get audited.  Bad papers that get no intention aren&#8217;t doing much damage, right?</p>
<p>Goldstein adds:</p>
<blockquote><p>
By the way, I was just having a conversation about <a href="https://web.archive.org/web/20181212191002/https://www.washingtonpost.com/news/monkey-cage/wp/2015/05/20/fake-study-on-changing-attitudes-sometimes-a-claim-that-is-too-good-to-be-true-isnt/?utm_term=.f39cc6d12f9e">your sensing that something was amiss with the LaCour study</a>. For years now I have used this quote of yours in a talk I give about putting numbers into perspective. I argue that it&#8217;s really important that people learn how to put numbers into perspective because if they don&#8217;t, they won&#8217;t notice that something is unusual and worthy of a deeper audit. You somehow sensed something was up with the Lacour result. <a href="https://goodauthority.org/news/pushing-at-an-open-door-when-can-personal-stories-change-minds-on-gay-rights/">You didn&#8217;t think it was fraud yet but you knew it was strange</a> because you know how to put such differences into perspective:</p>
<blockquote><p>A difference of 0.8 on a five-point scale . . . wow! You rarely see this sort of thing. Just do the math. On a 1-5 scale, the maximum theoretically possible change would be 4. But, considering that lots of people are already at “4” or “5” on the scale, it’s hard to imagine an average change of more than 2. And that would be massive. So we’re talking about a causal effect that’s a full 40% of what is pretty much the maximum change imaginable. Wow, indeed. And, judging by the small standard errors (again, see the graphs above), these effects are real, not obtained by capitalizing on chance or the statistical significance filter or anything like that.</p></blockquote>
</blockquote>
<p>My colleagues and I recently wrote <a href="https://sites.stat.columbia.edu/gelman/research/unpublished/hypothesizing_effect_size.pdf">a paper on this general topic of average effect sizes</a>.  It&#8217;s our contention that people generally are way too optimistic about possible effect sizes, in large part because they don&#8217;t think about variation.  If you ask someone to hypothesize an effect size, you&#8217;ll typically get a guess of the largest effect that might occur.</p>
<p>But what if you don&#8217;t really care about effect size&#8211;you just want to know about the effect?</p>
<p>For example, maybe you don&#8217;t believe that women during certain times of the month are three times more likely to <a href="https://web.archive.org/web/20260629132723/https://slate.com/technology/2013/07/statistics-and-psychology-multiple-comparisons-give-spurious-results.html">wear red</a> or pink shirts, but you are interested in some sort of evolutionary psychology theory of sexual display.  In that case, why should the effect size matter?  Why care that a study reported an estimate that was ridiculously implausible?</p>
<p>I have two to this questions, and thus two reasons why effect size is important even for problems where you don&#8217;t directly care about effect sizes:</p>
<p>1.  Effect sizes vary.  An treatment that has an effect (that is, a true effect, not just an estimated effect) of 0.1 for one group of people in one setting could have an effect of -0.2 in some other scenario.  A treatment effect in an experiment is the sum of all sorts of things, positive and negative, and there&#8217;s no logical reason to think the sign of the effect will be preserved.  Effect size matters.  The issue is not just that a smaller and more realistic effect size is less important; it&#8217;s also that smaller effects can be more easily produced by other factors, and this reduces the generality of any claims, even if the experiment at hand was done well.</p>
<p>2.  Experiments produce standard errors as well as estimates.  If the standard error from a study is large compared to any realistic effect size, then the study contains very little information.  Effect size is important in understanding the informativeness of an experiment, and to do this right you need to have some sense of what the true effect size could be.  You can&#8217;t just use a point estimate from the study itself, as this estimate will inherently be too noisy to use to judge the information in the study.  As I wrote <a href="https://sites.stat.columbia.edu/gelman/research/published/power_surgery_new_response.pdf">in this note</a> for the Annals of Surgery, Post-hoc power using observed estimate of effect size is too noisy to be useful.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/08/01/how-can-we-train-researchers-and-consumers-of-research-to-put-numbers-in-perspective-as-a-matter-of-course/feed/</wfw:commentRss>
			<slash:comments>16</slash:comments>
		
		
			</item>
		<item>
		<title>What do we learn from bestseller regressions?</title>
		<link>https://statmodeling.stat.columbia.edu/2026/07/31/what-do-we-learn-from-bestseller-regressions/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/07/31/what-do-we-learn-from-bestseller-regressions/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Fri, 31 Jul 2026 13:38:57 +0000</pubDate>
				<category><![CDATA[Economics]]></category>
		<category><![CDATA[Literature]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54254</guid>

					<description><![CDATA[Gaurav Sood writes: I was reading &#8216;The Bestseller Code.&#8217; The book reports results from some regressions of the form: bestseller or not ~ features of content This got me thinking about what you can recover from such an exercise. Say &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/07/31/what-do-we-learn-from-bestseller-regressions/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Gaurav Sood writes:</p>
<blockquote><p>I was reading &#8216;The Bestseller Code.&#8217; The book reports results from some regressions of the form:</p>
<p>bestseller or not ~ features of content</p>
<p>This got me thinking about what you can recover from such an exercise. </p>
<p>Say that there are two types of novels: type x and y. Let&#8217;s say that people prefer reading novels of type x. They are k% more likely to read a novel of type x than y.</p>
<p>At time t, the total number of novels is T, with novels of type x being 50%. And we can recover k% by regressing the number of people who read a novel on the type of novel.</p>
<p>Reader preference for type x novels leads writers to produce more novels of type x. At time t+1, the percentage of type of novels x being produced is 70%. But say that the total number of novels of type x that people can read is fixed at n_t. When we regress the number of people who read a novel on the type of novel, we can get that people are less likely to read novels of type x than y because we have more novels of type x at t+1.</p>
<p>I wrote <a href="https://www.gojiberries.io/what-do-we-learn-from-bestseller-regressions/">a brief post</a> based on the point here.</p></blockquote>
<p>I agree, and this sort of thing has always puzzled me. I&#8217;m sure economists have looked into the matter, but I don&#8217;t know the literature so I&#8217;m just guessing here.  Whenever people talk about the relative profitability of different genres, I wonder about what happens when the more profitable genre gets flooded with content.  That doesn&#8217;t mean that bestseller regressions are useless, just that they represent at best some equilibrium of a process that I don&#8217;t understand very well.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/07/31/what-do-we-learn-from-bestseller-regressions/feed/</wfw:commentRss>
			<slash:comments>7</slash:comments>
		
		
			</item>
		<item>
		<title>Posterior predictive checking is for non-Bayesians too!</title>
		<link>https://statmodeling.stat.columbia.edu/2026/07/30/posterior-predictive-checking-is-for-non-bayesians-too/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/07/30/posterior-predictive-checking-is-for-non-bayesians-too/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Thu, 30 Jul 2026 13:58:07 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Miscellaneous Statistics]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=51198</guid>

					<description><![CDATA[When I first started working on posterior predictive checking back in 1988, it was as a device for determining equivalent degrees of freedom for a chi-squared test for a model with constrained parameters&#8211;in that case, positivity restrictions in an image &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/07/30/posterior-predictive-checking-is-for-non-bayesians-too/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>When I first started working on posterior predictive checking back in 1988, it was as a device for determining equivalent degrees of freedom for a chi-squared test for a model with constrained parameters&#8211;in that case, positivity restrictions in an image reconstruction problem; <a href="https://stat.columbia.edu/~gelman/research/published/phd_thesis.pdf">see here for background</a>.</p>
<p>The idea is that the distribution of the test statistic depends on the true parameter vector&#8211;there is no simple pivotal quantity as would exist under a linear model with no constraints&#8211;but you can work out the distribution conditional on the true parameter vector, and then you can average this distribution over the posterior for the parameters.  </p>
<p>You can think of this as Bayesian&#8211;the marginal posterior distribution of the test statistic&#8211;but at the time I was thinking of it more as a generalization of the existing &#8220;plug-in&#8221; approach that would use a point estimate of the parameter vector.  From that perspective, posterior averaging is a technique for getting a distribution with better frequency properties than you&#8217;d get from the maximum likelihood estimate, and that&#8217;s because in an image reconstruction problem with positivity constraints and lots and lots of pixels, the maximum likelihood estimate will almost certainly be on the boundary of parameter space.  So just about any sort of averaging should get you closer to the true parameter value.</p>
<p>As noted above, I started working in this area in 1988.  Around 1991 I got my thoughts organized enough to write a paper and give a talk on the topic, and . . . it got a generally hostile reception!  The Bayesians didn&#8217;t like it because they didn&#8217;t like chi-squared tests or frequentist hypothesis testing more generally.  They wanted me to do Bayes factors, which even then I realized were generally a bad idea (for more on the topic, see chapters 6 and 7 of BDA3 (it was all in chapter 6 of the first two editions) or <a href="https://stat.columbia.edu/~gelman/research/published/avoiding.pdf">this article</a> from 1995).  The non-Bayesians didn&#8217;t like it because it was Bayesian, and because I just presented the method, without trying to justify it based on minimax properties or whatever.</p>
<p>Xiao-Li, Hal, and I finally published <a href="https://stat.columbia.edu/~gelman/research/published/A6n41.pdf">a version of the article</a> a few years later, and I resigned myself to the fact that posterior predictive checking was going to remain in the Bayesian world.  Predictive checking was a hard sell to the Bayesians, but, after a few decades of exposure to BDA and lots of applied examples, they started to get used to the idea.  I gave up on trying to promote the idea outside the Bayesian community.</p>
<p>But . . . it turns out that posterior predictive checking <em>has</em> been taken up for non-Bayesian uses!  Aki pointed me to <a href="https://easystats.github.io/performance/reference/check_predictions.html">this R package called &#8220;performance&#8221;</a> that &#8220;provides posterior predictive check methods for a variety of frequentist models.&#8221;  <a href="https://easystats.github.io/performance/articles/check_model.html">Here&#8217;s their vignette</a> on checking model assumption &#8211; linear models,&#8221; and <a href="https://joss.theoj.org/papers/10.21105/joss.03139">here&#8217;s the article</a>, performance: An R Package for Assessment, Comparison and Testing of Statistical Models, by Daniel Lüdecke, Mattan Ben-Shachar, Indrajeet Patil, Philip Waggoner, and Dominique Makowski.  It came out in 2021 and has already been cited <a href="https://scholar.google.com/scholar?hl=fr&#038;as_sdt=0%2C33&#038;q=performance%3A+An+R+Package+for+Assessment%2C+Comparison+and+Testing+of+Statistical+Models&#038;btnG=">over 2500 times</a>.</p>
<p>So it looks like people really are using posterior predictive checks for non-Bayesian models.  And it took less than 40 years for it to happen!</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/07/30/posterior-predictive-checking-is-for-non-bayesians-too/feed/</wfw:commentRss>
			<slash:comments>14</slash:comments>
		
		
			</item>
		<item>
		<title>&#8220;Over-coverage caught by pre-registration: 47 of 56 inside a stated 50% interval&#8221;</title>
		<link>https://statmodeling.stat.columbia.edu/2026/07/29/over-coverage-caught-by-pre-registration-47-of-56-inside-a-stated-50-interval/</link>
					<comments>https://statmodeling.stat.columbia.edu/2026/07/29/over-coverage-caught-by-pre-registration-47-of-56-inside-a-stated-50-interval/#comments</comments>
		
		<dc:creator><![CDATA[Andrew]]></dc:creator>
		<pubDate>Wed, 29 Jul 2026 21:12:27 +0000</pubDate>
				<category><![CDATA[Bayesian Statistics]]></category>
		<category><![CDATA[Economics]]></category>
		<guid isPermaLink="false">https://statmodeling.stat.columbia.edu/?p=54233</guid>

					<description><![CDATA[Alex Malinowski has a question about evaluating the calibration of interval forecasts: We publish interval forecasts under a pre-registration scheme: each forecast is serialised, hashed and timestamped into a Bitcoin block before publication, so the stated interval cannot be adjusted &#8230; <a href="https://statmodeling.stat.columbia.edu/2026/07/29/over-coverage-caught-by-pre-registration-47-of-56-inside-a-stated-50-interval/">Continue reading <span class="meta-nav">&#8594;</span></a>]]></description>
										<content:encoded><![CDATA[<p>Alex Malinowski has a question about evaluating the calibration of interval forecasts:</p>
<blockquote><p>
We publish interval forecasts under a pre-registration scheme: each forecast is serialised, hashed and timestamped into a Bitcoin block before publication, so the stated interval cannot be adjusted after the outcome is known. Across 56 resolved forecasts our stated 50% intervals contained 47 outcomes; at the 24-hour horizon, 36 of 44. Under honest 50% intervals, P(>=36 of 44) is about 1.3e-05.</p>
<p>The diagnosis, and the part I would most like criticised: walk-forward across 20,058 observations showed the miscalibration was conditional rather than uniform &#8211; 55.6% coverage across all days, 67.1% restricted to the calmest fifth. Our first hypothesis, that the sample carried too much old high-volatility history, was tested and rejected. The surviving explanation is that interval width ignored the current volatility regime. Conditioning the historical sample on regime measured at each window&#8217;s open (not its close, which leaks the outcome) brings coverage to 50.8%.</p>
<p>Two things I am unsure about. The instruments are correlated, so effective sample size is well below 56 and I have not done that properly. And at a 30-day horizon the conditional interval comes out wider than the unconditional one, which I have kept but cannot fully account for.</p></blockquote>
<p>I asked him what was the application, and he replied:</p>
<blockquote><p>Crypto prices. 24-hour and 7-day intervals on eight pairs (BTC, ETH, SOL, BNB, XRP, DOGE, ADA, LINK), scored against Binance closes.</p>
<p>I chose it as the substrate rather than the subject. Outcomes resolve within a day, the reference price is unambiguous, and there is no data-vintage or revision problem, so a coverage check accumulates evidence quickly and cheaply. The same construction is what we use for sports and prediction-market questions, but those resolve slowly and the sample is thin.</p>
<p>I realise crypto invites a certain reaction, and the reaction is mostly deserved. The calibration question doesn&#8217;t depend on it &#8211; the same test applies to any published range.</p></blockquote>
<p>He adds:</p>
<blockquote><p>If it helps frame it, the two things I&#8217;m least confident about are the ones I&#8217;d want readers to attack:</p>
<p>1. The eight instruments are correlated, so the effective sample size is well below 56 and I have not handled that properly. The binomial p-value I quoted assumes independence it doesn&#8217;t have.</p>
<p>2. At a 30-day horizon the regime-conditional interval comes out wider than the unconditional one. I kept the result because it&#8217;s inconvenient, but I can&#8217;t fully account for it beyond &#8220;calm periods have historically preceded larger monthly moves&#8221;, which feels like a description rather than an explanation.</p>
<p>Data and outcomes are CC-BY, one row per sealed forecast with its hash:<br />
https://huggingface.co/datasets/neuportal/neuportal-sealed-crypto-forecasts</p></blockquote>
<p>My only quick thought here is that it&#8217;s indeed a good idea to look at conditional calibration, but you have to be careful only to condition on things in the forecast, not on the outcome.  So when he writes, &#8220;67.1% restricted to the calmest fifth,&#8221; this is relevant as long as &#8220;calmest fifth&#8221; is something that can be determined before the outcome occurs.</p>
<p>More generally, yeah, it&#8217;s hard to evaluate the calibration of forecasts for correlated outcomes&#8211;this comes up in election prediction too.  My only answer is to try to model intermediate outcomes as much as possible to avoid getting stuck in a purely empirical mode of evaluations.</p>
<p>Maybe you in comments have other ideas?</p>
]]></content:encoded>
					
					<wfw:commentRss>https://statmodeling.stat.columbia.edu/2026/07/29/over-coverage-caught-by-pre-registration-47-of-56-inside-a-stated-50-interval/feed/</wfw:commentRss>
			<slash:comments>2</slash:comments>
		
		
			</item>
	</channel>
</rss>
